Data processing method and task platform

By obtaining the conformations of the ligands to be screened and the target protein, determining the binding channels and calculating the binding weights, the problem of false positive screening in the drug discovery process is solved, and accurate measurement and screening of ligand-protein binding is achieved.

CN121709012APending Publication Date: 2026-03-20ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411321831.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing computational methods are unable to effectively screen and filter false positive results in the drug discovery process, resulting in a large number of small molecules that can theoretically bind to proteins but cannot actually enter the protein interior.

Method used

By acquiring the conformations of the ligands to be screened and the target protein, the binding channels are determined, the channel energy and attribute information are calculated, the target binding channels are selected according to the binding weights, and the ease of binding between the ligands and proteins is screened using channel trajectory and energy analysis.

Benefits of technology

It provides an indicator to measure the ease with which proteins and ligands bind, effectively screening and filtering false positive results, and improving the accuracy of the drug discovery process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121709012A_ABST
    Figure CN121709012A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and a task platform. The data processing method comprises the following steps: acquiring a ligand to be screened and a target protein conformation; determining at least one binding channel corresponding to the target protein conformation; calculating channel energy information of each combination channel corresponding to the ligand to be screened, and obtaining channel attribute information corresponding to each combination channel; determining a combination weight corresponding to each combination channel according to the channel energy information and the channel attribute information of each combination channel; and selecting a target binding channel corresponding to the to-be-screened ligand and the target protein conformation according to each binding weight. According to the method, an index for measuring the binding difficulty degree of the protein and the ligand is obtained by utilizing the results of channel analysis and energy analysis through a probability statistics method. And false sun conditions can be effectively screened and filtered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of computer technology, and in particular to a data processing method. Background Technology

[0002] With the development of computational chemistry and bioinformatics technologies, virtual screening methods are being used more and more widely in drug development. Current computational methods for drug discovery focus on the final binding state of drugs and protein targets. However, this computational process only calculates whether small molecules and proteins can bind, without considering the binding process itself. This leads to a large number of false positives in computational screening results; that is, although small molecules and proteins can theoretically bind, in reality, the small molecules cannot enter the protein and bind.

[0003] Therefore, how to effectively screen and filter out false positives during the drug discovery process using computational methods has become an urgent problem for technical personnel. Summary of the Invention In view of this, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a task platform, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0004] According to a first aspect of the embodiments of this specification, a data processing method is provided, comprising: Obtain the conformation of the ligand to be screened and the target protein; Identify at least one binding channel corresponding to the conformation of the target protein; Calculate the channel energy information of each binding channel and the ligand to be screened, and obtain the channel attribute information of each binding channel; Based on the channel energy information and channel attribute information of each binding channel, determine the binding weight corresponding to each binding channel; The target binding channels corresponding to the conformation of the target protein and the ligand to be screened are selected based on the binding weights.

[0005] According to a second aspect of the embodiments of this specification, a data processing method is provided, applied to a cloud-side device, comprising: The receiving end device sends the conformation of the ligand to be screened and the target protein; Identify at least one binding channel corresponding to the conformation of the target protein; Calculate the channel energy information of each binding channel and the ligand to be screened, and obtain the channel attribute information of each binding channel; Based on the channel energy information and channel attribute information of each binding channel, determine the binding weight corresponding to each binding channel; The target binding channels corresponding to the conformation of the target protein and the ligand to be screened are selected according to each binding weight; The target is sent to the end-side device via the combined channel.

[0006] According to a third aspect of the embodiments of this specification, a task platform is provided, including a request interface and a response unit; The request interface is used to receive the ligand to be screened and the conformation of the target protein sent by the end device; The response unit is configured to: determine at least one binding channel corresponding to the conformation of the target protein; calculate the channel energy information of each binding channel and the ligand to be screened, and obtain the channel attribute information corresponding to each binding channel; determine the binding weight of each binding channel based on the channel energy information and channel attribute information of each binding channel; and select the target binding channel corresponding to the conformation of the target protein and the ligand to be screened based on the binding weight.

[0007] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above-described data processing method.

[0008] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the data processing method described above.

[0009] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the data processing method described above.

[0010] An embodiment of this specification provides a data processing method including: acquiring a ligand to be screened and a target protein conformation; determining at least one binding channel corresponding to the target protein conformation; calculating the channel energy information of each binding channel and the ligand to be screened, and acquiring the channel attribute information corresponding to each binding channel; determining the binding weight of each binding channel based on the channel energy information and channel attribute information; and selecting the target binding channel corresponding to the ligand to be screened and the target protein conformation based on the binding weight.

[0011] The methods provided in the embodiments of this specification analyze the energy changes of the ligand and the channel from the protein binding site to the core, thus measuring the ease of protein-ligand binding to a certain extent from the perspective of channel and energy. Using probabilistic statistical methods and the results of channel and energy analysis, an index for measuring the ease of protein-ligand binding is obtained. This effectively screens and filters false positives. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating a data processing method provided in one embodiment of this specification; Figure 2 This is a flowchart illustrating the data processing method applied to cloud-side devices according to one embodiment of this specification. Figure 3 This is a schematic diagram of the structure of a data processing apparatus provided in one embodiment of this specification; Figure 4 This is an architecture diagram of a data processing system provided in one embodiment of this specification; Figure 5 This is a schematic diagram of a task platform provided in one embodiment of this specification; Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0016] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0017] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0018] Drug discovery: This is a complex, multi-stage process aimed at identifying molecules with potential characteristics for developing effective drugs. In the early stages of this process, multiple steps are typically involved, some of which include target identification (identifying the biological target of the disease or health problem that needs to be treated; biological targets are usually proteins or molecules), high-throughput screening (testing a large number of compounds through large-scale experiments to identify those molecules that exhibit the desired pharmacological properties when interacting with the target), and compound optimization (once potential drug candidates are found, further improvements are made to enhance their efficiency, selectivity, and bioavailability).

[0019] Molecular docking: a computational chemistry approach used to predict how small molecules (often drug molecules) interact with biomolecules (often proteins) to understand their binding patterns and affinities.

[0020] Virtual screening: By computationally simulating the binding of a large number of compounds, the most promising drug candidates can be screened to reduce the number of laboratory experiments. Several software programs are available for this simulation, such as Auto Dock Vina and VinaGPU, which utilizes GPUs (Graphics Processing Units) for acceleration.

[0021] Wet experiments: These are actual experiments conducted in the laboratory using real substances (such as chemical reagents, biological specimens, or organisms) to detect the activity of small molecules and protein pocket binding in drug discovery.

[0022] In drug development, small molecules (ligands) interact with proteins at their active sites or binding sites (functional sites of proteins). The simulation of small molecule binding (entering the active site and forming a stable complex) and unbinding (releasing the ligand from the stable complex) is helpful for many practical applications. It allows for the search for ligands more likely to bind to specific proteins. Ligands can be modified to make their binding faster or more affinity-enhancing, or proteins can be modified to reduce or prevent ligand binding. Many proteins have binding sites embedded in their core, meaning that ligands must pass through a channel or tunnel before binding to the protein's functional site.

[0023] In this context, the method provided in the embodiments of this specification analyzes whether a ligand can potentially enter the protein core through tunnels or channels. Analyzing the channel and energy changes from the binding site to the complex surface can, to some extent, measure the ease with which a protein and ligand bind. Based on this, this specification provides a data processing method, and also relates to a data processing apparatus, a task platform, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0024] See Figure 1 , Figure 1 A flowchart of a data processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0025] Step 102: Obtain the conformation of the ligand to be screened and the target protein.

[0026] In the biopharmaceutical field, ligands are generally small molecules, ions, or biomolecules that can specifically bind to biological macromolecules (such as proteins and nucleic acids). Ligands serve as important components of drug targets or drug molecules during drug development. By specifically binding to target proteins, ligands can modulate the function of those proteins, thereby exerting therapeutic effects. Utilizing ligands in drug design allows for the prediction of drug-target interactions through techniques such as computer simulations and molecular docking, accelerating the drug development process.

[0027] Protein conformation refers to the arrangement of all atoms in a protein molecule in three-dimensional space, which determines the protein's function and properties. A protein's function is closely related to its conformation. The primary structure of a protein determines the formation of its secondary, tertiary, and quaternary structures, which in turn determine its function. Different conformational states may correspond to different biological activities or functional states.

[0028] There is a close relationship between ligands and protein conformation. Ligands can specifically bind to specific binding sites on proteins (also known as active pockets or ligand-binding pockets). By binding to proteins, ligands can regulate protein function. For example, some enzymes undergo conformational changes after binding to ligands, thereby activating or inhibiting their catalytic activity. Transport proteins undergo conformational changes after binding to specific ions or small molecules, thereby enabling transmembrane transport of substances.

[0029] In the methods provided in the embodiments of this specification, the ligands to be screened and the conformations of the target proteins obtained are determined through computer simulation calculations to identify ligand-protein pairs that can bind. That is, the ligands to be screened and the conformations of the target proteins are theoretically capable of binding. The method provided in the embodiments of this specification aims to provide an index for evaluating the ease with which the ligands to be screened and the conformations of the target proteins bind.

[0030] In one specific embodiment provided in this specification, obtaining the conformation of the ligand to be screened and the target protein includes: Obtain at least one ligand to be screened and the conformation of the target protein.

[0031] In practical applications, the number of ligands to be screened can be one, two, or more. Specifically, during computer simulations, multiple ligands may theoretically bind to the target protein. However, in practice, some ligands cannot enter the target protein, or some ligands have difficulty entering the protein, resulting in ligands failing to bind to the target protein in practice. Therefore, the method described in the embodiments of this specification provides a way to evaluate the ease with which each ligand to be screened binds to the conformation of the target protein.

[0032] In practical applications, the methods provided in this application are all implemented on the computer side. All ligands, protein conformations, etc. can be understood as the electronic information corresponding to the ligands, protein conformations, etc.

[0033] Step 104: Determine at least one binding channel corresponding to the conformation of the target protein.

[0034] After determining the conformation of the target protein, at least one binding channel is identified within that conformation. When ligands need to bind to the functional sites of a protein, they typically cannot jump directly from the protein surface to the core region. They need to pass through a series of specific channels or pathways, which are determined by the protein's three-dimensional structure. The channels or pathways from the protein surface to the core region are called binding channels.

[0035] Binding channels can include structures such as cracks, grooves, and pockets in proteins, which provide ligands with gateways to the protein's core region.

[0036] In practical applications, there are many ways to determine at least one binding channel corresponding to the conformation of a target protein. One can search from the core region to the surface according to a preset search rule (such as depth search, breadth search, etc.) to determine at least one binding channel.

[0037] In one specific embodiment provided in this specification, at least one binding channel corresponding to a target protein conformation is determined by the following method. All atoms of the target protein are abstracted as a set of spheres, denoted as the representation of input structure (RIS). The vertices and edges of a Voronoi diagram are constructed using Delaunay triangulation at the center of the RIS. The Voronoi vertices of the binding channel are found based on the determined starting point. A probe is used to locate the Voronoi vertices near the channel exit on the complex surface. The axis is identified as a simple path along the path composed of Voronoi vertices and edges. A binding channel is determined by a sphere centered on the axis and having a maximum radius that does not collide with the RIS, which moves along the central axis.

[0038] The A* Search Algorithm is used to calculate the minimum cost path between the starting Voronoi vertex and the exit Voronoi vertex of the channel, i.e., the channel's central axis. The heuristic function can be (Euclidean distance, average channel radius). The channel's cost and throughput are calculated using Equations 1 and 2 below.

[0039] Formula 1 Formula 2 Here, r(l) defines the radius of the largest sphere that does not collide with RIS. The parameter n controls the balance between width and length; the larger n is, the heavier the penalty on width. min Represents the bottleneck radius. The penalty imposed by m on the bottleneck radius is controlled by m; the smaller m is, The larger the value, the greater the penalty on the bottleneck radius. "Cost" represents cost, and "throughput" represents throughput.

[0040] The detected channels are clustered (by calculating the similarity of the central axis, distance, or other clustering metrics), redundant channels are removed, and the combined channels with the lowest cost in each cluster are selected and retained. The channel information corresponding to the combined channels is saved, and the combined channels with a throughput greater than a preset throughput threshold are identified from all clustering results for subsequent trajectory analysis.

[0041] Step 106: Calculate the channel energy information of each binding channel and the ligand to be screened, and obtain the channel attribute information corresponding to each binding channel.

[0042] After identifying each binding channel, it is necessary to further analyze the trajectory and energy of the ligand to be screened as it travels through the binding channel to the surface of the complex. Specifically, the channel energy required for the ligand to be screened to reach the surface of the complex through each binding channel is calculated, and the channel attribute information corresponding to each binding channel is obtained.

[0043] Channel energy information can be understood as the energy required for a ligand to be screened to reach the surface of the complex through a specific binding channel, such as the binding energy of the ligand on the protein surface or the binding energy of the ligand located within the active site. Channel attribute information can be understood as the attribute information of a specific binding channel, such as channel length and width.

[0044] In one specific embodiment provided in this specification, the calculation of channel energy information corresponding to each binding channel and the ligand to be screened includes: Select a combination channel to be processed, wherein the combination channel to be processed is any one of the combination channels; Calculate the target channel energy information corresponding to the ligand to be screened and the binding channel to be processed.

[0045] In this embodiment, one of the bonding channels is used as an example for explanation. The same processing method is used for the other bonding channels, and will not be repeated in this embodiment.

[0046] Select one binding channel to be processed from each binding channel. The binding channel to be processed can be any one of the binding channels, and each binding channel is independent of the others.

[0047] After identifying the binding channel to be processed, the target channel energy information of the ligand to be screened and the binding channel to be processed is calculated.

[0048] Specifically, calculating the target channel energy information corresponding to the ligand to be screened and the binding channel to be processed includes: Calculate the protein surface binding energy, active site binding energy, and trajectory maximum binding energy corresponding to the ligand to be screened and the binding channel to be treated; Calculate the product dissociation activation energy and reactant binding activation energy based on the protein surface binding energy, the active site binding energy, and the trajectory maximum binding energy. The highest binding energy of the trajectory, the product dissociation activation energy, and the reactant binding activation energy are determined as the target channel energy information.

[0049] In the embodiments provided in this specification, the channel to be processed is started from the starting point and fixed at equal distances d on the central axis to obtain a series of n points, α1, α2...αn. Then, the channel is cut at these points along the vertical direction of the central axis to obtain n tangent circles, denoted as θ1, θ2...θn.

[0050] Select ligand atom a from the ligand molecule λ to be screened. c , where a c For each ∈λ, a point is sampled on each tangent circle. The ligand molecule is placed on the tangent circle, and the ligand atom is placed on the sampling point. The sampling follows the following distribution: a Gaussian distribution with the center of the circle as the mean and the standard deviation being half the radius of the circle. If the sampling point is not inside the tangent circle, it is resampled until it is inside the tangent circle.

[0051] Using Autodock Vina, molecular docking was performed on molecules on each cross-section circle to find the lowest energy point. The conformation at that point was saved, and the energy was recorded to obtain the molecular trajectory and energy analysis of the ligand from the binding site to the complex surface. The binding energy E of the active site was obtained from the analysis results. Bound Highest binding energy of trajectory E Max Protein surface binding energy E Surface Among them, the binding energy E of the active site Bound The trajectory represents the binding energy of the ligand located within the active site, with the highest binding energy E. Max The protein surface binding energy E represents the highest binding energy in the trajectory. Surface Represents the binding energy of ligands on the protein surface.

[0052] After determining the protein surface binding energy, active site binding energy, and trajectory maximum binding energy, the product dissociation activation energy E can be further calculated. off Activation energy E for binding with reactants on Specifically, the product dissociation activation energy E off The highest binding energy E of the trajectory Max Binding energy E to active site Bound The difference is the activation energy E of reactant binding. on The highest binding energy E of the trajectory Max Protein surface binding energy E Surface The difference is the activation energy E of product dissociation. off Activation energy E for binding with reactants on This corresponds to the kinetics of the ligand to be screened crossing the binding channel.

[0053] Finally, the product is dissociated with activation energy E. off Activation energy E of reactant binding on The highest binding energy E of the trajectory MaxThe target channel energy information is used to determine the binding energy of the ligand to be screened and the channel to be processed.

[0054] In one specific embodiment provided in this specification, the channel attribute information corresponding to each combined channel is obtained, including: Select a combination channel to be processed, wherein the combination channel to be processed is any one of the combination channels; Obtain the channel length information, average channel radius, and channel bottleneck radius of the combined channel to be processed.

[0055] In one specific embodiment provided in this specification, the channel attribute information corresponding to each combined channel will also be acquired simultaneously. In this embodiment, we will still use one of the combined channels as an example for explanation.

[0056] Select one channel from the various bonding channels to be processed, and statistically analyze its channel length, average radius, and bottleneck radius. The channel length can be understood as the length of the bonding channel. The average radius can be understood as the average radius obtained by averaging the radii corresponding to the length of the bonding channel. The bottleneck radius can be understood as the minimum radius within the bonding channel to be processed.

[0057] Step 108: Determine the binding weight corresponding to each binding channel based on the channel energy information and channel attribute information of each binding channel.

[0058] After obtaining the channel energy and attribute information of each binding channel, the binding weight of the ligand to be screened through each binding channel can be determined by combining the channel energy and attribute information. In simple terms, taking channel attribute information as an example, the shorter the channel length, the better; the larger the average channel radius and the channel bottleneck radius, the better. That is, the shorter the channel and the larger the channel radius, the easier it is for the ligand to be screened to pass through the binding channel and enter the protein core more easily.

[0059] In one specific embodiment provided in this specification, the binding weight corresponding to each binding channel is determined based on the channel energy information and channel attribute information of each binding channel, including: Obtain a preset protein ligand calculation dataset, and obtain the reference channel energy information and reference channel attribute information corresponding to the preset protein ligand calculation dataset; Select a combination channel to be processed, wherein the combination channel to be processed is any one of the combination channels; Based on the energy information and attribute information of the channel to be processed, and the energy information and attribute information of the reference channel, the target binding weight corresponding to the channel to be processed is determined.

[0060] In practical applications, determining the binding weights of each binding channel requires using a pre-calculated protein ligand calculation dataset. This dataset is a validated dataset of ligand-protein bindings; for example, the CoreSet2016 dataset could be used. This dataset stores binding information for multiple ligands and proteins, such as the reference channel the ligand enters, reference channel attribute information, and reference channel energy information. In the method provided in the embodiments of this specification, channel energy information and channel attribute information are used for reference comparison. Therefore, after obtaining the pre-calculated protein ligand dataset, the reference channel energy information and reference channel attribute information corresponding to each protein and ligand stored in the dataset are further obtained.

[0061] Using the reference channel energy information and reference channel attribute information as a benchmark, and the channel energy information and channel attribute information of the channel to be processed as the comparison content, the two are compared and calculated separately to further determine the target binding weight corresponding to the channel to be processed.

[0062] Specifically, based on the energy information and attribute information of the channel to be processed, and the energy information and attribute information of the reference channel, the target binding weight corresponding to the channel to be processed is determined, including: The channel energy weights are determined based on the energy information of the channel to be processed and the energy information of the reference channel. The channel attribute weights are determined based on the channel attribute information to be processed and the reference channel attribute information; The target binding weight corresponding to the binding channel to be processed is determined based on the channel energy weight and the channel attribute weight.

[0063] During the calculation of the dataset between the target binding channel and the preset protein ligand, the energy information and attribute information of the target binding channel are compared with the energy information and attribute information of the reference channel. Furthermore, the channel energy information is compared with the reference channel energy information to obtain the channel energy weight; the channel attribute information is compared with the reference channel attribute information to obtain the channel attribute weight. Finally, the target binding weight corresponding to the target binding channel is obtained by combining the channel energy weight and the channel attribute weight.

[0064] The following sections will further explain how to obtain channel energy weights and channel attribute weights.

[0065] In one specific embodiment provided in this specification, the channel energy information includes the trajectory maximum binding energy, the product dissociation activation energy, and the reactant binding activation energy; Determining channel energy weights based on the energy information of the channel to be processed and the energy information of the reference channel includes: Calculate the weight of the highest binding energy of the trajectory based on the highest binding energy of the trajectory to be processed and the highest binding energy of the reference trajectory; Calculate the product dissociation activation energy weight based on the dissociation activation energy of the product to be treated and the dissociation activation energy of the reference product; Calculate the weight of the reactant binding activation energy based on the binding activation energy of the reactant to be treated and the binding activation energy of the reference reactant; The channel energy weight is calculated based on the highest binding energy weight of the trajectory, the product dissociation activation energy weight, and the reactant binding activation energy weight.

[0066] In practical applications, channel energy information includes the highest binding energy of the trajectory, the product dissociation activation energy, and the reactant binding activation energy.

[0067] The energy information of the channel to be processed is compared and calculated separately with the energy information of the reference channel in the preset protein ligand calculation dataset.

[0068] Taking the highest binding energy of a trajectory as an example, the highest binding energy of the trajectory to be processed is obtained. The number of trajectories in the preset protein ligand calculation dataset whose reference highest binding energy is greater than the highest binding energy of the trajectory to be processed is counted. The proportion of these highest binding energies in the preset protein ligand calculation dataset is then used as the weight of the highest binding energy of the trajectory. For example, if the highest binding energy of the trajectory to be processed is T, and there are 1 million data points in the preset protein ligand calculation dataset, and 680,000 of these data points have reference highest binding energies greater than the highest binding energy T, then the weight of the highest binding energy of the trajectory is 0.68. In the method provided in the embodiments of this specification, the weight of the highest binding energy of the trajectory is denoted as P_Emax.

[0069] In the process of calculating the product dissociation activation energy weight based on the dissociation activation energy of the product to be processed and the dissociation activation energy of the reference product, the number of reference products with dissociation activation energies lower than those of the products to be processed in the preset protein ligand calculation dataset is counted, and the proportion of this number in the preset protein ligand calculation dataset is counted. This proportion is used as the product dissociation activation energy weight. In the method provided in the embodiments of this specification, the product dissociation activation energy weight is denoted as P_Eoff.

[0070] In the process of calculating the reactant binding activation energy weight based on the binding activation energy of the reactant to be treated and the reference reactant binding activation energy, the number of reactant binding activation energies greater than the binding activation energy of the reactant to be treated in the preset protein ligand calculation dataset is counted, and the proportion of this number in the preset protein ligand calculation dataset is counted. This proportion is used as the reactant binding activation energy weight. In the method provided in the embodiments of this specification, the product dissociation activation energy weight is denoted as P_Eon.

[0071] Given the maximum binding energy weight P_Emax, the product dissociation activation energy weight P_Eoff, and the product dissociation activation energy weight P_Eon, the channel energy weight can be calculated by multiplying these three values. In the method provided in the embodiments of this specification, the channel energy weight is denoted as P1, and P1 = P_Emax * P_Eoff * P_Eon.

[0072] In another specific embodiment provided in this specification, the channel attribute information includes channel length information, channel average radius, and channel bottleneck radius; Determining channel attribute weights based on the channel attribute information to be processed and the reference channel attribute information includes: Calculate the channel length information weight based on the channel length information to be processed and the reference channel length information; Calculate the average radius weight of the channel based on the average radius of the channel to be processed and the average radius of the reference channel; Calculate the channel bottleneck radius weight based on the bottleneck radius of the channel to be processed and the bottleneck radius of the reference channel; The channel attribute weights are calculated based on the channel length information weight, the channel average radius weight, and the channel bottleneck radius weight.

[0073] In the process of calculating the channel length information weight based on the channel length information to be processed and the reference channel length information, the number of reference channel length information values ​​greater than the channel length information to be processed in the preset protein ligand calculation dataset is counted, and the proportion of this number in the preset protein ligand calculation dataset is counted. This proportion is used as the channel length information weight. In the method provided in the embodiments of this specification, the channel length information weight is denoted as P_length.

[0074] In the process of calculating the channel average radius weight based on the average radius of the channel to be processed and the average radius of the reference channel, the number of reference channels with an average radius smaller than the average radius of the channel to be processed in the preset protein ligand calculation dataset is counted, and the proportion of this number in the preset protein ligand calculation dataset is counted. This proportion is used as the channel average radius weight. In the method provided in the embodiments of this specification, the channel average radius weight is denoted as P_radius_mean.

[0075] In the process of calculating the channel bottleneck radius weight based on the bottleneck radius of the channel to be processed and the bottleneck radius of the reference channel, the number of reference channel bottleneck radii smaller than the bottleneck radius of the channel to be processed in the preset protein ligand calculation dataset is counted, and the proportion of this number in the preset protein ligand calculation dataset is counted. This proportion is used as the channel bottleneck radius weight. In the method provided in the embodiments of this specification, the channel bottleneck radius weight is denoted as P_radius_min.

[0076] Given the channel length information weight P_length, the channel average radius weight P_radius_mean, and the channel bottleneck radius weight P_radius_min, the channel attribute weight can be calculated by multiplying the three. In the method provided in the embodiments of this specification, the channel attribute weight is denoted as P2, and then P2 = P_Emax * P_radius_mean * P_radius_min.

[0077] Having determined the channel energy weight P1 and channel attribute weight P2 through the above steps, the binding weight can be further determined using the channel energy weight P1 and channel attribute weight P2. In the method provided in the embodiments of this specification, the binding weight is denoted as P_GT, then P_GT=P1*P2, and further, P_GT=P_Emax*P_Eoff*P_Eon*P_Emax*P_radius_mean*P_radius_min.

[0078] In practical applications, the larger the binding weight P_GT corresponding to the binding channel to be processed, the greater the probability that the ligand to be screened will bind to the conformation of the target protein through the binding channel to be processed.

[0079] Step 110: Select the target binding channel corresponding to the conformation of the target protein and the ligand to be screened according to each binding weight.

[0080] After performing channel trajectory analysis and binding energy analysis on each binding channel using the methods described above, the binding weight corresponding to each binding channel can be obtained. The larger the binding weight, the greater the probability of ligand binding to protein. Based on this, the target binding channel corresponding to the target protein conformation can be selected from each binding channel according to the binding weight. In the method provided in the embodiments of this specification, the target binding channel represents an indicator of whether the target ligand can bind to the target protein conformation. If a target binding channel exists and the binding weight of the target binding channel is greater than a threshold, it indicates that the target ligand can bind to the target protein conformation. Conversely, even if a target binding channel exists, if the weight corresponding to the target binding channel is less than the threshold, the target ligand cannot enter the protein through the target binding channel, and the target ligand cannot bind to the target protein conformation.

[0081] In one specific embodiment provided in this specification, the target binding channel corresponding to the conformation of the target protein and the ligand to be screened is selected according to each binding weight, including: Sort the combined channels from largest to smallest according to their respective combined weights. Based on the sorting results, a preset number of binding channels are selected as the target binding channels for the ligands to be screened and the conformation of the target protein.

[0082] As mentioned above, the higher the binding weight, the greater the probability of ligand binding to protein. Therefore, the binding channels are sorted in descending order of their corresponding binding weights. Based on the sorting results, the highest-ranking binding channel is selected as the target binding channel for the ligand to be screened and the corresponding conformation of the target protein. In practical applications, the top-ranked binding channel is selected as the target binding channel for the ligand to be screened and the corresponding conformation of the target protein.

[0083] In one specific embodiment provided in this specification, when at least one ligand to be screened is obtained, the method further includes: Determine the binding weight of the target binding channel for each ligand to be screened; The ligands to be screened are sorted according to their binding weights to obtain the binding screening results of each ligand to the target protein conformation.

[0084] In the scenario provided in this embodiment, multiple ligands to be screened may be received. Using the method provided in this specification, the ease with which each ligand binds to the target protein conformation can be further determined. Specifically, the target binding channel corresponding to each ligand to be screened can be determined using the above method, and the binding weight corresponding to the target binding channel is used as the protein ligand binding weight for each ligand to be screened binding to the target protein conformation. The higher the protein ligand binding weight, the higher the probability of the corresponding ligand to be screened binding to the target protein conformation.

[0085] Each ligand to be screened is ranked by its binding weight, and the binding screening results corresponding to the conformation of the target protein are obtained. Specifically, ligands that obviously cannot bind are screened out, thereby effectively avoiding false positives where the ligand can theoretically bind to the protein but cannot actually enter the protein through the channel.

[0086] The methods provided in the embodiments of this specification fully utilize the results obtained from channel trajectory analysis and energy analysis to analyze the ease with which ligands bind from their binding sites to the protein core, providing an index to measure the ease of ligand-protein binding. Through statistical mixed analysis, reliable binding trajectory analysis and false positive filtering are achieved. By statistically analyzing various distributions on the real complex structure, a ranking index is established to determine the ease of binding between each ligand to the target protein conformation. This enriches the screening content of computer-based virtual screening.

[0087] See Figure 2 , Figure 2 This specification illustrates a process flowchart of a data processing method applied to cloud-side devices according to an embodiment of the present invention, which specifically includes the following steps: Step 202: Receive the information on the ligand to be screened and the conformation information of the target protein sent by the receiving end device.

[0088] Step 204: Determine at least one binding channel corresponding to the conformational information of the target protein.

[0089] Step 206: Calculate the channel energy information corresponding to each binding channel and the information of the ligand to be screened, and obtain the channel attribute information corresponding to each binding channel.

[0090] Step 208: Determine the binding weight corresponding to each binding channel based on the channel energy information and channel attribute information of each binding channel.

[0091] Step 210: Select the target binding channel corresponding to the ligand information to be screened and the conformation information of the target protein according to each binding weight.

[0092] Step 212: Send the target combined channel to the end-side device.

[0093] In practical applications, a large number of ligands and proteins are typically processed, requiring substantial computing resources. Some end-side devices may lack the necessary processing capabilities. Therefore, the method provided in the embodiments of this specification can also be implemented on a cloud-side device. After receiving the conformation of the ligand to be screened and the target protein from the end-side device, the cloud-side device performs the above-mentioned processing steps to obtain the target binding channel corresponding to the conformation of the ligand to be screened and the target protein, and finally sends the target binding channel to the end-side device.

[0094] In one specific embodiment provided in this specification, the conformation of the ligand to be screened and the target protein sent by the receiving end-side device includes: At least one ligand to be screened and the conformation of the target protein are sent by the receiving end device; Accordingly, the method further includes: Determine the binding weight of the target binding channel for each ligand to be screened; The ligands to be screened are sorted according to their binding weights to obtain the binding screening results of each ligand to be screened and the conformation of the target protein. The combined filtering results are sent to the end-side device.

[0095] In this embodiment, the cloud-side device receives multiple ligands to be screened and target protein conformations. Using the method described above, the target binding channels corresponding to the multiple ligands to be screened can be calculated. The binding weight corresponding to each target binding channel is used as the binding weight between each ligand to be screened and the target protein conformation, thereby determining the binding screening result of the ease of binding between each ligand to be screened and the target protein conformation. The binding screening result is then sent to the end-side device for reference.

[0096] The methods provided in the embodiments of this specification fully utilize the results obtained from channel trajectory analysis and energy analysis to analyze the ease with which ligands bind from their binding sites to the protein core, providing an index to measure the ease of ligand-protein binding. Through statistical mixed analysis, reliable binding trajectory analysis and false positive filtering are achieved. By statistically analyzing various distributions on the real complex structure, a ranking index is established to determine the ease of binding between each ligand to the target protein conformation. This enriches the screening content of computer-based virtual screening.

[0097] Corresponding to the above method embodiments, this specification also provides data processing apparatus embodiments. Figure 3 A schematic diagram of the structure of a data processing apparatus according to one embodiment of this specification is shown. Figure 3 As shown, the device includes: The acquisition module 302 is configured to acquire the conformation of the ligand to be screened and the target protein. The channel determination module 304 is configured to determine at least one binding channel corresponding to the conformation of the target protein. The calculation module 306 is configured to calculate the channel energy information of each binding channel and the ligand to be screened, and to obtain the channel attribute information corresponding to each binding channel. The weight determination module 308 is configured to determine the binding weight corresponding to each binding channel based on the channel energy information and channel attribute information of each binding channel. The selection module 310 is configured to select the target binding channel corresponding to the conformation of the target protein and the ligand to be screened according to each binding weight.

[0098] Optionally, the computing module 306 is further configured to: Select a combination channel to be processed, wherein the combination channel to be processed is any one of the combination channels; Obtain the channel length information, average channel radius, and channel bottleneck radius of the combined channel to be processed.

[0099] Optionally, the computing module 306 is further configured to: Select a combination channel to be processed, wherein the combination channel to be processed is any one of the combination channels; Calculate the target channel energy information corresponding to the ligand to be screened and the binding channel to be processed.

[0100] Optionally, the computing module 306 is further configured to: Calculate the protein surface binding energy, active site binding energy, and trajectory maximum binding energy corresponding to the ligand to be screened and the binding channel to be treated; Calculate the product dissociation activation energy and reactant binding activation energy based on the protein surface binding energy, the active site binding energy, and the trajectory maximum binding energy. The highest binding energy of the trajectory, the product dissociation activation energy, and the reactant binding activation energy are determined as the target channel energy information.

[0101] Optionally, the weight determination module 308 is further configured to: Obtain a preset protein ligand calculation dataset, and obtain the reference channel energy information and reference channel attribute information corresponding to the preset protein ligand calculation dataset; Select a combination channel to be processed, wherein the combination channel to be processed is any one of the combination channels; Based on the energy information and attribute information of the channel to be processed, and the energy information and attribute information of the reference channel, the target binding weight corresponding to the channel to be processed is determined.

[0102] Optionally, the weight determination module 308 is further configured to: The channel energy weights are determined based on the energy information of the channel to be processed and the energy information of the reference channel. The channel attribute weights are determined based on the channel attribute information to be processed and the reference channel attribute information; The target binding weight corresponding to the binding channel to be processed is determined based on the channel energy weight and the channel attribute weight.

[0103] Optionally, the channel energy information includes the highest trajectory binding energy, the product dissociation activation energy, and the reactant binding activation energy; The weight determination module 308 is further configured as follows: Calculate the weight of the highest binding energy of the trajectory based on the highest binding energy of the trajectory to be processed and the highest binding energy of the reference trajectory; Calculate the product dissociation activation energy weight based on the dissociation activation energy of the product to be treated and the dissociation activation energy of the reference product; Calculate the weight of the reactant binding activation energy based on the binding activation energy of the reactant to be treated and the binding activation energy of the reference reactant; The channel energy weight is calculated based on the highest binding energy weight of the trajectory, the product dissociation activation energy weight, and the reactant binding activation energy weight.

[0104] Optionally, the channel attribute information includes channel length information, average channel radius, and channel bottleneck radius; The weight determination module 308 is further configured as follows: Calculate the channel length information weight based on the channel length information to be processed and the reference channel length information; Calculate the average radius weight of the channel based on the average radius of the channel to be processed and the average radius of the reference channel; Calculate the channel bottleneck radius weight based on the bottleneck radius of the channel to be processed and the bottleneck radius of the reference channel; The channel attribute weights are calculated based on the channel length information weight, the channel average radius weight, and the channel bottleneck radius weight.

[0105] Optionally, the selection module 310 is further configured to: Sort the combined channels from largest to smallest according to their respective combined weights. Based on the sorting results, a preset number of binding channels are selected as the target binding channels for the ligands to be screened and the conformation of the target protein.

[0106] Optionally, the acquisition module 302 is further configured to: Obtain at least one ligand to be screened and the conformation of the target protein; Accordingly, the device also includes a screening module configured to: Determine the binding weight of the target binding channel for each ligand to be screened; The ligands to be screened are sorted according to their binding weights to obtain the binding screening results of each ligand to the target protein conformation.

[0107] The apparatus provided in the embodiments of this specification fully utilizes the results obtained from channel trajectory analysis and energy analysis to analyze the ease with which ligands bind from their binding sites to the protein core, providing an index to measure the ease of ligand-protein binding. Through statistical mixed analysis of data, reliable binding trajectory analysis and false positive filtering are achieved. By statistically analyzing various distributions on the real complex structure, a ranking index is established to determine the ease of binding between each ligand to the target protein conformation. This enriches the screening content of computer-based virtual screening.

[0108] The above is an illustrative scheme of a data processing apparatus according to this embodiment. It should be noted that the technical solution of this data processing apparatus and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing apparatus, please refer to the description of the technical solution of the data processing method described above.

[0109] See Figure 4 , Figure 4 This specification illustrates an architecture diagram of a data processing system according to one embodiment of the present specification. The data processing system may include a client 100 and a server 200. Client 100 is used to send the ligand to be screened and the conformation of the target protein to server 200; Server 200 is used to determine at least one binding channel corresponding to the conformation of the target protein; calculate the channel energy information of each binding channel and the ligand to be screened, and obtain the channel attribute information of each binding channel; determine the binding weight of each binding channel based on the channel energy information and channel attribute information of each binding channel; select the target binding channel corresponding to the conformation of the target protein and the ligand to be screened based on each binding weight; and send the target binding channel to client 100. Client 100 is also used to receive target combination channels sent by server 200.

[0110] The data processing system may include multiple clients 100 and a server 200. Clients 100 can be referred to as edge devices, and server 200 can be referred to as cloud devices. Multiple clients 100 can establish communication connections through server 200. In the complex generation scenario, server 200 is used to provide a service for determining the binding channels of ligands and protein conformations among multiple clients 100. Multiple clients 100 can act as either senders or receivers, communicating through server 200.

[0111] Users can interact with server 200 through client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the complex generation scenario, users can publish data streams to server 200 through client 100, server 200 can generate target combination channels based on the data streams, and push the target combination channels to other clients that have established communication.

[0112] In this system, client 100 and server 200 establish a connection via a network. The network provides the medium for communication between client 100 and server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. Data transmitted by client 100 may need to undergo encoding, transcoding, compression, or other processing before being published to server 200.

[0113] Client 100 can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. Client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by server 200, such as a real-time communication (RTC) SDK. Client 100 can be deployed on electronic devices and depends on the device or certain apps on the device to run. Electronic devices may have displays and support information browsing, such as personal mobile terminals like mobile phones, tablets, and personal computers. Various other types of applications can also be configured on electronic devices, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.

[0114] Server 200 may include servers providing various services, such as servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server integrated with blockchain. The server can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0115] It is worth noting that the data processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the data processing methods provided in the embodiments of this specification. In other embodiments, the data processing methods provided in the embodiments of this specification may also be executed jointly by the client and the server.

[0116] Figure 5 A schematic diagram of a task platform provided in an embodiment of this specification is shown. The task platform includes a request interface 502 and a response unit 504, wherein: Request interface 502 is used to receive the ligand to be screened and the conformation of the target protein sent by the end-side device.

[0117] The response unit 504 is used to determine at least one binding channel corresponding to the conformation of the target protein; calculate the channel energy information of each binding channel and the ligand to be screened, and obtain the channel attribute information corresponding to each binding channel; determine the binding weight of each binding channel based on the channel energy information and channel attribute information of each binding channel; and select the target binding channel corresponding to the conformation of the target protein and the ligand to be screened based on the binding weight.

[0118] Figure 6 A structural block diagram of a computing device 600 according to one embodiment of this specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.

[0119] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0120] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0121] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.

[0122] The processor 620 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-described data processing method.

[0123] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.

[0124] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.

[0125] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.

[0126] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.

[0127] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the data processing method described above.

[0128] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0129] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0130] It should be noted that the above description describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0131] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0132] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A data processing method, comprising: Obtain the conformation of the ligand to be screened and the target protein; Identify at least one binding channel corresponding to the conformation of the target protein; Calculate the channel energy information of each binding channel and the ligand to be screened, and obtain the channel attribute information of each binding channel; Based on the channel energy information and channel attribute information of each binding channel, determine the binding weight corresponding to each binding channel; The target binding channels corresponding to the conformation of the target protein and the ligand to be screened are selected based on the binding weights.

2. The method as described in claim 1, wherein obtaining the channel attribute information corresponding to each combined channel includes: Select a combination channel to be processed, wherein the combination channel to be processed is any one of the combination channels; Obtain the channel length information, average channel radius, and channel bottleneck radius of the combined channel to be processed.

3. The method of claim 1, comprising calculating the channel energy information corresponding to each binding channel and the ligand to be screened, including: Select a combination channel to be processed, wherein the combination channel to be processed is any one of the combination channels; Calculate the target channel energy information corresponding to the ligand to be screened and the binding channel to be processed.

4. The method of claim 3, comprising calculating the target channel energy information corresponding to the ligand to be screened and the binding channel to be processed, including: Calculate the protein surface binding energy, active site binding energy, and trajectory maximum binding energy corresponding to the ligand to be screened and the binding channel to be treated; Calculate the product dissociation activation energy and reactant binding activation energy based on the protein surface binding energy, the active site binding energy, and the trajectory maximum binding energy. The highest binding energy of the trajectory, the product dissociation activation energy, and the reactant binding activation energy are determined as the target channel energy information.

5. The method according to any one of claims 1-4, wherein the binding weight corresponding to each binding channel is determined based on the channel energy information and channel attribute information of each binding channel, comprising: Obtain a preset protein ligand calculation dataset, and obtain the reference channel energy information and reference channel attribute information corresponding to the preset protein ligand calculation dataset; Select a combination channel to be processed, wherein the combination channel to be processed is any one of the combination channels; Based on the energy information and attribute information of the channel to be processed, and the energy information and attribute information of the reference channel, the target binding weight corresponding to the channel to be processed is determined.

6. The method as described in claim 5, wherein the target binding weight corresponding to the channel to be processed is determined based on the channel energy information and channel attribute information corresponding to the channel to be processed, and the reference channel energy information and the reference channel attribute information, comprising: The channel energy weights are determined based on the energy information of the channel to be processed and the energy information of the reference channel. The channel attribute weights are determined based on the channel attribute information to be processed and the reference channel attribute information; The target binding weight corresponding to the binding channel to be processed is determined based on the channel energy weight and the channel attribute weight.

7. The method of claim 6, wherein the channel energy information includes the maximum binding energy of the trajectory, the product dissociation activation energy, and the reactant binding activation energy; Determining channel energy weights based on the energy information of the channel to be processed and the energy information of the reference channel includes: Calculate the weight of the highest binding energy of the trajectory based on the highest binding energy of the trajectory to be processed and the highest binding energy of the reference trajectory; Calculate the product dissociation activation energy weight based on the dissociation activation energy of the product to be treated and the dissociation activation energy of the reference product; Calculate the weight of the reactant binding activation energy based on the binding activation energy of the reactant to be treated and the binding activation energy of the reference reactant; The channel energy weight is calculated based on the highest binding energy weight of the trajectory, the product dissociation activation energy weight, and the reactant binding activation energy weight.

8. The method as described in claim 6, wherein the channel attribute information includes channel length information, channel average radius, and channel bottleneck radius; Determining channel attribute weights based on the channel attribute information to be processed and the reference channel attribute information includes: Calculate the channel length information weight based on the channel length information to be processed and the reference channel length information; Calculate the average radius weight of the channel based on the average radius of the channel to be processed and the average radius of the reference channel; Calculate the channel bottleneck radius weight based on the bottleneck radius of the channel to be processed and the bottleneck radius of the reference channel; The channel attribute weights are calculated based on the channel length information weight, the channel average radius weight, and the channel bottleneck radius weight.

9. The method of claim 1, wherein the target binding channel corresponding to the conformation of the target protein and the ligand to be screened is selected according to each binding weight, comprising: Sort the combined channels from largest to smallest according to their respective combined weights. Based on the sorting results, a preset number of binding channels are selected as the target binding channels for the ligands to be screened and the conformation of the target protein.

10. The method of claim 1, wherein obtaining the conformation of the ligand to be screened and the target protein comprises: Obtain at least one ligand to be screened and the conformation of the target protein; Accordingly, the method further includes: Determine the binding weight of the target binding channel for each ligand to be screened; The ligands to be screened are sorted according to their binding weights to obtain the binding screening results of each ligand to the target protein conformation.

11. A data processing method applied to cloud-side equipment, comprising: The receiving end device sends information on the ligands to be screened and the conformational information of the target protein; Determine at least one binding channel corresponding to the conformational information of the target protein; Calculate the channel energy information corresponding to each binding channel and the ligand information to be screened, and obtain the channel attribute information corresponding to each binding channel; Based on the channel energy information and channel attribute information of each binding channel, determine the binding weight corresponding to each binding channel; Based on each binding weight, the target binding channel corresponding to the ligand information to be screened and the conformation information of the target protein is selected; The target is sent to the end-side device via the combined channel.

12. The method of claim 11, wherein receiving the ligand to be screened and the conformation of the target protein sent by the receiving end-side device comprises: At least one ligand to be screened and the conformation of the target protein are sent by the receiving end device; Accordingly, the method further includes: Determine the binding weight of the target binding channel for each ligand to be screened; The ligands to be screened are sorted according to their binding weights to obtain the binding screening results of each ligand to be screened and the conformation of the target protein. The combined filtering results are sent to the end-side device.

13. A task platform, comprising a request interface and a response unit; The request interface is used to receive the ligand to be screened and the conformation of the target protein sent by the end device; The response unit is used to determine at least one binding channel corresponding to the conformation of the target protein; Calculate the channel energy information of each binding channel and the ligand to be screened, and obtain the channel attribute information of each binding channel; Based on the channel energy information and channel attribute information of each binding channel, determine the binding weight corresponding to each binding channel; The target binding channels corresponding to the conformation of the target protein and the ligand to be screened are selected based on the binding weights.

14. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 12.

15. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 12.

16. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 12.