Drug structure-activity relationship analysis method and device based on multi-level structure
Through the multi-layered drug structure-activity relationship analysis method, the problem that traditional methods are difficult to integrate multi-source data is solved, and the quantitative evaluation of the relationship between compounds and target activity is achieved, which improves the accuracy and efficiency of drug design.
Patent Information
- Application Number
- CN202510278030.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-10
AI Technical Summary
In the development of small molecule drugs, traditional structure-activity relationship analysis methods are difficult to consider the macroscopic and microscopic structure effects of molecules at the same time, and cannot effectively integrate data from different sources, resulting in a lack of accuracy and efficiency in drug design.
The multi-layered drug structure-activity relationship analysis method is used to characterize the compounds by clustering, characterizing the original structure and fragment structure, and the weights of different hierarchical structures are calculated, and the weights are dynamically adjusted based on the R&D status of the compounds, combined with target correlation evaluation, and quantitative evaluation of the activity relationship between the compounds and the targets is achieved.
It improves the accuracy and efficiency of drug structure-effect relationship analysis, can better identify key pharmacodynamic groups, provide more reliable drug screening standards, and optimize drug design process.
Smart Images

Figure CN120280031A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of drug research and development, and more particularly, to a method, apparatus, computing device, and storage medium for analyzing the structure-activity relationship of drugs based on a multi-level structure. Background Art
[0002] Small molecule drugs are usually organic compounds with a molecular weight below 1000 daltons. In the process of small molecule drug research and development, the analysis of the structure-activity relationship (SAR) is a core link. SAR refers to the relationship between the chemical structure of a drug or other bioactive substance and its biological activity, which can provide a theoretical basis for new drug research and development. Traditional SAR analysis methods focus on the intuitive association between the chemical structure of a compound and its biological activity. However, with the rapid increase in drug research and development data and the continuous improvement of complexity, traditional SAR analysis methods face many challenges in practical applications.
[0003] On the one hand, since the biological activity of a molecule is not only determined by its macroscopic structure (such as molecular size, shape, functional groups), but also affected by more microscopic aspects (such as electronic effects, stereochemical properties, solvent effects, etc.). On the other hand, in addition to the structure and biological activity data of compounds, drug research and development also involve a large amount of other data, such as clinical trial data, experimental methods, toxicity data, pharmacokinetic information, etc. The activity of a compound may show significant differences under different experimental methods, and the clinical trial status will also affect the drug efficacy. Moreover, there may be certain conflicts or inconsistencies in different experimental results and data sources. In this context, how to establish a reasonable data integration mechanism, especially how to handle data differences from different platforms, different experimental methods, and even different researchers, and effectively integrate and analyze them, has become an important problem to be solved.
[0004] In the prior art solutions, for the similarity search method based on molecular fingerprints, the molecular structure is converted into binary fingerprints, and the structure-activity relationship is analyzed by calculating the similarity between molecular fingerprints. This method can capture the local structural features of molecules, but it depends on a single-level structural representation and it is difficult to consider both the basic skeleton structure and local substituent features of molecules simultaneously. For the SAR analysis method based on substructure matching, by identifying the common skeleton and variable substituents in molecules, the influence of different substituents on activity is analyzed. However, the evaluation of the importance of molecular fragments (such as functional groups, side chains, etc.) at different development stages is relatively vague, and there is a lack of quantitative evaluation of the credibility of target association. For the activity prediction method based on machine learning, machine learning algorithms are used to establish a prediction model between molecular descriptors and activity. Although this method can establish a relationship model between molecular descriptors and activity, it is difficult to integrate activity data from different sources, resulting in the omission or misjudgment of some potential targets. The model usually lacks dynamic adaptation to the changing requirements and goals during the drug development process and fails to provide targeted strategies or evaluations at different stages. For example, in the early stage, it may be more necessary to predict and screen compounds, while in the later stage, more attention may be paid to the toxicity and pharmacokinetic properties of the drug. Therefore, accurate drug design cannot be achieved solely through the activity prediction model. Summary of the Invention
[0005] To improve the accuracy and efficiency of SAR analysis, the embodiments described in this article provide a method, apparatus, computing device, and computer-readable storage medium storing a computer program for analyzing the structure-activity relationship of drugs based on a multi-level structure.
[0006] According to a first aspect of the present disclosure, there is provided a method for analyzing the structure-activity relationship of drugs based on a multi-level structure, including: performing a multi-level molecular structure characterization on a compound, where the multi-level molecular structure includes a clustering structure, an original structure, and a fragment structure; calculating the weights of the clustering structure, the original structure, and the fragment structure based on the R & D status of the compound; and determining the activity relationship between the compound and the target, and calculating a target association credibility score based on the activity relationship, the evidence supporting the activity relationship, and the R & D status weight of the compound.
[0007] In some embodiments of the present disclosure, performing a multi-level molecular structure characterization on a compound, where the multi-level molecular structure includes a clustering structure, an original structure, and a fragment structure includes: collecting experimental data, clustering the molecular structures in the experimental data using the Butina clustering method to obtain the clustering structure of the molecules; converting the molecular structure into the standard SMILES format, and performing stereochemical processing and centering processing to obtain the original structure of the molecules; decomposing the molecular structure into fragments based on molecular fragmentation rules or chemical synthesis feasibility, screening and de-duplicating the fragments to obtain the fragment structure of the molecules.
[0008] In some embodiments of the present disclosure, the Butina clustering method is used to cluster the molecular structures in the test data, and the obtained clustering structure of the molecules includes: using the Murcko skeleton extraction method to remove the side chains in the molecular structure to obtain a molecular skeleton that only contains the ring system of the molecule and the atoms connecting the ring systems; converting the molecular skeleton into a fingerprint and calculating the similarity between different fingerprints; setting a similarity threshold and a clustering size threshold, and dividing the molecular skeletons into multiple clusters based on the similarity; calculating the average similarity between each molecular skeleton in the cluster and other members, and selecting the molecular skeleton with the highest average similarity as the representative structure of the cluster.
[0009] In some embodiments of the present disclosure, converting the molecular structure into the standard SMILES format and performing stereochemical processing and centering processing to obtain the original structure of the molecule includes: converting the molecular structure into the standard SMILES format; retaining the known stereochemical information of the molecular structure, removing the undefined stereocenter markers, and unifying the representation of cis-trans isomers; and removing the inorganic salt ions in the molecular structure, standardizing the charge state and the protonation state.
[0010] In some embodiments of the present disclosure, decomposing the molecular structure into fragments based on molecular fragmentation rules or chemical synthesis feasibility, screening and removing duplicates from the fragments to obtain the fragment structure of the molecule includes: setting molecular fragmentation rules, where the rules include: the fragment contains at least one ring system, and the number of heavy atoms in the fragment is not less than 8, and the size of the largest fragment does not exceed 70% of the number of heavy atoms in the original molecule; using the rule-based RECAP method to decompose the molecule into fragments according to the reactivity and synthesis pathway of the molecule, or using the BRICS method based on chemical synthesis feasibility to generate fragments by cutting the rotatable bonds in the molecule; setting screening criteria for the fragments including the molecular weight range, the number of rotatable bonds, and the number of hydrogen bond donors / acceptors, screening out the fragments that meet the screening criteria, and removing duplicates from the screened fragments to obtain a plurality of fragment structures.
[0011] In some embodiments of the present disclosure, calculating the weights of the clustering structure, the original structure, and the fragment structure based on the R & D status of the compound includes: calculating the R & D status weight W dev (c) of the compound through the following formula:
[0012] W dev (c) = BaseW·(1 + ∑(Phase i ·K i )) or
[0013] W dec (c) = Base W ·(1 + max(Phase)·K phase ) In the formula, BaseW Denotes the base weight value, set to 1.0, Phase i Is the R & D stage identification variable, with values: pre - clinical = 0, clinical phase I = 1, clinical phase II = 2, clinical phase III = 3, marketing approval = 4, K i Is the influence factor for each stage, K 临床前 = 0.1; K 临床I期 = 0.2; K 临床II期 = 0.4; K 临床III期 = 0.6; K 上市批准 = 1.0;
[0014] Calculate the weight W of the clustering structure according to the R & D status weight, R & D stage distribution factor of each member in the molecular clustering, and the number of internal members of the clustering cluster (cls):
[0015] W cbuster (cls)= ∑(W dev (c i ))·(1 + ln(N members ))·F phase
[0016] In the formula, W dev(ci) Is the sum of the R & D status weights of the molecular members of the clustering, N members Is the number of molecular members of the clustering, F phase Is the R & D stage distribution factor:
[0017]
[0018] Among them, N approved Is the number of molecules with marketing approval, N phaseIII Is the number of molecules in clinical phase III, N phaseII Is the number of molecules in clinical phase II, N phaseI Is the number of molecules in clinical phase I;
[0019] Calculate the weight W of the fragment structure by statistically analyzing the occurrence frequency and R & D status weighted average value of the fragment in different R & D stages frag (f):
[0020] W frag (f)= F global (f)·W state (f), where,
[0021] In the formula, F global (f) is the global frequency of fragment f, N occur (f) is the number of occurrences of fragment f, N totalis the total number of molecules, I(f, c i ) is an indicator function. When molecule c i contains fragment f, I(f, c i ) is 1; otherwise it is 0. W state (f) is the weighted average of the R & D status, and W dev (c i ) is the R & D status weight of molecule c i ; A top - down and bottom - up weight transfer system is formed based on the R & D status weight, the clustering structure weight, and the fragment structure weight.
[0022] In some embodiments of the present disclosure, determining the activity relationship between a compound and a target, and calculating the target - associated confidence score based on the activity relationship, the evidence supporting the activity relationship, and the R & D status weight includes: determining whether there is an activity relationship between the compound and the target by using a threshold calculation method according to the activity data including the half - inhibitory concentration, the half - effective concentration, the inhibition constant, and the dissociation constant:
[0023]
[0024] In the formula, IC 50 is the half - inhibitory concentration, EC 50 is the half - effective concentration, K i is the inhibition constant, K d is the dissociation constant, 1 indicates an activity relationship, and 0 indicates no activity relationship;
[0025] Calculating the confidence score of the target - associated based on the R & D status weight of the compound, the determination result of the activity relationship, and the number of evidences supporting the activity relationship:
[0026] Conf(c, t) = W dev (c)·R target (c, t)·(1 + ln(N evidence ));
[0027] Wherein, W dev (c) is the R & D status weight of the compound, R target (c, t) is the determination result of the target relationship, and N evidence is the number of evidences supporting the target - associated relationship. The number of evidences includes the number of independent experiments, the number of literatures or patents to which the experiments belong.
[0028] According to the second aspect of the present disclosure, a drug structure-activity relationship analysis device based on a multi-level structure is provided, including a multi-level structure characterization module, a weight calculation module and a target association evaluation module, wherein the multi-level structure characterization module is used to perform multi-level molecular structure characterization of a compound, and the multi-level molecular structure includes a cluster structure, an original structure and a fragment structure; the weight calculation module is used to calculate the weights of the cluster structure, the original structure and the fragment structure based on the research and development status of the compound; the target association evaluation module is used to determine the activity relationship between the compound and the target, and calculate the target association credibility score based on the activity relationship, the evidence supporting the activity relationship and the research and development status weight of the compound.
[0029] According to a third aspect of the present disclosure, a computing device is provided, the computing device comprising at least one processor; and at least one memory storing a computer program. When the computer program is executed by the at least one processor, the computing device executes the steps of the multi-level structure-based drug structure-activity relationship analysis method described in the first aspect of the present disclosure.
[0030] According to a fourth aspect of the present disclosure, a computer-readable storage medium storing a computer program is provided, wherein the computer program, when executed by a processor, implements the steps of the multi-level structure-based drug structure-activity relationship analysis method according to the first aspect of the present disclosure.
[0031] According to the drug structure-activity relationship analysis method and device based on multi-level structure provided by the present disclosure, by clustering, characterizing the original structure and fragment structure of the compound, it is helpful to fully understand the overall picture of the molecule and its potential pharmacodynamic characteristics. According to the research and development status of the compound, the weights of different hierarchical structures are dynamically calculated, which can improve the timeliness of the analysis and the accuracy of identifying key pharmacophores, thereby better guiding drug development, especially in complex drug systems, it can clearly identify the structural parts that can effectively bind to the target. By calculating the target association credibility score, it is possible to achieve a quantitative evaluation of the activity relationship between the compound and the target, providing a more objective standard for drug screening and more reliable decision support. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly described below. It should be noted that the drawings described below only relate to some embodiments of the present disclosure, but are not intended to limit the present disclosure, wherein:
[0033] Figure 1 is an exemplary flow chart of a multi-level structure-based drug structure-activity relationship analysis method 100 according to an embodiment of the present disclosure;
[0034] Figure 2Shows an exemplary schematic diagram of a multi-level molecular structure characterization according to an embodiment of the present disclosure;
[0035] Figure 3 Is a schematic structural diagram of a drug structure-activity relationship analysis device 300 based on a multi-level structure according to an embodiment of the present disclosure;
[0036] Figure 4 Is a schematic block diagram of a computing device according to an embodiment of the present disclosure.
[0037] It should be noted that the elements in the drawings are schematic and not drawn to scale. Detailed implementation manners
[0038] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of the present disclosure without creative efforts also belong to the scope protected by the present disclosure.
[0039] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the subject matter of the present disclosure belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the specification and the relevant art, and will not be interpreted in an idealized or overly formal form unless clearly defined otherwise herein. Additionally, terms such as "first" and "second" are only used to distinguish one component (or a part of a component) from another component (or another part of a component).
[0040] Aiming at the problems that the prior art cannot effectively process the association information between compound structures and targets at different levels, resulting in the inability to fully explore and utilize the structure-activity relationship, and the inability to effectively adapt to the changing needs during the R & D process, the present disclosure proposes a drug structure-activity relationship analysis method based on a multi-level structure, aiming to analyze the molecular structure of compounds at multiple levels, be able to more accurately capture different characteristics of compounds, and then conduct in-depth analysis of their association with targets. And based on the changes in different stages and information during the drug R & D process, dynamically adjust the weights between compounds and targets, be able to accurately reflect the structural characteristics and target association situations in different stages, thereby improving the accuracy of the structure-activity relationship map. By constructing a quantitative target association assessment, a systematic and quantitative analysis of the relationship between compounds and targets can be achieved, further optimizing the drug screening and design process and improving the drug R & D efficiency.
[0041] Figure 1 Exemplary flowchart showing the method 100 for analyzing the structure-activity relationship of drugs based on a multi-level structure according to an embodiment of the present disclosure.
[0042] At Figure 1 block S102, multi-level molecular structure characterization is performed on the compound, and the multi-level molecular structure includes a clustering structure, an original structure, and a fragment structure.
[0043] Relevant academic papers and research reports can be consulted through academic databases such as Google Scholar, PubMed, etc., or publicly available datasets such as ChEMBL, PubChem, etc. to obtain existing experimental data. In the embodiments of the present disclosure, the molecular structures in the experimental data are hierarchically expressed by computational methods, mainly divided into three levels: the clustering structure layer, the original structure layer, and the fragment structure layer.
[0044] Among them, the clustering structure represents the main categories of the molecular skeleton and is used to extract the core structural features of the compound. The Butina clustering method can be used to cluster the molecular structures in the experimental data to obtain the clustering structure of the molecules.
[0045] In an embodiment of the present disclosure, the Murcko skeleton extraction method can be used to remove the side chains in the molecular structure to obtain a molecular skeleton that only contains the ring system of the molecule and the atoms connecting the ring systems. By removing the side chains and substituents (these non-ring parts are usually the variable parts of the molecule and do not affect the basic skeleton structure of the molecule), the important ring systems and their connection relationships are retained, and the part representing the molecular skeleton is extracted. For example, the original compound is C9H 12 O2 (assumed to be a certain drug molecule, containing a ring system and multiple functional groups). After Murcko skeleton extraction, the obtained skeleton may be just a benzene ring or a cyclohexene structure, removing all non-ring parts such as methyl, hydroxyl, etc.
[0046] Then, the extracted molecular skeleton is converted into a fingerprint, and the similarity between different fingerprints is calculated. For example, the molecular skeleton is converted into a Morgan fingerprint (this fingerprint is generated based on the local chemical environment of each atom in the molecule through an iterative process). The generation of the Morgan fingerprint is based on the following steps:
[0047] Initialization: Fingerprint generation begins by assigning a unique label (usually the atomic number or other identifier) to each atom in the molecule;
[0048] Iterative update: In each iteration, the label of each atom is updated to the hash value of the combination of its own label and the labels of neighboring atoms. This process is usually repeated multiple times, and the "environment" of the atoms considered in each iteration expands to more distant neighbors. In this way, the label of each atom gradually contains more and more information about its chemical environment;
[0049] Radius and bit vector: A key parameter of the Morgan algorithm is the "radius", which refers to how far neighbors around an atom are to be considered. For example, a radius of 1 means only directly connected atoms are considered, and a radius of 2 extends to atoms beyond two bonds. Eventually, the final label of each atom is used to update a bit vector (usually of a fixed length, such as 1024 or 2048 bits). If an atom's label maps to a specific bit through a hash function, then that bit is set to 1;
[0050] Result: The final bit vector represents the chemical structure and properties of the entire molecule. Due to its environment-based generation method, the Morgan fingerprint can capture the context information of atoms in a molecule, which makes it very useful in predicting molecular properties and behaviors.
[0051] RDKit is an open-source cheminformatics and machine learning software, mainly used for processing data on molecules and chemical reactions. RDKit provides a series of tools and algorithms that allow users to process molecular structures, perform chemical reaction simulations, analyze chemical data, and conduct molecular searches and matches. In some embodiments of the present disclosure, the RDKit library can be used to generate Morgan fingerprints and calculate the pairwise similarity between molecules. Cosine similarity or Tanimoto coefficient can be used to calculate the similarity between fingerprints. For example, Tanimoto(A,B) = |A∪B| / |A∩B|, where |A∩B| is the number of common features of the two fingerprints, and |A∪B| is the total number of all features in the two fingerprints. The similarity score is a number between 0 and 1, with 1 indicating exactly the same and 0 indicating completely different.
[0052] Set a similarity threshold and a cluster size threshold, and divide the molecular skeletons into multiple clusters based on the similarity. The similarity threshold controls which molecules are grouped into the same class, while the cluster size threshold is used to limit the number of molecules included in each cluster. Through these conditions, molecules with similar structures can be effectively organized and classified. Finally, calculate the average similarity between each molecular skeleton within a cluster and other members, and select the molecular skeleton with the highest average similarity as the representative structure of the cluster.
[0053] The original structure is the complete molecular structure, maintaining the original information of the molecule, including all atoms and bonding patterns. To ensure the consistency and comparability of the structural data, the molecular structure is converted into the standard SMILES format (a linear string format for representing chemical molecular structures), and stereochemical processing and centering processing are performed to construct the original structure.
[0054] During the standardization process, known stereochemical information of the molecular structure is retained, undefined stereocenter markers are removed, and confusion or errors caused by lack of stereochemical information are avoided. Also, the representation of cis-trans isomers is unified. For example, all cis-trans isomers are labeled in a standard cis or trans form to ensure a unified representation in different situations. In addition, to ensure consistency in aspects such as the charge and ions of the molecular structure, inorganic salt ions are removed, and the charge state and protonation state (i.e., the hydrogen ionization state of certain groups in the molecule) are standardized to ensure that functional groups such as amino and carboxyl groups have a consistent representation under the same pH conditions. After standardization, the molecular structure will conform to a unified format and standard.
[0055] Fragment structures are chemical structure fragments obtained through the molecular fragmentation method, which helps to understand the details inside the molecule, identify the key functional units of the molecule, and is helpful for molecule optimization and synthesis. In some embodiments of the present disclosure, the molecular structure is decomposed into fragments based on molecular fragmentation rules or chemical synthesis feasibility, and the fragments are screened and duplicate entries are removed to obtain the fragment structure of the molecule.
[0056] For example, the molecular fragmentation rules include: the fragment contains at least one ring system, and the number of heavy atoms (usually referring to elements such as carbon, nitrogen, oxygen, and phosphorus) in the fragment is at least 8, and the size of the largest fragment cannot exceed 70% of the number of heavy atoms of the original molecule, so that the fragmented molecule maintains diversity among fragments of different sizes.
[0057] Two main methods can be used to generate molecular fragments: using the rule-based RECAP method to decompose the molecule into fragments according to the reactivity and synthesis pathway of the molecule, or using the BRICS method based on chemical synthesis feasibility to generate fragments by cutting the rotatable bonds in the molecule.
[0058] After fragment generation, the fragments are screened and de-duplicated to construct a fragment structure. Specifically, screening criteria for the fragments are set, such as molecular weight range, number of rotatable bonds, number of hydrogen bond donors / acceptors, etc., and the fragments that meet the screening criteria are selected. Generally, the molecular weight of the fragments should be within a certain range. Fragments that are too large may be too complex and not easy to synthesize, while fragments that are too small may not have enough chemical information. A common molecular weight range can be set to 150–350 daltons, which can be adjusted according to specific requirements. This can ensure the feasibility of the selected fragments in drug design and synthesis. The number of rotatable bonds affects the flexibility of the molecule. There should not be too many rotatable bonds in the fragment, because the presence of rotatable bonds may lead to a decrease in the stability of the fragment and affect its feasibility as a drug candidate. Generally, the number of rotatable bonds in the fragment is limited to 2–5. Hydrogen bonds are an important way of intermolecular interaction. In fragment screening, it is necessary to ensure that the fragment has an appropriate number of hydrogen bond donors and acceptors. The number of hydrogen bonds should be balanced. Too many donors or acceptors may lead to an unstable molecule or a decrease in the affinity for the target. Therefore, the number of hydrogen bond donors / acceptors can be set within a specific range according to the drug activity target. To avoid interference from similar fragments to subsequent analysis and screening, the selected fragments are de-duplicated to obtain multiple fragment structures. For example, by comparing the SMILES or fingerprint representations of the fragments, similar or identical fragments are identified and removed.
[0059] Figure 2 Shows an exemplary schematic diagram of a multi-level molecular structure characterization according to an embodiment of the present disclosure. In Figure 2 the example, the clustering structure is the representative structure of the molecular backbone, retaining the ring system and the atoms connecting the ring systems. The original structure maintains the complete molecular structure through normalization. The fragment structure is multiple component fragments after the original structure is disassembled.
[0060] Subsequently, in step S104, the weights of the clustering structure, the original structure, and the fragment structure are calculated based on the R & D status of the compound.
[0061] To identify the backbone structure, the original structure, and the fragment structure with drug R & D value or potential, the importance of different hierarchical structures can be quantitatively calculated to effectively guide drug design and screening. When calculating the importance of different hierarchical structures, the R & D status of the compound is a core influencing factor. The embodiments of the present disclosure establish a weight transfer system that influences each other from top to bottom and from bottom to top with the R & D status as the core.
[0062] In some embodiments of the present disclosure, updates on the R & D status of compounds can be centrally obtained from clinical, approval agencies, or commercial databases in various countries. The R & D status refers to the different stages a drug is in during the drug development process. Different stages reflect the different degrees of a drug from laboratory research to clinical application. Therefore, each stage has a corresponding basic weight to quantify its influence in drug R & D, indicating that regardless of the stage, the drug has a certain R & D potential. Depending on whether the scenario where a drug has different R & D statuses in different therapeutic areas is ignored, the weight of the compound R & D status W is calculated through the following formula dev (c):
[0063] W dev (c) = Base W ·(1 + ∑(Phase i ·K i )) or
[0064] W dev (c) = Base W ·(1 + max(Phase)·K phase )
[0065] In the formula, Base W represents the basic weight value, set to 1.0, Phase i is the R & D stage identification variable, with values: pre - clinical = 0; clinical phase I = 1; clinical phase II = 2; clinical phase III = 3; marketing approval = 4, and K i is the influencing factor for each stage. Each stage will have different influencing factors according to its clinical or experimental progress. The influencing factor for the pre - clinical stage (Phase = 0) can be set as: K 临床前 = 0.1; the influencing factor for clinical phase I (Phase = 1): K 临床I期 = 0.2; the influencing factor for clinical phase II (Phase = 2): K 临床II期 = 0.4; the influencing factor for clinical phase III (Phase = 3): K 临床III期 = 0.6; the influencing factor for marketing approval (Phase = 4): K 上市批准 = 1.0.
[0066] For example, if a certain compound is in clinical phase II (Phase = 2), its weight is: W dev = 1.0×(1 + 2×0.4) = 1.8, indicating that the compound has a relatively high potential in R & D.
[0067] The clustering structure is the result of grouping molecules based on the similarity of their skeletal structures. The calculation of the clustering structure weight takes into account the R & D status of the molecules in each cluster, the R & D stage distribution factor, and the number of internal members of the cluster. According to the R & D status weight, the R & D stage distribution factor of each member in the molecular cluster, and the number of internal members of the cluster, the weight W of the clustering structure is calculated. cluster (cls):
[0068] W cluster (cls) = ∑(W dev (c i ))·(1 + ln(N members ))·F phase
[0069] In the formula, W dev(ci) is the sum of the R & D status weights of the molecules in the cluster members, N members is the number of molecules in the cluster members, and F phase is the R & D stage distribution factor:
[0070]
[0071] Among them, N approved is the number of molecules approved for marketing, N phaseIII is the number of molecules in clinical phase III, N phaseII is the number of molecules in clinical phase II, N phaseI is the number of molecules in clinical phase I.
[0072] The calculation steps include: First, based on the formula for calculating the R & D status weight of compounds, calculate the sum of the R & D status weights of the molecules in the cluster members. The R & D stage distribution factor is calculated according to the R & D stage distribution of the cluster members, reflecting the proportion or importance of drugs or candidate molecules at different stages. This calculation method reflects the cumulative effect of the R & D status of the members, the logarithmic gain of the number of members, and the additional contribution of high-stage members. Taking a clustering structure containing 4 compounds as an example, Compound A: an approved drug (Phase = 4); Compound B: in clinical phase III (Phase = 3), Compound C: in clinical phase II (Phase = 2); Compound D: preclinical (Phase = 0). First, calculate the R & D status weight W dev :
[0073] Compound A: W dev (A) = 1.0*(1 + 4*1.0) = 5.0
[0074] Compound B: W dev (B) = 1.0*(1 + 3*0.6) = 2.8
[0075] Compound C: W dev(C) = 1.0 * (1 + 2 * 0.4) = 1.8
[0076] Compound D: W dev (D) = 1.0 * (1 + 0 * 0.1) = 1.0
[0077] Then, calculate the R & D stage distribution factor F phase = (1 + 0.8 * 1 + 0.6 * 1 + 0 * 1) / 4 = 0.6. Finally, calculate the final clustering weight W cluster = (5.0 + 2.8 + 1.8 + 1.0) * (1 + log(4)) * 0.6 = 15.18.
[0078] The weight of the clustering structure affects the weight evaluation of its member molecules, and the R & D status weight of the molecules in turn affects the weight calculation of the fragments. The statistical distribution of the fragments will feedback and affect the overall importance evaluation of the molecules. By statistically analyzing the occurrence frequency and the weighted average of the R & D status of the fragments in different R & D stages, calculate the weight of the fragment structure: W frag (f) = F global (f) · W state (f),
[0079] In the formula, F global (f) is the global frequency of fragment f, N occur (f) is the number of occurrences of fragment f, N total is the total number of molecules, I(f, c i ) is the indicator function. When molecule c i contains fragment f, I(f, c i ) is 1, otherwise it is 0. W state (f) is the weighted average of the R & D status, W dev (c i ) is the R & D status weight of molecule c i . This calculation method reflects the universality (global frequency) of the fragments, the influence of the R & D status of the molecules where they are located, and the importance of each occurrence on average.
[0080] Assume that the total number of molecules (N total ) = 100. Consider a specific fragment f that appears in drug molecules at 3 different R & D progress stages: Molecule C1: (Phase = 3); Molecule C2: (Phase = 2); Molecule C3: (Phase = 0). First, calculate the R & D status weight W dev of each molecule containing this fragment based on the R & D status weight calculation formula of the compound:
[0081] C1 (Phase III): W dev(C1) = 1.0*(1+3*0.6) = 2.8
[0082] C2(stage II):W dev (C2) = 1.0*(1+2*0.4) = 1.8
[0083] C3(preclinical):W dev (C3) = 1.0*(1+0*0.1) = 1.0
[0084] Then, calculate the global frequency F global (f): F global (f) = N occur (f) / N total =3 / 100=0.03, calculate the state weighted average W state (f):W state (f) = ∑(W dev (c i) *I(f,c i )) / N occur (f) = (2.8 + 1.8 + 1.0) / 3 = 1.87. Finally, the segment weight W is calculated. frag (f):W frag (f) = F global (f)*Wstate(f)=0.03*1.87=0.0561.
[0085] The weight calculation formulas of the above-mentioned different hierarchical structures can reflect the synergistic effect of weights. A top-down and bottom-up weight transfer system is formed based on the R&D status weight, cluster structure weight and fragment structure weight. Specifically, the top-down weight transfer is reflected in that the cluster structure weight acts as a correction factor to affect the weight evaluation of its member molecules. Specifically, the weight of the cluster structure will affect the weight evaluation of all member molecules in the cluster, thereby correcting the potential value of each molecule. The R&D status weight of the molecule is passed as a calculation factor to the weight calculation of the fragments that constitute the molecule. The higher the R&D status weight of a molecule, the greater the progress made by the molecule in the R&D process and the stronger its R&D potential, thereby affecting the weight evaluation of its related fragments.
[0086] Bottom-up weight feedback is reflected in the following: the weight statistical distribution of the fragments affects the importance assessment of the original molecule through the R&D stage distribution factor feedback. The importance of the fragments is fed back to the level of the entire molecule, which can help optimize the potential assessment of the molecule. Member molecules with high R&D status weights increase the overall weight of the cluster to which they belong through the distribution factor. This means that if a molecule progresses well in the R&D process, it not only increases its importance in the entire molecular structure, but also increases the overall value of the cluster to which the molecule belongs.
[0087] By focusing on the weight evaluation of the scaffold structure, scaffold structures with potential for drug development can be screened out. These scaffolds usually have high importance in different molecules. By evaluating the fragment weights, fragments that make important contributions to the drug efficacy can be identified. These fragments are the important basis for drug activity. Based on the weight system, it is possible to evaluate whether a newly designed molecule has high development potential. By evaluating the weights between fragments and their interactions, the feasibility of different fragment combinations can be predicted, thus helping to construct drug candidate molecules with good activity.
[0088] Return Figure 1 As shown, finally in step S106, the activity relationship between the compound and the target is determined, and based on the activity relationship, the evidence supporting the activity relationship, and the weight of the R & D status, the target association credibility score is calculated.
[0089] In some embodiments of the present disclosure, a threshold calculation method is used for target activity determination and prediction of target association. According to activity data such as half maximal inhibitory concentration, half maximal effective concentration, inhibition constant, dissociation constant, etc., it is determined whether there is an activity relationship between the compound and the target:
[0090]
[0091] In the formula, IC 50 is the half maximal inhibitory concentration, EC 50 is the half maximal effective concentration, K i is the inhibition constant, K d is the dissociation constant, 1 indicates an activity relationship, and 0 indicates no activity relationship. For example, for compound C1 against target T1: IC 50 = 100 nM, 100 nM = 0.1 μM, which is less than 1 μM. Therefore, R target (C1, T1) = 1, that is, there is an activity relationship.
[0092] Then, based on the R & D status weight of the compound, the activity relationship determination result, and the number of evidences supporting the activity relationship (the number of evidences can be the number of independent experiments or the number of literature and patents to which the experiments belong), the credibility score of the target association is calculated:
[0093] Conf(c, t) = W dev (c)·R target (c, t)·(1 + ln(N evidence ))
[0094] Among them, W dev (c) is the R & D status weight of the compound, R target (c, t) is the determination result of the target relationship, and N evidenceThe number of evidences supporting the target association relationship, where the number of evidences includes the number of independent experiments and the number of literatures or patents to which the trials belong.
[0095] For example, the R & D status is clinical phase II, W dev = 1.8, and the activity data includes IC 50 = 100 nM, R target = 1. If the number of trial evidences supporting this activity data is 2, then the target association credibility score: Conf(C1,T1) = 1.8 * 1 * (1 + log(2)) = 1.8 * 1.301 = 2.34.
[0096] The above-mentioned target association evaluation supports multi-source data integration, which helps to comprehensively consider information from various aspects such as experimental results, literature reports, and theoretical models, and improves the depth and breadth of analysis. Therefore, the structure-activity relationship can combine the structures and activities of different compounds with target information, provide comprehensive pharmacodynamic prediction and structural optimization guidance, and help researchers make more accurate decisions in the drug design process.
[0097] Figure 3 FIG. 16 is a schematic structural diagram of a drug structure-activity relationship analysis device 300 based on a multi-level structure according to an embodiment of the present disclosure. Referring to Figure 3 as shown, the device 300 includes a multi-level structure characterization module 310, a weight calculation module 320, and a target association evaluation module 330.
[0098] Among them, the multi-level structure characterization module 310 can perform multi-level molecular structure characterization on compounds, and the multi-level molecular structure includes a clustering structure, an original structure, and a fragment structure.
[0099] First, the multi-level structure characterization module 310 collects experimental data including compound structures and related biological activities. Based on similarity analysis, multiple compounds are divided into different clusters, and the compounds within the clusters have similar structural characteristics, obtaining a clustering structure. The original structure refers to the basic structural unit of each compound. The fragment structure is to disassemble the compound into smaller structural units, usually fragments that can react independently or participate in biological actions.
[0100] The weight calculation module 320 can calculate the weights of the clustering structure, the original structure, and the fragment structure based on the R & D status of the compound.
[0101] The drug R & D process usually goes through different stages, such as screening, optimization, pre - clinical stage, etc. The characteristics of compounds in different R & D stages, such as their biological activities and stabilities, will vary. By combining the R & D status of molecules, the weights of each clustering structure, original structure, and fragment structure are calculated. These weights reflect the importance of each structure in drug development. For example, compounds in the early stage may have lower weights, while structures in the optimization stage may have more practical application value. A weight transfer system is established to simulate the mutual relationships and influences between structures. Through this weight transfer method, the relationships between structures can be dynamically adjusted, and their changes can be traced during the R & D process.
[0102] The target - related evaluation module 330 can determine the activity relationship between a compound and a target, and calculate the target - related credibility score based on the activity relationship, the evidence supporting the activity relationship, and the R & D status weight of the compound. By calculating the target - related credibility score, a quantitative evaluation of the activity relationship between the compound and the target can be achieved, providing a more objective standard for drug screening.
[0103] Figure 4 is a schematic block diagram of a computing device according to an embodiment of the present disclosure. As Figure 4 shown, the computing device 400 may include a processor 410 and a memory 420 storing a computer program. When the computer program is executed by the processor 410, the computing device 400 can execute the steps of the drug structure - activity relationship analysis method 100 based on a multi - level structure as Figure 1 shown. In one example, the computing device 400 can be a computer device or a cloud computing node.
[0104] In an embodiment of the present disclosure, the processor 410 can be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi - core processor architecture, etc. The memory 420 can be any type of memory implemented using data storage technology, including but not limited to random access memory, read - only memory, semiconductor - based memory, flash memory, disk memory, etc.
[0105] In addition, in an embodiment of the present disclosure, the computing device 400 may also include an input device 430, such as a keyboard, a mouse, etc., for inputting test data. Additionally, the computing device 400 may further include an output device 440, such as a display, etc., for outputting the structure - activity relationship knowledge graph.
[0106] In other embodiments of the present disclosure, a computer - readable storage medium storing a computer program is also provided, wherein the computer program can implement the steps of the drug structure - activity relationship analysis method 100 based on a multi - level structure as Figure 1 shown when executed by a processor.
[0107] In summary, the method and apparatus for analyzing the structure-activity relationship of drugs based on a multi-level structure according to the embodiments of the present disclosure can help to comprehensively understand the overall picture of a molecule and its potential pharmacodynamic characteristics by performing clustering on compounds and multi-level characterization of the original structure and fragment structure. According to the R & D status of the compounds, the weights of different-level structures are dynamically calculated, which can improve the timeliness of analysis and the accuracy of identifying key pharmacophores, thereby better guiding drug development. Especially in a complex drug system, the structural part that can effectively bind to the target can be clearly identified. By calculating the target association credibility score, a quantitative evaluation of the activity relationship between the compound and the target can be achieved, providing a more objective standard for drug screening and more reliable decision-making support.
[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of apparatuses and methods according to multiple embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, the segment of a program, or the part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0109] Unless the context clearly indicates otherwise, the singular forms of the words used in this specification and the appended claims include the plural, and vice versa. Thus, when referring to the singular, the plural of the corresponding term is generally included. Similarly, the terms "comprising" and "including" shall be interpreted as inclusive rather than exclusive. Likewise, the term "including" and "or" shall be interpreted as inclusive, unless such an interpretation is explicitly prohibited in this specification. Where the term "example" is used in this specification, especially when it is located after a group of terms, the "example" is merely exemplary and explanatory, and should not be considered exclusive or extensive.
[0110] Further aspects and scopes of adaptability become apparent from the description provided herein. It should be understood that the various aspects of the present application can be implemented alone or in combination with one or more other aspects. It should also be understood that the description herein and the specific embodiments are for illustrative purposes only and are not intended to limit the scope of the present application.
[0111] The above has described several embodiments of the present disclosure in detail. However, it is obvious that those skilled in the art can make various modifications and variations to the embodiments of the present disclosure without departing from the spirit and scope of the present disclosure. The protection scope of the present disclosure is defined by the appended claims.
Claims
1. A method for analyzing the structure-activity relationship of drugs based on a multi-level structure, characterized in that, The method includes: Performing multi-level molecular structure characterization on a compound, where the multi-level molecular structure includes a clustering structure, a raw structure, and a fragment structure; Calculating the weights of the clustering structure, the raw structure, and the fragment structure based on the R & D status of the compound; and Determining the activity relationship between the compound and the target, and calculating the target association credibility score based on the activity relationship, the evidence supporting the activity relationship, and the R & D status weight of the compound.
2. The method for analyzing the structure-activity relationship of drugs based on a multi-level structure according to claim 1, characterized in that, The performing multi-level molecular structure characterization on a compound, where the multi-level molecular structure includes a clustering structure, a raw structure, and a fragment structure includes: Collecting experimental data, clustering the molecular structures in the experimental data using the Butina clustering method to obtain the clustering structure of the molecules; Converting the molecular structure into the standard SMILES format, and performing stereochemical processing and centering processing to obtain the raw structure of the molecules; and Decomposing the molecular structure into fragments based on molecular fragmentation rules or chemical synthesis feasibility, screening and removing duplicates of the fragments to obtain the fragment structure of the molecules.
3. The method for analyzing the structure-activity relationship of drugs based on a multi-level structure according to claim 2, characterized in that The using the Butina clustering method to cluster the molecular structures in the experimental data to obtain the clustering structure of the molecules includes: Using the Murcko skeleton extraction method to remove the side chains in the molecular structure to obtain a molecular skeleton that only contains the ring system of the molecule and the atoms connecting the ring systems; Converting the molecular skeleton into a fingerprint and calculating the similarity between different fingerprints; Setting a similarity threshold and a clustering size threshold, and dividing the molecular skeletons into multiple clusters based on the similarity; and Calculating the average similarity of each molecular skeleton within the cluster to other members, and selecting the molecular skeleton with the highest average similarity as the representative structure of the cluster.
4. The method for analyzing the structure-activity relationship of drugs based on a multi-level structure according to claim 2, wherein The converting the molecular structure into the standard SMILES format, and performing stereochemical processing and centering processing to obtain the raw structure of the molecules includes: Converting the molecular structure into the standard SMILES format; Retaining the known stereochemical information of the molecular structure, removing the undefined stereocenter markers, and unifying the representation of cis-trans isomers; and Removing the inorganic salt ions of the molecular structure, and standardizing the charge state and protonation state.
5. The method for analyzing the structure-activity relationship of drugs based on a multi-level structure according to claim 2, wherein The decomposing the molecular structure into fragments based on molecular fragmentation rules or chemical synthesis feasibility, screening and removing duplicates of the fragments to obtain the fragment structure of the molecules includes: Setting molecular fragmentation rules, where the rules include: the fragment contains at least one ring system, and the number of heavy atoms in the fragment is not less than 8, and the size of the largest fragment does not exceed 70% of the number of heavy atoms in the original molecule; Using the rule-based RECAP method to decompose the molecule into fragments according to the reactivity and synthesis pathway of the molecule, or using the BRICS method based on chemical synthesis feasibility to generate fragments by cutting the rotatable bonds in the molecule; Setting screening criteria for the fragments including molecular weight range, number of rotatable bonds, number of hydrogen bond donors / acceptors, screening out the fragments that meet the screening criteria, and removing duplicates of the screened fragments to obtain multiple fragment structures.
6. The method for analyzing the structure-activity relationship of drugs based on a multi-level structure according to claim 1, characterized in that, The calculating the weights of the clustering structure, the raw structure, and the fragment structure based on the R & D status of the compound includes: Calculate the weight W of the R & D status of the compound through the following formula dev (c): W dev (c) = Base W ·(1 + ∑(Phase i ·K i )) or W dev (c) = ·Base W ·(1 + max(Phase)·K phase ) In the formula, Base W represents the basic weight value, which is set to 1.0, and Phase i is the R & D stage identification variable, and its values are: pre - clinical = 0, clinical phase I = 1, clinical phase II = 2, clinical phase III = 3, marketing approval = 4, and K i is the influencing factor for each stage, and K 临床前 = 0.1; K 临床I期 = 0.2; K 临床II期 = 0.4; K 临床III期 = 0.6; K 上市批准 = 1.0; Calculate the cluster structure weight W according to the R & D status weight of each member in the cluster, the R & D stage distribution factor, and the number of internal members in the cluster cluster (cls): W cluster (cls) = ∑(W dev (c i ))·(1 + ln(N members ))·F phase Wherein, W dev(ci) is the sum of the R & D status weights of the clustering member molecules, N members is the number of clustering member molecules, and F phase is the R & D stage distribution factor: Among them, N approved is the number of molecules with marketing approval, N phaseIII is the number of molecules in clinical phase III, N phaseII is the number of molecules in clinical phase II, N phaseI is the number of molecules in clinical phase I; Calculate the fragment structure weight W by statistically counting the occurrence frequency of fragments in different R & D stages and the weighted average of R & D status frag (f): W frag (f) = F global (f)·W state (f), where where F global (f) is the global frequency of segment f, N occur (f) is the number of occurrences of segment f, N total is the total number of molecules, I(f, c i ) is an indicator function, when molecule c i contains segment f, I(f, c i ) is 1, otherwise 0, W state (f) is the weighted average of the R & D status, W dev (c i ) is the R & D status weight of molecule c i ; and A top-down and bottom-up weight transfer system is formed based on the R & D status weight, the clustering structure weight, and the fragment structure weight.
7. The method for analyzing the structure-activity relationship of drugs based on a multi-level structure according to claim 1, characterized in that Determining the activity relationship between the compound and the target, and calculating the target association credibility score based on the activity relationship, the evidence supporting the activity relationship, and the R & D status weight, including: According to the activity data including the half inhibitory concentration, the half effective concentration, the inhibition constant, and the dissociation constant, using a threshold calculation method to determine whether there is an activity relationship between the compound and the target: wherein, IC 50 is the half inhibitory concentration, EC 50 is the half effective concentration, K i is the inhibition constant, K d is the dissociation constant, 1 indicates an active relationship, and 0 indicates no active relationship; Calculating the target association credibility score based on the R & D status weight of the compound, the determination result of the activity relationship, and the number of evidences supporting the activity relationship: Conf(c,t) = W dev (c)·R target (c,t)·(1 + ln(N evidence )); Among them, W dev (c) is the R & D status weight of the compound, R target (c, t) is the determination result of the target relationship, N evidence is the number of evidences supporting the target association relationship, and the number of evidences includes the number of independent experiments and the number of literatures or patents to which the experiments belong.
8. A drug structure-activity relationship analysis device based on a multi-level structure, characterized in that The device includes: A multi-level structure characterization module for performing multi-level molecular structure characterization on the compound, where the multi-level molecular structure includes a clustering structure, an original structure, and a fragment structure; A weight calculation module for calculating the weights of the clustering structure, the original structure, and the fragment structure based on the R & D status of the compound; and A target association evaluation module for determining the activity relationship between the compound and the target, and calculating the target association credibility score based on the activity relationship, the evidence supporting the activity relationship, and the R & D status weight of the compound.
9. A computing device, characterized in that, Including: At least one processor; And At least one memory storing a computer program; Wherein, when the computer program is executed by the at least one processor, the computing device is caused to execute the steps of the multi-level structure-based drug structure-activity relationship analysis method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the multi-level structure-based drug structure-activity relationship analysis method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Repositioning drug discovery method based on integration of a plurality of transcriptome data sets and drug target information
CN108694991A
Data analysis method and device for target drug
CN112382362A
Intelligent drug molecule generation method based on reinforcement learning and docking
CN113488116A
Drug-target interaction prediction method fusing multi-layer drug structure information
CN114067905A
Method and device for screening molecules and application thereof
CN114300067A
Cited By
N-type polymer thermoelectric performance prediction method based on structural decoupling and neural network
CN122157859A
N-type polymer thermoelectric performance prediction method based on structural decoupling and neural network
CN122157859B