Method and device for analyzing structure-activity relationship of drugs based on multi-level structure

By employing a multi-level structure-activity relationship (SCR) analysis method, compounds are clustered, and their original and fragment structures are characterized. The weights are dynamically adjusted based on the compound's development status, which solves the problem of integrating multi-source data in traditional methods. This enables quantitative assessment of the relationship between compounds and target activity, improving the accuracy and efficiency of drug design.

CN120280031BActive Publication Date: 2025-12-30BEIJING YAODU PHARMACEUTICAL TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510278030.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-12-30
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

In the development of small molecule drugs, existing technologies, particularly traditional structure-activity relationship analysis methods, struggle to simultaneously consider the influence of macroscopic and microscopic molecular structures and cannot effectively integrate data from different sources, resulting in a lack of accuracy and efficiency in drug design.

Method used

A multi-level structure-activity relationship analysis method is adopted. By clustering compounds, characterizing their original structure and fragment structure, calculating the weights of different levels of structure, and dynamically adjusting the weights based on the compound's development status, combined with target correlation assessment, a quantitative assessment of the activity relationship between the compound and the target is achieved.

Benefits of technology

It improves the accuracy and efficiency of structure-activity relationship analysis, enabling better identification of key pharmacophores, providing objective drug screening criteria, and guiding drug development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280031B_ABST
    Figure CN120280031B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a multi-level structure-based drug structure-activity relationship analysis method and device, the method comprising: performing multi-level molecular structure characterization on a compound, the multi-level molecular structure comprising cluster structure, original structure and fragment structure; calculating the weights of the cluster structure, the original structure and the fragment structure based on the research and development state of the compound; and determining the activity relationship between the compound and a target point, and calculating a target point association credibility score based on the activity relationship, evidence supporting the activity relationship and the research and development state weight of the compound.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of drug development, and more specifically, to a method, apparatus, computing device, and storage medium for drug structure-activity relationship analysis based on a multi-level structure. Background Technology

[0002] Small molecule drugs are typically organic compounds with a molecular weight below 1000 Daltons. Structure-Activity Relationship (SAR) analysis is a core component of small molecule drug development. SAR refers to the relationship between the chemical structure of a drug or other physiologically active substance and its physiological activity, providing a theoretical basis for new drug development. Traditional SAR analysis methods focus on the intuitive correlation between a compound's chemical structure and biological activity. However, with the rapid increase and growing complexity of drug development data, traditional SAR analysis methods face many challenges in practical applications.

[0003] On the one hand, the biological activity of a molecule is determined not only by its macroscopic structure (such as molecular size, shape, and functional groups) but also by more microscopic factors (such as electronic effects, stereochemical properties, and solvent effects). On the other hand, in addition to the structural and biological activity data of compounds, drug development involves a large amount of other data, such as clinical trial data, experimental methods, toxicity data, and pharmacokinetic information. The activity of a compound may vary significantly under different experimental methods, and the status of clinical trials can also affect efficacy. Furthermore, different experimental results and data sources may conflict or be inconsistent. In this context, establishing a reasonable data integration mechanism, especially how to handle data differences from different platforms, experimental methods, and even different researchers, and effectively integrate and analyze them, has become an important problem that needs to be solved.

[0004] Existing technologies employ molecular fingerprint-based similarity search methods, converting molecular structures into binary fingerprints and analyzing structure-activity relationships by calculating the similarity between molecular fingerprints. While this method can capture local structural features, it relies on a single-level structural characterization and struggles to simultaneously consider the basic skeletal structure and local substituent features. Substructure matching-based SAR analysis identifies common skeletal structures and variable substituents within molecules, analyzing the impact of different substituents on activity. However, the assessment of the importance of molecular fragments (such as functional groups and side chains) at different development stages is somewhat vague, lacking a quantitative assessment of the reliability of target associations. Machine learning-based activity prediction methods utilize machine learning algorithms to establish predictive models between molecular descriptors and activity. Although this method can establish a model of the relationship between molecular descriptors and activity, integrating activity data from different sources is difficult, leading to the omission or misjudgment of certain potential targets. These models typically lack dynamic adaptation to the constantly changing needs and goals in drug development, failing to provide targeted strategies or assessments at different stages. For example, in early stages, there may be a greater need to predict and screen compounds, while in later stages, the focus may shift to drug toxicity and pharmacokinetic properties. Therefore, accurate drug design cannot be achieved solely through activity prediction models. Summary of the Invention

[0005] To improve the accuracy and efficiency of SAR analysis, the embodiments described herein provide a method, apparatus, computing device, and computer-readable storage medium storing computer programs for drug structure-activity relationship analysis based on a multi-level structure.

[0006] According to a first aspect of this disclosure, a method for analyzing the structure-activity relationship of a drug based on a multi-level structure is provided, comprising: characterizing the multi-level molecular structure of a compound, wherein the multi-level molecular structure includes a cluster structure, a primitive structure, and a fragment structure; calculating the weights of the cluster structure, the primitive structure, and the fragment structure based on the compound's development status; and determining the activity relationship between the compound and a target, and calculating a target association confidence score based on the activity relationship, evidence supporting the activity relationship, and the compound's development status weight.

[0007] In some embodiments of this disclosure, multi-level molecular structure characterization of compounds, including clustered structures, original structures, and fragment structures, includes: collecting experimental data; using the Butina clustering method to cluster the molecular structures in the experimental data to obtain the clustered molecular structures; converting the molecular structures into the standard SMILES format and performing stereochemical processing and centering processing to obtain the original molecular structures; and decomposing the molecular structures into fragments based on molecular fragmentation rules or chemical synthesis feasibility, screening and deduplicating the fragments to obtain the fragment structures of the molecules.

[0008] In some embodiments of this disclosure, the molecular structures in the experimental data are clustered using the Butina clustering method to obtain the molecular cluster structure, which includes: removing side chains from the molecular structure using the Murcko backbone extraction method to obtain a molecular backbone containing only the ring system and atoms connecting the ring system; converting the molecular backbone into fingerprints and calculating the similarity between different fingerprints; setting a similarity threshold and a cluster size threshold, and dividing the molecular backbone into multiple clusters based on the similarity; calculating the average similarity between each molecular backbone in the cluster and other members, and selecting the molecular backbone with the highest average similarity as the representative structure of the cluster.

[0009] In some embodiments of this disclosure, converting the molecular structure to the standard SMILES format and performing stereochemical and centering processes to obtain the original molecular structure includes: converting the molecular structure to the standard SMILES format; retaining known stereochemical information of the molecular structure, removing undefined stereocenter markers, and standardizing the representation of cis-trans isomers; and removing inorganic salt ions from the molecular structure, and standardizing charge states and protonated states.

[0010] In some embodiments of this disclosure, the molecular structure is decomposed into fragments based on molecular fragmentation rules or chemical synthesis feasibility, and the fragments are screened and deduplicated to obtain fragment structures of the molecule. This includes: setting molecular fragmentation rules, which include: the fragment must contain at least one ring system, the number of heavy atoms in the fragment must be no less than 8, and the size of the largest fragment must not exceed 70% of the number of heavy atoms in the original molecule; using the rule-based RECAP method to decompose the molecule into fragments according to the reactivity and synthetic pathway of the molecule, or using the BRICS method based on chemical synthesis feasibility to generate fragments by cleaving the rotatable bonds in the molecule; setting screening criteria for the fragments, including molecular weight range, number of rotatable bonds, and number of hydrogen bond donors / acceptors, screening out fragments that meet the screening criteria, and deduplicating the screened fragments to obtain multiple fragment structures.

[0011] In some embodiments of this disclosure, calculating the weights of the cluster structure, original structure, and fragment structure based on the compound's development status includes: calculating the compound development status weight W using the following formula. dev (c):

[0012] W dev (c)=BaseW·(1+∑(Phase i ·K i ))or

[0013] W dec (c) = Base W ·(1+max(Phase)·K phase In the formula, BaseW This represents the base weight value, set to 1.0. (Phase) i The variable representing the research and development stage is: preclinical = 0, Phase I clinical = 1, Phase II clinical = 2, Phase III clinical = 3, market approval = 4, K. i K represents the influencing factors at each stage. 临床前 =0.1; K 临床I期 =0.2; K 临床II期 =0.4; K 临床III期 =0.6; K 上市批准 =1.0;

[0014] The weight W of the cluster structure is calculated based on the R&D status weight, R&D stage distribution factor, and number of members within each member of the molecular cluster. cluster (cls):

[0015] W cbuster (cls)=∑(W dev (c i ))·(1+ln(N members ))·F phase

[0016] In the formula, W dev(ci) N is the sum of the research and development status weights of the cluster members. members F represents the number of cluster members. phase Distribution factor for the R&D stage:

[0017]

[0018] Where, N approved N represents the number of molecules approved for market launch. phaseIII N represents the number of molecules in Phase III clinical trials. phaseII N represents the number of molecules in Phase II clinical trials. phaseI This represents the number of molecules in Phase I clinical trials.

[0019] The weight W of the fragment structure is calculated by statistically analyzing the frequency of occurrence of fragments at different R&D stages and the weighted average of R&D status. frag (f):

[0020] W frag (f)=F global (f)·W state (f), where,

[0021] In the formula, F global (f) represents the global frequency of segment f, N occur (f) represents the number of occurrences of segment f, N totalI(f,c) represents the total number of molecules. i ) is an indicator function, when the molecule c i When fragment f is included, I(f,c) i W is 1 if it is not 0 otherwise. state (f) is the weighted average of the R&D status, W dev (c i ) is the molecule c i The R&D status weights; based on the R&D status weights, cluster structure weights, and fragment structure weights, a top-down and bottom-up weight transfer system is formed.

[0022] In some embodiments of this disclosure, determining the activity relationship between the compound and the target, and calculating the target association confidence score based on the activity relationship, evidence supporting the activity relationship, and R&D status weights, includes: determining whether there is an activity relationship between the compound and the target using a threshold calculation method based on activity data including half-maximal inhibitory concentration, half-maximal effective concentration, inhibition constant, and dissociation constant.

[0023]

[0024] In the formula, IC 50 The half-maximal inhibitory concentration (IC50) and EC50 50 It is the half-maximal effective concentration, K i It is the suppression constant, K d It is the dissociation constant, where 1 indicates an active relationship and 0 indicates no active relationship;

[0025] Based on the compound's research and development status weight, the activity relationship determination result, and the amount of evidence supporting the activity relationship, a confidence score for target association is calculated:

[0026] Conf(c,t) = W dev (c)·R target (c,t)·(1+ln(N) evidence ));

[0027] Among them, W dev (c) represents the research and development status weight of the compound, R target (c,t) represents the target-point relationship judgment result, N evidence The amount of evidence supporting the association with the target includes the number of independent experiments and the number of documents or patents related to the experiments.

[0028] According to a second aspect of this disclosure, a drug structure-activity relationship analysis device based on a multi-level structure is provided, comprising a multi-level structure characterization module, a weight calculation module, and a target association assessment module. The multi-level structure characterization module is used to characterize the multi-level molecular structure of a compound, including clustered structures, original structures, and fragment structures. The weight calculation module is used to calculate the weights of the clustered structures, original structures, and fragment structures based on the compound's development status. The target association assessment module is used to determine the activity relationship between the compound and a target, and to calculate a target association confidence score based on the activity relationship, evidence supporting the activity relationship, and the compound's development status weights.

[0029] According to a third aspect of this disclosure, a computing device is provided, comprising at least one processor and at least one memory storing a computer program. When the computer program is executed by the at least one processor, the computing device causes the computing device to perform the steps of the drug structure-activity relationship analysis method based on a multi-level structure as described in the first aspect of this disclosure.

[0030] According to a fourth aspect of this disclosure, a computer-readable storage medium storing a computer program is provided, wherein the computer program, when executed by a processor, implements the steps of the drug structure-activity relationship analysis method based on a multi-level structure according to a first aspect of this disclosure.

[0031] The drug structure-activity relationship analysis method and apparatus based on multi-level structures provided in this disclosure, through multi-level characterization of compounds by clustering, original structure, and fragment structure, helps to comprehensively understand the overall molecule and its potential pharmacological characteristics. Dynamically calculating the weights of different structural levels according to the compound's development status improves the timeliness of analysis and the accuracy of identifying key pharmacophores, thereby better guiding drug development. Especially in complex drug systems, it can clearly identify structural parts that can effectively bind to the target. By calculating the target association confidence score, a quantitative assessment of the activity relationship between the compound and the target can be achieved, providing a more objective standard for drug screening and offering more reliable decision support. Attached Figure Description

[0032] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure, wherein:

[0033] Figure 1 This is an exemplary flowchart of a drug structure-activity relationship analysis method 100 based on a multi-level structure according to an embodiment of the present disclosure;

[0034] Figure 2An exemplary schematic diagram of a multi-level molecular structure characterization according to an embodiment of the present disclosure is shown;

[0035] Figure 3 This is a schematic diagram of the structure of a multi-level structure-activity relationship analysis device 300 according to an embodiment of the present disclosure;

[0036] Figure 4 This is a schematic block diagram of a computing device according to embodiments of the present disclosure.

[0037] It should be noted that the elements in the attached diagram are schematic and not drawn to scale. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are also within the scope of protection of this disclosure.

[0039] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having meanings consistent with their meanings in the context of the specification and in the relevant art, and shall not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. Furthermore, terms such as “first” and “second” are used only to distinguish one component (or part of a component) from another component (or another part of a component).

[0040] To address the limitations of existing technologies in effectively processing the correlation information between compound structures and targets at different levels, which leads to insufficient exploration and utilization of structure-activity relationships (SARs) and an inability to effectively adapt to the ever-changing needs during drug development, this disclosure proposes a multi-level structure-based drug SAR analysis method. This method aims to more accurately capture different characteristics of compounds by analyzing their molecular structures at multiple levels, thereby enabling in-depth analysis of their correlation with targets. Furthermore, by dynamically adjusting the weights between compounds and targets based on different stages and information changes in the drug development process, it accurately reflects the structural features and target correlations at different stages, thus improving the accuracy of SAR mapping. By constructing a quantitative target correlation assessment, a systematic and quantitative analysis of the relationship between compounds and targets can be achieved, further optimizing the drug screening and design process and improving drug development efficiency.

[0041] Figure 1 An exemplary flowchart of a drug structure-activity relationship analysis method 100 based on a multi-level structure according to an embodiment of the present disclosure is shown.

[0042] exist Figure 1 At box S102, the compound is characterized by multi-level molecular structure, which includes cluster structure, original structure and fragment structure.

[0043] Existing experimental data can be obtained by consulting relevant academic papers and research reports in academic databases such as Google Scholar and PubMed, or by accessing publicly available datasets such as ChEMBL and PubChem. This embodiment of the present disclosure uses computational methods to represent the molecular structures in the experimental data in a hierarchical manner, mainly divided into three levels: a clustered structure layer, a primitive structure layer, and a fragment structure layer.

[0044] In this context, cluster structures represent the main categories of the molecular skeleton and are used to extract the core structural features of compounds. The Butina clustering method can be used to cluster the molecular structures in the experimental data to obtain the molecular cluster structures.

[0045] In one embodiment of this disclosure, the Murcko skeleton extraction method can be used to remove side chains from the molecular structure, yielding a molecular skeleton containing only the ring system and the atoms connecting the ring systems. By removing side chains and substituents (these acyclic parts are typically variable parts of the molecule and do not affect the basic skeletal structure), the important ring systems and their connections are preserved, extracting the portion representing the molecular skeleton. For example, the original compound is C9H. 12 O2 (hypothetically a drug molecule containing a ring system and multiple functional groups) after Murcko skeleton extraction may result in a skeleton that is only a benzene ring or cyclohexene structure, with all acyclic parts such as methyl and hydroxyl groups removed.

[0046] Then, the extracted molecular backbone is converted into a fingerprint, and the similarity between different fingerprints is calculated. For example, the molecular backbone is converted into a Morgan fingerprint (this fingerprint is generated through an iterative process based on the local chemical environment of each atom within the molecule). The generation of the Morgan fingerprint is based on the following steps:

[0047] Initialization: Fingerprint generation begins by assigning a unique label (usually an atom number or other identifier) ​​to each atom in the molecule;

[0048] Iterative Update: In each iteration, the label of each atom is updated to the hash value of the combination of its own label and the labels of its neighboring atoms. This process is typically repeated multiple times, with each iteration considering the "environment" of the atom to include more distant neighbors. In this way, the label of each atom gradually incorporates more and more information about its chemical environment;

[0049] Radius and Bit Vector: A key parameter of Morgan's algorithm is the "radius," which determines how far away the atom's neighbors are to be considered. For example, a radius of 1 means only directly connected atoms are considered, while a radius of 2 extends to atoms beyond two bonds. Ultimately, the final tag for each atom is used to update a bit vector (typically a fixed length, such as 1024 or 2048 bits). If an atom's tag is mapped to a specific bit via a hash function, that bit is set to 1.

[0050] Results: The final position vectors represent the chemical structure and properties of the entire molecule. Due to its context-based generation method, Morgan fingerprints are able to capture contextual information about atoms in a molecule, making them extremely useful for predicting molecular properties and behavior.

[0051] RDKit is an open-source cheminformatics and machine learning software primarily used for processing molecular and chemical reaction data. RDKit provides a suite of tools and algorithms that allow users to process molecular structures, perform chemical reaction simulations, analyze chemical data, and perform molecular searches and matches. In some embodiments of this disclosure, the RDKit library can be used to generate Morgan fingerprints and calculate pairwise similarities between molecules. Cosine similarity or the Tanimoto coefficient can be used to calculate the similarity between fingerprints; for example, Tanimoto(A,B)=∣A∪B∣∣A∩B∣, where ∣A∩B∣ is the number of common features between the two fingerprints, and ∣A∪B∣ is the total number of all features in the two fingerprints. The similarity score is a number between 0 and 1, where 1 indicates identical fingerprints and 0 indicates completely different fingerprints.

[0052] A similarity threshold and a cluster size threshold are set, and the molecular skeleton is divided into multiple clusters based on the similarity. The similarity threshold controls which molecules are classified into the same category, while the cluster size threshold limits the number of molecules in each cluster. These conditions enable the effective organization and classification of structurally similar molecules. Finally, the average similarity of each molecular skeleton within a cluster with other members is calculated, and the molecular skeleton with the highest average similarity is selected as the representative structure of that cluster.

[0053] The original structure is the complete molecular structure, preserving the original information of the molecule, including all atoms and bonding patterns. To ensure the consistency and comparability of structural data, the molecular structure is converted to the standard SMILES format (a linear string format for representing chemical molecular structures), and stereochemical processing and centering are performed to reconstruct the original structure.

[0054] During the standardization process, known stereochemical information of the molecular structure is retained, while undefined stereocenter markers are removed to avoid confusion or errors caused by a lack of stereo information. The representation of cis-trans isomers is standardized; for example, all cis-trans isomers are labeled with a standard cis or trans form to ensure consistent representation under different conditions. Furthermore, to ensure consistency in charge, ions, etc., of the molecular structure, inorganic salt ions are removed, and charge states and protonation states (i.e., the hydrogen ionization states of certain groups in the molecule) are standardized to ensure consistent representation of functional groups such as amino and carboxyl groups under the same pH conditions. After standardization, the molecular structure will conform to a unified format and standard.

[0055] Fragment structures are chemical structural fragments obtained through molecular fragmentation methods. They help to understand the internal details of molecules, identify key functional units, and facilitate molecular optimization and synthesis. In some embodiments of this disclosure, molecular structures are decomposed into fragments based on molecular fragmentation rules or the feasibility of chemical synthesis. These fragments are then screened and deduplicated to obtain fragment structures of the molecule.

[0056] For example, molecular fragmentation rules include: a fragment must contain at least one ring system, and the number of heavy atoms (usually carbon, nitrogen, oxygen, phosphorus, etc.) in the fragment must be at least 8. The size of the largest fragment cannot exceed 70% of the number of heavy atoms in the original molecule, so that the fragmented molecule maintains diversity among fragments of different sizes.

[0057] There are two main methods for generating molecular fragments: using the rule-based RECAP method to break down molecules into fragments based on their reactivity and synthetic pathways, or using the BRICS method based on the feasibility of chemical synthesis to generate fragments by breaking rotatable bonds in the molecule.

[0058] After fragment generation, fragments are screened and deduplicated to construct the fragment structure. Specifically, screening criteria are set, such as molecular weight range, number of rotational bonds, and number of hydrogen bond donors / acceptors, to select fragments that meet the criteria. Typically, the molecular weight of the fragment should be within a certain range. Fragments that are too large may be overly complex and difficult to synthesize, while fragments that are too small may lack sufficient chemical information. A common molecular weight range is 150–350 Daltons, which can be adjusted according to specific needs. This ensures the feasibility of the screened fragments in drug design and synthesis. The number of rotational bonds affects the flexibility of the molecule. Fragments should not contain too many rotational bonds, as their presence may reduce fragment stability and affect their feasibility as drug candidates. Typically, the number of rotational bonds in a fragment is limited to 2–5. Hydrogen bonds are an important mode of intermolecular interaction. In fragment screening, it is necessary to ensure that the fragment has an appropriate number of hydrogen bond donors and acceptors. The number of hydrogen bonds should be balanced; too many donors or acceptors may lead to unstable molecules or weakened affinity for the target. Therefore, the number of hydrogen bond donors / acceptors can be set within a specific range based on the drug activity target. To avoid interference from similar segments in subsequent analysis and selection, the selected segments are deduplicated to obtain multiple segment structures. For example, the SMILES or fingerprint representations of the segments are compared to identify and remove similar or identical segments.

[0059] Figure 2 An exemplary schematic diagram of a multi-level molecular structure characterization according to an embodiment of the present disclosure is shown. Figure 2 In the examples, the clustered structure represents the molecular skeleton, preserving the ring systems and the atoms connecting the ring systems. The original structure maintains the complete molecular structure through normalization. The fragmented structure consists of multiple constituent fragments after the original structure has been broken down.

[0060] Subsequently, in step S104, the weights of the cluster structure, the original structure, and the fragment structure are calculated based on the research and development status of the compound.

[0061] To identify scaffold, original, and fragment structures with drug development value or potential, the importance of structures at different levels can be quantitatively calculated, effectively guiding drug design and screening. When calculating the importance of structures at different levels, the compound's development status is a core influencing factor. This disclosure establishes a top-down and bottom-up weighting system based on the development status.

[0062] In some embodiments of this disclosure, updates to the compound development status can be obtained centrally from clinical, approval, or commercial databases in various countries. Development status refers to the different stages a drug is in during its development process. Different stages reflect the varying degrees a drug has progressed from laboratory research to clinical application. Therefore, each stage has a corresponding basic weight to quantify its influence in drug development, indicating that the drug possesses certain development potential regardless of the stage. Depending on whether the scenario of a drug having different development statuses in different therapeutic areas is ignored, the compound development status weight W is calculated using the following formula. dev (c):

[0063] W dev (c) = Base W ·(1+∑(Phase i ·K i ))or

[0064] W dev (c) = Base W ·(1+max(Phase)·K phase )

[0065] In the formula, Base W This represents the base weight value, set to 1.0. (Phase) i The variable representing the research and development stage has the following values: preclinical = 0; Phase I clinical = 1; Phase II clinical = 2; Phase III clinical = 3; market approval = 4, K i These are the impact factors for each stage. Each stage will have different impact factors depending on its clinical or experimental progress. An impact factor for the preclinical stage (Phase = 0) can be set: K. 临床前 =0.1; Clinical Phase I (Phase = 1) Impact Factor: K 临床I期 =0.2; Impact factor for Phase II clinical trials (Phase = 2): K 临床II期 =0.4; Impact factor for Phase III clinical trials (Phase = 3): K 临床III期 =0.6; Impact factor for market approval (Phase = 4): K 上市批准 =1.0.

[0066] For example, if a compound is in Phase II clinical trials (Phase = 2), its weight is: W dev =1.0×(1+2×0.4)=1.8, indicating that the compound has high potential in research and development.

[0067] Clustering structure is the result of grouping molecules based on the similarity of their skeletal structures. The calculation of cluster structure weights considers the R&D status, R&D stage distribution factor, and the number of members within each cluster. Based on the R&D status weight, R&D stage distribution factor, and the number of members within each cluster, the weight W of the cluster structure is calculated. cluster (cls):

[0068] W cluster (cls)=∑(W dev (c i ))·(1+ln(N members ))·F phase

[0069] In the formula, W dev(ci) N is the sum of the research and development status weights of the cluster members. members F represents the number of cluster members. phase Distribution factor for the R&D stage:

[0070]

[0071] Where, N approved N represents the number of molecules approved for market launch. phaseIII N represents the number of molecules in Phase III clinical trials. phaseII N represents the number of molecules in Phase II clinical trials. phaseI This represents the number of molecules in Phase I clinical trials.

[0072] The calculation steps include: First, based on the formula for calculating the compound development status weights, the sum of the development status weights of the cluster members is calculated. The development stage distribution factor is calculated according to the development stage distribution of the cluster members, reflecting the proportion or importance of drugs or candidate molecules at different stages. This calculation method reflects the cumulative effect of the member's development status, the logarithmic gain of the number of members, and the additional contribution of members at higher stages. Taking a cluster structure containing 4 compounds as an example, compound A: marketed drug (Phase = 4); compound B: Phase III clinical trial (Phase = 3); compound C: Phase II clinical trial (Phase = 2); compound D: preclinical trial (Phase = 0). First, the development status weight W of each compound is calculated. dev :

[0073] Compound A:W dev (A) = 1.0 * (1 + 4 * 1.0) = 5.0

[0074] Compound B:W dev (B) = 1.0 * (1 + 3 * 0.6) = 2.8

[0075] Compound C:W dev(C) = 1.0 * (1 + 2 * 0.4) = 1.8

[0076] Compound D:W dev (D) = 1.0 * (1 + 0 * 0.1) = 1.0

[0077] Then, calculate the distribution factor F of the R&D stage. phase = (1 + 0.8 * 1 + 0.6 * 1 + 0 * 1) / 4 = 0.6. Finally, calculate the final cluster weight W. cluster =(5.0+2.8+1.8+1.0)*(1+log(4))*0.6=15.18.

[0078] The weights of cluster structures influence the weight assessment of their member molecules, and the research and development status weights of molecules, in turn, affect the fragment weight calculation. The statistical distribution of fragments feeds back into the overall importance assessment of molecules. The fragment structure weights are calculated by statistically analyzing the frequency of fragment occurrence at different research and development stages and the weighted average of research and development status: W. frag (f)=F global (f)·W state (f),

[0079] In the formula, F global (f) represents the global frequency of segment f, N occur (f) represents the number of occurrences of segment f, N total I(f,c) represents the total number of molecules. i ) is an indicator function, when the molecule c i When fragment f is included, I(f,c) i W is 1 if it is 1, otherwise it is 0. state (f) is the weighted average of the R&D status, W dev (c i ) is the molecule c i The research and development status weights are calculated using this method. This calculation method reflects the universality of the fragment (global frequency), the influence of the research and development status of the molecule it belongs to, and the importance of each occurrence on average.

[0080] Assuming the total number of molecules (N) total () = 100. Consider a specific fragment f, which appears in drug molecules at three different stages of development: molecule C1 (Phase = 3); molecule C2 (Phase = 2); molecule C3 (Phase = 0). First, calculate the development state weight W of each molecule containing this fragment based on the compound development state weight calculation formula. dev :

[0081] C1 (Phase III): W dev(C1) = 1.0 * (1 + 3 * 0.6) = 2.8

[0082] C2 (Phase II): W dev (C2) = 1.0 * (1 + 2 * 0.4) = 1.8

[0083] C3 (preclinical): W dev (C3) = 1.0 * (1 + 0 * 0.1) = 1.0

[0084] Then, calculate the global frequency F. global (f): F global (f)=N occur (f) / N total =3 / 100 = 0.03, calculate the state-weighted average W state (f): W state (f)=∑(W dev (c i) *I(f,c i )) / N occur (f) = (2.8 + 1.8 + 1.0) / 3 = 1.87. The final segment weight W is calculated. frag (f): W frag (f)=F global (f)*Wstate(f)=0.03*1.87=0.0561.

[0085] The weight calculation formulas for the different hierarchical structures described above reflect the synergistic effect of weights. A top-down and bottom-up weight transfer system is formed based on R&D state weights, cluster structure weights, and fragment structure weights. Specifically, the top-down weight transfer manifests as follows: cluster structure weights act as correction factors influencing the weight evaluation of its member molecules. In other words, the weight of the cluster structure affects the weight evaluation of all member molecules within that cluster, thereby correcting the potential value of each molecule. The R&D state weight of a molecule is transferred as a calculation factor to the weight calculation of the fragments that constitute that molecule. The higher the R&D state weight of a molecule, the greater the progress made in the R&D process and the stronger its R&D potential, thus affecting the weight evaluation of its related fragments.

[0086] The bottom-up weighting feedback manifests as follows: the statistical distribution of fragment weights influences the importance assessment of the original molecule through the distribution factor feedback during the development stage. The importance of fragments is fed back to the entire molecule hierarchy, helping to optimize the molecule's potential assessment. Member molecules with high development stage weights increase the overall weight of their respective clusters through the distribution factor. This means that if a molecule progresses well during development, it not only increases its own importance within the overall molecular structure but also enhances the overall value of the cluster to which it belongs.

[0087] By focusing on weighted assessment of skeletal structures, promising skeletal structures for drug development can be screened out; these skeletal structures typically hold high importance in different molecules. Evaluation of fragment weights can identify fragments that significantly contribute to drug efficacy, forming a crucial basis for drug activity. Based on this weighted system, the research potential of newly designed molecules can be assessed. By evaluating the weights between fragments and their interactions, the feasibility of different fragment combinations can be predicted, thereby aiding in the construction of drug candidate molecules with good activity.

[0088] return Figure 1 As shown, in step S106, the activity relationship between the compound and the target is determined, and the target association confidence score is calculated based on the activity relationship, the evidence supporting the activity relationship, and the research and development status weight.

[0089] In some embodiments of this disclosure, a threshold calculation method is used to determine target activity and predict target association. The presence of an activity relationship between the compound and the target is determined based on activity data such as half-maximal inhibitory concentration (MCC), half-maximal effective concentration (MCIC), inhibition constant, and dissociation constant.

[0090]

[0091] In the formula, IC 50 The half-maximal inhibitory concentration (IC50) and EC50 50 It is the half-maximal effective concentration, K i It is the suppression constant, K d This is the dissociation constant; 1 indicates an activity relationship, and 0 indicates no activity relationship. For example, compound C1 against target T1: IC 50 =100nM, 100nM = 0.1μM, which is less than 1μM, therefore R target (C1,T1)=1, indicating an activity relationship.

[0092] Then, based on the compound's R&D status weight, the activity relationship determination result, and the number of pieces of evidence supporting the activity relationship (the number of pieces of evidence can be the number of independent trials or the number of literature or patents to which the trials belong), the confidence score of the target association is calculated:

[0093] Conf(c,t) = W dev (c)·R target (c,t)·(1+ln(N) evidence )).

[0094] Among them, W dev (c) represents the research and development status weight of the compound, R target (c,t) represents the target-point relationship judgment result, N evidenceThe amount of evidence supporting the association with the target includes the number of independent experiments and the number of documents or patents related to the experiments.

[0095] For example, if the research and development status is Phase II clinical trials, W dev =1.8, activity data includes IC 50 =100nM, R target =1, the number of experimental evidences supporting the activity data is 2, then the target association confidence score is: Conf(C1,T1)=1.8*1*(1+log(2))=1.8*1.301=2.34.

[0096] The aforementioned target association assessment supports the integration of multi-source data, helping to comprehensively consider information from experimental results, literature reports, and theoretical models, thereby enhancing the depth and breadth of the analysis. Therefore, structure-activity relationship analysis can combine the structure, activity, and target information of different compounds, providing comprehensive guidance for efficacy prediction and structure optimization, and helping researchers make more accurate decisions in the drug design process.

[0097] Figure 3 This is a schematic diagram of a multi-level structure-activity relationship analysis device 300 according to an embodiment of the present disclosure. (Refer to...) Figure 3 As shown, the device 300 includes a multi-level structure characterization module 310, a weight calculation module 320, and a target correlation evaluation module 330.

[0098] The multi-level structure characterization module 310 can perform multi-level molecular structure characterization on compounds, including cluster structure, original structure and fragment structure.

[0099] First, the multi-level structure characterization module 310 collects experimental data containing compound structures and related biological activities. Based on similarity analysis, multiple compounds are divided into different clusters. Compounds within a cluster have similar structural features, resulting in cluster structures. The original structure refers to the basic structural unit of each compound. Fragment structures break down compounds into smaller structural units, typically fragments capable of reacting independently or participating in biological processes.

[0100] The weight calculation module 320 can calculate the weights of cluster structures, original structures, and fragment structures based on the research and development status of compounds.

[0101] Drug development typically involves different stages, such as screening, optimization, and preclinical stages. Compounds at different development stages exhibit varying biological activities, stability, and other properties. By considering the molecular development status, the weights of each cluster structure, original structure, and fragment structure are calculated. These weights reflect the importance of each structure in drug development. For example, compounds from earlier stages may have lower weights, while structures from the optimization stage may have greater practical application value. A weight transfer system is established to simulate the interactions and influences between structures. This weight transfer method allows for dynamic adjustment of the relationships between structures and tracking their changes throughout the development process.

[0102] The target association assessment module 330 can determine the activity relationship between a compound and a target, and calculate a target association confidence score based on the activity relationship, the evidence supporting that activity relationship, and the compound's development status weight. By calculating the target association confidence score, a quantitative assessment of the activity relationship between a compound and a target can be achieved, providing a more objective standard for drug screening.

[0103] Figure 4 This is a schematic block diagram of a computing device according to embodiments of the present disclosure. Figure 4 As shown, the computing device 400 may include a processor 410 and a memory 420 storing a computer program. When the computer program is executed by the processor 410, the computing device 400 is made capable of performing tasks such as... Figure 1 The steps of the drug structure-activity relationship analysis method 100 based on a multi-level structure are shown. In one example, the computing device 400 may be a computer device or a cloud computing node.

[0104] In embodiments of this disclosure, processor 410 may be, for example, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a processor based on a multi-core processor architecture, etc. Memory 420 may be any type of memory implemented using data storage technologies, including but not limited to random access memory, read-only memory, semiconductor-based memory, flash memory, disk storage, etc.

[0105] Furthermore, in embodiments of this disclosure, the computing device 400 may also include an input device 430, such as a keyboard or mouse, for inputting experimental data. Additionally, the computing device 400 may also include an output device 440, such as a display, for outputting a structure-property relationship knowledge graph.

[0106] In other embodiments of this disclosure, a computer-readable storage medium storing a computer program is also provided, wherein the computer program, when executed by a processor, is capable of performing the following functions: Figure 1 The steps of the drug structure-activity relationship analysis method 100 based on multi-level structure are shown.

[0107] In summary, the drug structure-activity relationship analysis method and apparatus based on multi-level structures according to embodiments of this disclosure, through multi-level characterization of compounds by clustering, original structure, and fragment structure, helps to comprehensively understand the overall molecule and its potential pharmacological characteristics. Dynamically calculating the weights of different structural levels according to the compound's development status improves the timeliness of analysis and the accuracy of identifying key pharmacophores, thereby better guiding drug development, especially in complex drug systems, where it can clearly identify structural parts that can effectively bind to the target. By calculating the target association confidence score, a quantitative assessment of the activity relationship between the compound and the target can be achieved, providing a more objective standard for drug screening and offering more reliable decision support.

[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses and methods according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0109] Unless otherwise expressly indicated by the context, the singular form of words used herein and in the appended claims includes the plural form, and vice versa. Thus, when referring to the singular, the plural form of the corresponding term is generally included. Similarly, the terms “comprising” and “including” shall be interpreted as including rather than exclusively. Likewise, the terms “including” and “or” shall be interpreted as including unless such interpretation is expressly prohibited herein. Where the term “example” is used herein, particularly when it follows a set of terms, the “example” is merely exemplary and illustrative and should not be considered exclusive or extensive.

[0110] Further aspects and scope of adaptation become apparent from the description provided herein. It should be understood that various aspects of this application may be implemented individually or in combination with one or more other aspects. It should also be understood that the descriptions and specific embodiments herein are for illustrative purposes only and are not intended to limit the scope of this application.

[0111] Several embodiments of this disclosure have been described in detail above. However, it is obvious that those skilled in the art can make various modifications and variations to the embodiments of this disclosure without departing from the spirit and scope of this disclosure. The scope of protection of this disclosure is defined by the appended claims.

Claims

1. A method for analyzing a structure-activity relationship of a drug based on a multi-level structure, characterized by, The method comprises: performing multi-level molecular structure characterization on the compound; the multi-level molecular structure comprises a cluster structure, an original structure, and a fragment structure, the molecular structure is decomposed into fragments based on a molecular fragmentation rule or chemical synthesis feasibility, the fragments are screened and deduplicated to obtain the fragment structure of the molecule; the cluster structure represents a main category of a molecular skeleton of the compound, is used to extract a core structural feature of the compound, and the original structure is a complete molecular structure, maintains original information of the molecule, and comprises all atoms and bonding modes; calculating weights of the cluster structure, the original structure, and the fragment structure based on a research and development state of the compound, comprising: The compound development status weight is calculated by the following equation : or where Base W represents the base weight value, set to 1.0, Phase i is a variable identifying the development stage of the compound in the different therapeutic areas, taking the values: preclinical = 0, clinical phase I = 1, clinical phase II = 2, clinical phase III = 3, approved on the market = 4, K i is the impact factor for each phase, K 临床前 = 0.1; K 临床I期 = 0.2; K 临床II期 = 0.4; K 临床III期 = 0.6; K 上市批准 = 1.0; According to the research and development state weight of each member in the cluster, the research and development stage distribution factor, and the number of members inside the cluster, a cluster structure weight is calculated : where ∑(W dev (c i )) is the sum of the development status weights of the cluster member molecules, N members is the number of cluster member molecules, F phase is the development stage distribution factor: ; wherein, the number of molecules in market approval, the number of molecules in clinical phase III, the number of molecules in clinical phase II, the number of molecules in clinical phase I; The fragment structure weight is calculated by counting the frequency of the fragment in different R&D stages and the weighted average of R&D status : wherein , ; In the formula, Let N be the global frequency of segment f. occur (f) represents the number of occurrences of segment f, N total I(f,c) represents the total number of molecules. i ) is an indicator function, when the molecule c i When fragment f is included, I(f,c) i W is 1 if it is not 0 otherwise. state (f) is the weighted average of the R&D status, W dev (c i ) is the molecule c i The weight of the R&D status; and forming a top-down and bottom-up weight transfer system based on the research and development state weight, the cluster structure weight, and the fragment structure weight; focusing on weight evaluation of a skeleton structure, screening out a skeleton structure that has potential for drug research and development; identifying a fragment that has an important contribution to drug efficacy through evaluation of a fragment weight; based on the weight system, evaluating whether a new compound has high research and development potential, predicting feasibility of different fragment combinations through evaluation of the weights of the fragments and interactions therebetween, and helping to construct a drug candidate molecule with good activity; and determining an activity relationship between the compound and a target point, and calculating a target point correlation credibility score based on the activity relationship, evidence supporting the activity relationship, and a research and development state weight of the compound.

2. The method of QSAR analysis based on a multi-level structure according to claim 1, wherein The multi-level molecular structure characterization on the compound comprises: collecting test data, clustering molecular structures in the test data by using a Butina clustering method to obtain a cluster structure of the molecule; converting the molecular structure into a standard SMILES format, and performing stereochemistry processing and centering processing to obtain an original structure of the molecule.

3. The method of claim 2, wherein the method is based on a multi-level structure. The clustering of the molecular structures in the test data by using the Butina clustering method comprises: removing side chains in the molecular structure by using a Murcko skeleton extraction method to obtain a molecular skeleton comprising only a ring system of the molecule and atoms connecting the ring system; converting the molecular skeleton into a fingerprint, and calculating similarity between different fingerprints; setting a similarity threshold and a cluster size threshold, dividing the molecular skeleton into multiple clusters based on the similarity; and calculating an average similarity of each molecular skeleton in the cluster to other members, and selecting a molecular skeleton with the highest average similarity as a representative structure of the cluster.

4. The method of claim 2, wherein the method is based on a multi-level structure. The conversion of the molecular structure into a standard SMILES format and the stereochemistry processing and centering processing comprise: converting the molecular structure into a standard SMILES format; retaining known stereochemistry information of the molecular structure, removing undefined stereo center markers, and unifying representations of cis-trans isomers; and removing inorganic salt ions of the molecular structure, normalizing charge states and protonation states.

5. The multi-hierarchy-based pharmacophore analysis method according to claim 2, wherein, The decomposition of the molecular structure into fragments based on a molecular fragmentation rule or chemical synthesis feasibility, the screening and deduplication of the fragments, and the obtaining of the fragment structure of the molecule comprise: Setting a molecule fragmentation rule, the rule including: the fragment contains at least one ring system, and the number of heavy atoms in the fragment is not less than 8, and the size of the largest fragment is not more than 70% of the number of heavy atoms of the original molecule; Decomposing the molecule into fragments according to the reactivity and synthesis route of the molecule using a rule-based RECAP method, or generating fragments by cutting off rotatable bonds in the molecule using a BRICS method based on chemical synthesis feasibility; Setting a screening standard for the fragments, including a molecular weight range, a number of rotatable bonds, and a number of hydrogen bond donors / acceptors, screening the fragments that meet the screening standard, and performing a deduplication process on the screened fragments to obtain a plurality of fragment structures.

6. The multi-hierarchy-based pharmacophore analysis method according to claim 1, wherein, The method includes: Determining whether there is an activity relationship between the compound and the target point according to activity data including a half-inhibitory concentration, a half-effective concentration, an inhibition constant, and a dissociation constant; ; In the formula, IC 50 The half-maximal inhibitory concentration (IC50) and EC50 50 It is the half-maximal effective concentration, K i It is the suppression constant, K d It is the dissociation constant, where 1 indicates an active relationship and 0 indicates no active relationship; Calculating a target point association credibility score based on the activity relationship determination result, the evidence supporting the activity relationship, and the compound development state weight: ; wherein, W dev (c) is the weight of the research and development state of the compound, R target (c, t) is the target point relationship determination result, N evidence is the number of evidences supporting the target point correlation relationship, and the number of evidences includes the number of independent experiments and the number of literatures to which the experiments belong.

7. A drug structure-activity relationship analysis device based on a multi-level structure, characterized in that, The device includes: A multi-level structure representation module for multi-level molecular structure representation of a compound; the multi-level molecular structure includes a cluster structure, an original structure, and a fragment structure, the molecular structure is decomposed into fragments based on a molecule fragmentation rule or chemical synthesis feasibility, the fragments are screened and deduplicated to obtain a fragment structure of the molecule; the cluster structure represents a main category of a molecular skeleton of the compound and is used to extract core structural features of the compound, and the original structure is a complete molecular structure that maintains original information of the molecule, including all atoms and bonding modes; A weight calculation module for calculating weights of the cluster structure, the original structure, and the fragment structure based on a development state of the compound, including: The compound development status weight is calculated by the following equation : or where Base W represents the base weight value, set to 1.0, Phase i is a variable identifying the development stage of the compound in the different therapeutic areas, taking the values: preclinical = 0, clinical phase I = 1, clinical phase II = 2, clinical phase III = 3, approved on the market = 4, K i is the impact factor for each phase, K 临床前 = 0.1; K 临床I期 = 0.2; K 临床II期 = 0.4; K 临床III期 = 0.6; K 上市批准 = 1.0; According to the research and development state weight of each member in the cluster, the research and development stage distribution factor, and the number of members inside the cluster, a cluster structure weight is calculated : where ∑(W dev (c i )) is the sum of the development status weights of the cluster member molecules, N members is the number of cluster member molecules, F phase is the development stage distribution factor: ; wherein, the number of molecules in market approval, the number of molecules in clinical phase III, the number of molecules in clinical phase II, the number of molecules in clinical phase I; The fragment structure weight is calculated by counting the frequency of the fragment in different R&D stages and the weighted average of R&D status : wherein , ; In the formula, Let N be the global frequency of segment f. occur (f) represents the number of occurrences of segment f, N total I(f,c) represents the total number of molecules. i ) is an indicator function, when the molecule c i When fragment f is included, I(f,c) i W is 1 if it is not 0 otherwise. state (f) is the weighted average of the R&D status, W dev (c i ) is the molecule c i The weight of the R&D status; and Forming a top-down and bottom-up weight transfer system based on the development state weight, the cluster structure weight, and the fragment structure weight; focusing on weight evaluation of the skeleton structure to screen out skeleton structures that have potential for drug development; identifying fragments that make important contributions to drug efficacy through evaluation of the fragment weights; based on the weight system, evaluating whether a new compound has high development potential, predicting the feasibility of different fragment combinations by evaluating the weights of the fragments and their interactions, and helping to construct drug candidate molecules with good activity; and A target point association evaluation module for determining an activity relationship between a compound and a target point, and calculating a target point association credibility score based on the activity relationship, evidence supporting the activity relationship, and a development state weight of the compound.

8. A computing device, comprising: The method includes: At least one processor; And At least one memory storing a computer program; When the computer program is executed by the at least one processor, the computer program causes the computer device to perform the steps of the multi-level structure-based drug structure-activity relationship analysis method according to any one of claims 1 to 6.

9. A computer readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the multi-hierarchy-based drug structure-activity relationship analysis method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Repositioning drug discovery method based on integration of a plurality of transcriptome data sets and drug target information

    CN108694991A

  • Data analysis method and device for target drug

    CN112382362A

  • Drug molecule skeleton replacing and screening method based on deep transfer learning model

    CN115881244A