Method, system and medium for realizing improvement of ligand activity based on molecular structure analysis

Through the method based on molecular structure analysis, algorithms and models are used to quickly determine the activity optimization parameters of ligand compounds, which solves the problem of low molecular structure optimization efficiency and reliability of ligand compounds in the prior art, and achieves efficient and reliable improvement of the biological activity of ligand compounds.

CN118629540BActive Publication Date: 2025-06-13SUN YAT SEN UNIVERSITY CANCER CENTER (CANCER HOSPITAL AFFILIATED TO SUN YAT SEN UNIVERSITY CANCER RESEARCH INSTITUTE OF SUN YAT SEN UNIVERSITY) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410732914.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2025-06-13
Estimated Expiration
2044-06-06

AI Technical Summary

Technical Problem

In the prior art, the molecular structure optimization efficiency and reliability of ligand compounds are low, and it is difficult to efficiently improve the biological activity of ligand compounds.

Method used

Through a method based on molecular structure analysis, algorithms and models are used to quickly determine the activity optimization parameters of ligand compounds, and optimize the molecular structure of ligand compounds to improve their biological activity.

Benefits of technology

The molecular structure of ligand compounds is optimized with higher efficiency and reliability, improve their biological activity, and reduce their dependence on researchers' work experience and participation in manual operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118629540B_ABST
    Figure CN118629540B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and medium for improving ligand activity based on molecular structure analysis, which relates to the technical field of computational chemistry. The method includes: determining the set of positions of substitution groups on the ligand compound, determining the substitution group parameters of each ligand-derived compound, for each of the ligand-derived compounds, determining the interaction relationship between it and the receptor compound, and then according to this interaction relationship and the substitution group parameters, determining the corresponding preferred substitution group structure at each target position, and further determining the activity optimization parameters corresponding to the ligand compound. The activity optimization parameters are used to indicate that the corresponding target preferred substitution group structure is set at the target position of the ligand compound. It can be seen that the present invention can optimize the molecular structure of the ligand compound with higher efficiency and reliability to improve the biological activity of the ligand compound, and reduces the dependence on the work experience of researchers and the participation of manual operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computational chemistry technology, and in particular to a method, system and medium for improving ligand activity based on molecular structure analysis. Background Art

[0002] In the field of biomedicine, the matching relationship between ligand compounds and receptor compounds makes targeted therapy possible. Among them, if the molecular structure of the ligand compound changes, its matching ability with the receptor compound will also change, that is, the activity of the ligand compound will change. For example, if the molecular structure of the lead drug changes, its affinity with the target protein will also change. Analyzing the molecular structure of the ligand compound, thereby optimizing the molecular structure of the ligand compound and improving the activity of the ligand compound is one of the important ways to develop new targeted drugs.

[0003] In the prior art, the molecular structure of the ligand compound is optimized. On the one hand, the molecular structure optimization strategy of the ligand compound is found based on the researcher's own experience and through a large number of trial and error methods. On the other hand, the ligand compound is subjected to three-dimensional quantitative structure-activity relationship analysis with the help of analytical tools such as 3D-QSAR, thereby assisting researchers in conducting experiments. However, changes in the molecular structure of the ligand compound often produce various restrictive effects, and it is often difficult to predict whether the activity of the optimized ligand compound will be improved. Therefore, the efficiency and reliability of the molecular structure optimization methods of artificial or semi-artificial ligand compounds in the prior art are still very low.

[0004] Therefore, how to achieve molecular structure optimization of ligand compounds with higher efficiency and reliability in the process of targeted drug development and reduce the reliance on manual labor in the process of molecular structure optimization of ligand compounds is a problem that needs to be solved in the process of targeted drug development. Summary of the invention

[0005] The present invention provides a method, system and medium for improving ligand activity based on molecular structure analysis. Based on the molecular structure analysis of ligand-derived compounds, a series of algorithms and models are used to quickly provide activity optimization parameters for indicating the improvement of the biological activity of ligand compounds. The molecular structure of ligand compounds can be optimized with higher efficiency and reliability to improve the biological activity of ligand compounds, and the dependence on the work experience of researchers and the participation of manual operations can be reduced.

[0006] In order to solve the above technical problems, the first aspect of the present invention discloses a method for improving ligand activity based on molecular structure analysis, the method comprising:

[0007] Based on the molecular structure information of the ligand compound, determine the set of substituent positions of the ligand compound, where the set of substituent positions includes at least one target position at which the corresponding substituent structure can be substituted;

[0008] Obtain a plurality of ligand-derived compounds corresponding to the ligand compound, and based on the molecular structure information of the ligand-derived compounds and the set of substituent positions, determine the substituent parameters corresponding to the ligand-derived compounds, where the substituent parameters are used to represent the target substituent structure at each of the target positions of the corresponding ligand-derived compound;

[0009] Obtain the molecular structure information of the receptor compound. For each of the ligand-derived compounds, input the molecular structure information of the ligand-derived compound and the molecular structure information of the receptor compound into a pre-trained molecular docking model to obtain a first molecular docking result output by the molecular docking model, and determine the interaction relationship between the ligand-derived compound and the receptor compound based on the first molecular docking result;

[0010] For each of the ligand-derived compounds, determine the substituent activity parameter corresponding to the ligand-derived compound according to the substituent parameter corresponding to the ligand-derived compound and the interaction relationship between the ligand-derived compound and the receptor compound, where the substituent activity parameter is used to represent the affinity between each target substituent structure at each of the target positions of the ligand-derived compound and the receptor compound;

[0011] According to the substituent activity parameters corresponding to all the ligand-derived compounds, determine the corresponding dominant substituent structure at each of the target positions, where the affinity between the dominant substituent structure and the receptor compound satisfies a preset activity condition; according to the corresponding dominant substituent structure at each of the target positions, determine the activity optimization parameter corresponding to the ligand compound, where the activity optimization parameter is used to indicate to set the corresponding target dominant substituent structure at the target position of the ligand compound.

[0012] A second aspect of the present invention discloses a system for improving ligand activity based on molecular structure analysis, the system includes:

[0013] A position determination module, configured to determine the set of substituent positions of the ligand compound based on the molecular structure information of the ligand compound, where the set of substituent positions includes at least one target position at which the corresponding substituent structure can be substituted;

[0014] A parameter determination module, configured to obtain a plurality of ligand-derived compounds corresponding to the ligand compound, and determine substitution group parameters corresponding to the ligand-derived compounds according to the molecular structure information of the ligand-derived compounds and the set of substitution group positions, where the substitution group parameters are used to represent the target substitution group structures at each of the target positions of the corresponding ligand-derived compounds;

[0015] A first molecular docking module, configured to obtain the molecular structure information of the receptor compound. For each of the ligand-derived compounds, input the molecular structure information of the ligand-derived compound and the molecular structure information of the receptor compound into a pre-trained molecular docking model, obtain a first molecular docking result output by the molecular docking model, and determine the interaction relationship between the ligand-derived compound and the receptor compound based on the first molecular docking result;

[0016] An activity analysis module, configured to, for each of the ligand-derived compounds, determine substitution group activity parameters corresponding to the ligand-derived compounds according to the substitution group parameters corresponding to the ligand-derived compounds and the interaction relationship between the ligand-derived compounds and the receptor compound, where the substitution group activity parameters are used to represent the affinity magnitude between each target substitution group structure at each of the target positions of the ligand-derived compounds and the receptor compound;

[0017] An activity optimization module, configured to determine a dominant substitution group structure corresponding to each of the target positions according to the substitution group activity parameters corresponding to all the ligand-derived compounds, where the affinity between the dominant substitution group structure and the receptor compound meets a preset activity condition; and determine activity optimization parameters corresponding to the ligand compound according to the dominant substitution group structure corresponding to each of the target positions, where the activity optimization parameters are used to indicate to set a corresponding target dominant substitution group structure at the target positions of the ligand compound.

[0018] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the activity optimization module determines the activity optimization parameters corresponding to the ligand compound according to the dominant substitution group structure corresponding to each of the target positions includes:

[0019] For each of the advantageous substituent group structures corresponding to each of the target positions, a characteristic ligand-derived compound and a comparative ligand-derived compound corresponding to the advantageous substituent group structure are determined from the plurality of ligand-derived compounds. The characteristic ligand-derived compound has the advantageous substituent group structure disposed at the target position, and the comparative ligand-derived compound does not have the advantageous substituent group structure disposed at the target position and has the same target substituent group structure as the characteristic ligand-derived compound at other target positions.

[0020] For each of the advantageous substituent group structures corresponding to each of the target positions, the characteristic ligand-derived compound and the comparative ligand-derived compound corresponding to the advantageous substituent group structure are respectively combined with the receptor compound to obtain a characteristic conjugate and a comparative conjugate corresponding to the advantageous substituent group structure. According to the molecular structure information of the characteristic conjugate and the molecular structure information of the comparative conjugate corresponding to the advantageous substituent group structure, the RMSD values of the characteristic conjugate and the comparative conjugate corresponding to the advantageous substituent group relative to the receptor compound are respectively calculated, the RMSF values of the characteristic conjugate and the comparative conjugate corresponding to the advantageous substituent group are respectively calculated, and the hydrogen bond data of the characteristic conjugate and the comparative conjugate corresponding to the advantageous substituent group are respectively calculated. The hydrogen bond data includes the acceptor of the hydrogen bond, the bond length of the hydrogen bond, the bond angle of the hydrogen bond, and the occurrence frequency of the hydrogen bond. The binding free energy data of the characteristic conjugate and the comparative conjugate corresponding to the advantageous substituent group are respectively calculated. The binding free energy data includes polar binding energy data and non-polar binding energy data.

[0021] The RMSD values, RMSF values, hydrogen bond data, and binding free energy data of the characteristic conjugates and the comparative conjugates corresponding to all the advantageous substituent group structures corresponding to all the target positions are input into a pre-trained activity optimization model. The activity optimization model is trained using multiple sets of activity training data sets. Each set of the activity training data sets includes RMSD values, RMSF values, hydrogen bond data, binding free energy data, and the target position-substituent group structure activity relationship corresponding to the training data set. The target position-substituent group structure activity relationship result output by the activity optimization model is obtained, and the activity optimization parameters corresponding to the ligand compound are determined based on the target position-substituent group structure activity relationship result.

[0022] As an alternative implementation manner, in the second aspect of the present invention, the system further includes:

[0023] A ligand-derived module, configured to determine a plurality of optimized ligand-derived compounds based on the activity optimization parameters, where the optimized ligand-derived compounds include compounds obtained by setting corresponding dominant substituent group structures at the target positions of the ligand compound according to the indication of the activity optimization parameters;

[0024] A solvent analysis module, configured to obtain the proportion information of the receptor compound in the solvent, the type information of the solvent substance, and the proportion information of the solvent substance, where the solvent is composed of the receptor compound at a first proportion and the solvent substance at a second proportion;

[0025] A second molecular docking module, configured to input the molecular structure information of each optimized ligand-derived compound, the molecular structure information of the receptor compound, and the molecular structure information of each solvent substance into the molecular docking model for each optimized ligand-derived compound, obtain the second molecular docking result output by the molecular docking model, and determine the interaction relationship between the optimized ligand-derived compound and the receptor compound, and the interaction relationship between the optimized ligand-derived compound and each solvent substance based on the second molecular docking result;

[0026] An activity screening module, configured to determine the binding ability parameter corresponding to each optimized ligand-derived compound according to the proportion information of the receptor compound, the proportion information of the solvent substance, the interaction relationship between the optimized ligand-derived compound and the receptor compound, and the interaction relationship between the optimized ligand-derived compound and each solvent substance, where the binding ability parameter is used to measure the ability of the corresponding optimized ligand-derived compound to bind to the receptor compound in the solvent to form a complex; determine whether the binding ability parameter is greater than a preset parameter threshold, and if the binding ability parameter is greater than the preset parameter threshold, determine that the optimized ligand-derived compound is a high-activity ligand-derived compound.

[0027] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the parameter determination module obtains a plurality of ligand-derived compounds corresponding to the ligand compound includes:

[0028] Obtain a plurality of known ligand-derived compounds corresponding to the ligand compound, where the known ligand-derived compounds include derivatives of the existing ligand compound;

[0029] According to the known ligand-derived compounds and the substituent group position set, determine a plurality of potential substituent group structures corresponding to each target position of the ligand compound;

[0030] Generate a plurality of combined ligand-derived compounds according to all the potential substituent group structures corresponding to all the target positions, wherein at least one of the same target positions in any two of the combined ligand-derived compounds is provided with different potential substituent group structures;

[0031] Input the molecular structure information of all the combined ligand-derived compounds into a pre-trained verification model, which is trained using multiple sets of verification training data sets. Each set of verification training data sets includes the molecular structure information of a training ligand-derived compound with a preset training substituent group structure set at each target position, and the label data corresponding to the training ligand-derived compound, where the label data is used to indicate whether the training ligand-derived compound meets the property verification condition and the stability verification condition. Among them, the property verification condition is that the training ligand-derived compound can bind to the receptor compound to form a complex, and the stability verification condition is that the training ligand-derived compound can exist stably; obtain the result data output by the verification model, and determine the potential ligand-derived compounds that simultaneously meet the property verification condition and the stability verification condition from all the combined ligand-derived compounds according to the result data;

[0032] Determine a plurality of ligand-derived compounds corresponding to the ligand compound according to the known ligand-derived compound and the potential ligand-derived compound.

[0033] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the activity optimization module determines the dominant substituent group structure corresponding to each target position according to the substituent group activity parameters corresponding to all the ligand-derived compounds includes:

[0034] Determine the affinity index data corresponding to different target substituent group structures at each target position on the ligand compound according to the substituent group activity parameters corresponding to all the ligand-derived compounds;

[0035] For each target position on the ligand compound, determine all the dominant substituent group structures corresponding to the target position from different target substituent group structures according to the affinity index data corresponding to different target substituent group structures at the target position;

[0036] And, the affinity between the dominant substituent group structure and the receptor compound meets a preset activity condition, which specifically includes:

[0037] For each of the target positions on the ligand compound, the average value of the affinity index data corresponding to the preferred substituent group structure is greater than the average value of the affinity index data corresponding to all the target substituent group structures at this target position.

[0038] As an alternative embodiment, in the second aspect of the present invention, the system further includes:

[0039] An index determination module, configured to determine the affinity index data corresponding to each target substituent group structure on different ligand-derived compounds according to the substituent group activity parameters corresponding to all the ligand-derived compounds;

[0040] A ligand screening module, for each of the target substituent group structures, to determine the preferred ligand-derived compound and the inferior ligand-derived compound corresponding to this target substituent group structure according to the affinity index data corresponding to this target substituent group structure on all the ligand-derived compounds, wherein the average value of the affinity index data corresponding to this target substituent group structure on the preferred ligand-derived compound is greater than the average value of the affinity index data corresponding to this target substituent group structure on all the ligand-derived compounds; the average value of the affinity index data corresponding to this target substituent group structure on the inferior ligand-derived compound is less than or equal to the average value of the affinity index data corresponding to this target substituent group structure on all the ligand-derived compounds;

[0041] A combination analysis module, configured to determine the preferred substituent group combination according to all the target substituent group structures on the preferred ligand-derived compounds; and determine the inferior substituent group combination according to all the target substituent group structures on the inferior ligand-derived compounds;

[0042] A combination optimization module, configured to determine the activity combination optimization parameters corresponding to the ligand compound according to the preferred substituent group combination and the inferior substituent group combination, and the activity combination optimization parameters are used to indicate that corresponding target substituent group structures are respectively set at at least two target positions of the ligand compound.

[0043] As an alternative embodiment, in the second aspect of the present invention, the system further includes:

[0044] A common structure determination module is configured to, for each of the target positions on the ligand compound, determine the common structural features corresponding to all the dominant substituent group structures corresponding to the target position according to all the dominant substituent group structures corresponding to the target position, calculate the RMSF value of the common structure according to the molecular structure information of the common structure feature, and determine whether the RMSF value of the common structure is greater than a preset fluctuation threshold. If it is determined that the RMSF value of the common structure is greater than the preset fluctuation threshold, then determine that the common structure is a dominant common structure;

[0045] An optimization and update module is configured to update the activity optimization parameters corresponding to the ligand compound according to all the dominant common structures corresponding to each of the target positions, so that the activity optimization parameters are further used to indicate that a corresponding target dominant common structure is set at the target position of the ligand compound.

[0046] The third aspect of the present invention discloses another ligand activity improvement implementation system based on molecular structure analysis. The system includes:

[0047] A memory storing executable program code;

[0048] A processor coupled to the memory;

[0049] The processor calls the executable program code stored in the memory and executes the ligand activity improvement implementation method disclosed in the first aspect of the present invention.

[0050] The fourth aspect of the present invention discloses a computer storage medium storing computer instructions, which are used to execute the ligand activity improvement implementation method disclosed in the first aspect of the present invention when the computer instructions are called.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] Implementing the ligand activity improvement implementation method based on molecular structure analysis in the embodiments of the present invention can calculate the interaction relationship between each ligand-derived compound and the receptor compound, and then determine the affinity magnitudes corresponding to different substituent group structures at each target position according to the information of the substituent groups on the ligand compound, and further determine the dominant substituent groups, so as to optimize the molecular structure of the ligand compound with higher efficiency and reliability to improve the biological activity of the ligand compound. In addition, based on the molecular structure information of the receptor compound, the ligand compound and the ligand-derived compound, the present invention automatically analyzes the intermolecular interaction and matching relationship by using corresponding algorithms and models, and further calculates the activity optimization parameters for representing the molecular structure optimization strategy, reducing the dependence on the work experience of researchers and the participation of manual operations. Brief Description of the Drawings

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0054] Figure 1 is a schematic flowchart of a method for improving ligand activity based on molecular structure analysis disclosed in an embodiment of the present invention;

[0055] Figure 2 is a chemical structural formula diagram of a CDK6 inhibitor in an embodiment of the present invention;

[0056] Figure 3 is a chemical structural formula diagram of a CDK6 inhibitor derivative compound in an embodiment of the present invention;

[0057] Figure 4 is a schematic flowchart of another method for improving ligand activity based on molecular structure analysis disclosed in an embodiment of the present invention;

[0058] Figure 5 is a schematic structural diagram of a system for improving ligand activity based on molecular structure analysis disclosed in an embodiment of the present invention;

[0059] Figure 6 is a schematic structural diagram of another system for improving ligand activity based on molecular structure analysis disclosed in an embodiment of the present invention;

[0060] Figure 7 is a schematic structural diagram of yet another system for improving ligand activity based on molecular structure analysis disclosed in an embodiment of the present invention. Detailed Embodiments

[0061] To enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0062] References herein to "embodiments" mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0063] Embodiment 1

[0064] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for improving ligand activity based on molecular structure analysis disclosed in an embodiment of the present invention. Among them, Figure 1 the method for improving ligand activity based on molecular structure analysis described can be applied to a system for improving ligand activity based on molecular structure analysis. The system for improving ligand activity based on molecular structure analysis can be integrated in a cloud server or a local server, and the embodiments of the present invention do not make limitations. As Figure 1 shown, the method for improving ligand activity based on molecular structure analysis can include the following operations:

[0065] 101. Determine a set of substitution group positions of the ligand compound based on the molecular structure information of the ligand compound.

[0066] In the embodiments of the present invention, specifically, the set of substitution group positions includes at least one target position where the corresponding substitution group structure can be substituted. In the field of biochemistry, receptor compounds are generally substances on the surface of human cells, hormones, and various regulatory factors, etc. Ligand compounds are mainly targeted drugs used to match receptor compounds, such as inhibitors, lead drug molecules, etc. For example, the receptor compound is cyclin-dependent kinase 6 (CDK6), and the ligand compound is a Figure 2 shown CDK6 inhibitor: benzamide (4-substituted N-phenylpyrimidin-2-amine).

[0067] In the molecular structure of the above 4-substituted N-phenylpyrimidin-2-amine, R1, R2, R3, R4, and X respectively represent the target positions where the corresponding substitution group structures can be substituted. R1, R2, R3, R4, and X constitute the set of substitution group positions of the molecule of the 4-substituted N-phenylpyrimidin-2-amine derivative compound.

[0068] 102. Obtain a plurality of ligand-derived compounds corresponding to the ligand compound, and determine the substitution group parameters corresponding to the ligand-derived compounds according to the molecular structure information and the set of substitution group positions of the ligand-derived compounds.

[0069] In the embodiments of the present invention, the substitution group parameters are used to represent the target substitution group structures at each target position of the corresponding ligand-derived compound. A series of derivatives exist for the ligand compound on the premise of ensuring its properties and stability, which are called ligand-derived compounds in the embodiments of the present invention. For example, the 4-substituted N-phenylpyrimidin-2-amine-derived compounds may include 72 types shown in Table 1:

[0070] Table 1 Information related to 4-substituted N-phenylpyrimidin-2-amine-derived compounds

[0071]

[0072]

[0073]

[0074] In the embodiments of the present invention, the substitution group parameters corresponding to the ligand-derived compounds can be expressed as {NH-Cyclopentyl; CF3; H; NCOCH3; N}, indicating that the target substitution group structures at the target positions of R1, R2, R3, R4, and X in the 4-substituted N-phenylpyrimidin-2-amine-derived compounds are NH-Cyclopentyl, CF3, H, NCOCH3, and N respectively.

[0075] 103. Obtain the molecular structure information of the receptor compound. For each ligand-derived compound, input the molecular structure information of the ligand-derived compound and the molecular structure information of the receptor compound into a pre-trained molecular docking model to obtain a first molecular docking result output by the molecular docking model, and determine the interaction relationship between the ligand-derived compound and the receptor compound based on the first molecular docking result.

[0076] In the embodiments of the present invention, specifically, the molecular docking model analyzes the electric field distribution according to the molecular structure of the compound, and analyzes the intermolecular interactions from the perspectives of hydrogen bonds, van der Waals bonds, covalent bonds, and ionic bonds. The molecular docking model can be obtained through the following steps:

[0077] Construct a docking training dataset: multiple pairs of different molecular structures and the intermolecular interactions of each pair of molecular structures;

[0078] Train an initial molecular docking model using a docking training dataset to obtain a molecular docking model trained to convergence. This molecular docking model is used to output the interaction relationship between any molecular structures after inputting any molecular structure.

[0079] For example, input the molecular structure information of the 19th 4-substituted N-phenylpyrimidin-2-amine derivative compound in Table 1 and the molecular structure information of the receptor compound into the molecular docking model, then the interaction relationship between the 19th 4-substituted N-phenylpyrimidin-2-amine derivative compound and the receptor compound can be obtained. Among them, this interaction relationship can be reflected by the information of the chemical bonds formed at different positions between the molecular structures. Among them, the chemical structure of the 19th 4-substituted N-phenylpyrimidin-2-amine derivative compound is as Figure 3 shown.

[0080] 104. For each ligand derivative compound, determine the substitution group activity parameter corresponding to the ligand derivative compound according to the substitution group parameter corresponding to the ligand derivative compound and the interaction relationship between the ligand derivative compound and the receptor compound.

[0081] In the embodiments of the present invention, specifically, the substitution group activity parameter is used to represent the affinity between each target substitution group structure at each target position of the ligand derivative compound and the receptor compound. For the operation in step 104, the following is an example:

[0082] For the 4-substituted N-phenylpyrimidin-2-amine derivative compound No. 19 in Table 1, according to its substituent group parameters {NH-Cyclopentyl; CF3; H; NCOCH3; N}, and the interaction relationship between the 4-substituted N-phenylpyrimidin-2-amine derivative compound No. 19 and CDK6 obtained in Step 102, the substituent group activity parameters of the 4-substituted N-phenylpyrimidin-2-amine derivative compound No. 19 are determined to be: {12.345; 64.987; 3.764; 55.352}. Among them, the interaction relationship can be reflected by the information of the chemical bonds formed at different positions between the molecular structures. The substituent group activity parameters {12.345; 64.987; 3.764; 55.352} represent the affinity magnitudes between the target substituent group structures at the R1, R2, R3, R4, and X target positions of the 4-substituted N-phenylpyrimidin-2-amine derivative compound No. 19 and the receptor compound respectively.

[0083] 105. Determine the dominant substituent group structure corresponding to each target position according to the substituent group activity parameters corresponding to all ligand derivative compounds; determine the activity optimization parameters corresponding to the ligand compound according to the dominant substituent group structure corresponding to each target position.

[0084] In the embodiments of the present invention, specifically, the affinity between the dominant substituent group structure and the receptor compound satisfies a preset activity condition; the activity optimization parameter is used to indicate setting the corresponding target dominant substituent group structure at the target position of the ligand compound, so as to achieve the effect of improving the biological activity of the ligand compound.

[0085] In the embodiments of the present invention, optionally, among the substituent group activity parameters corresponding to a large number of obtained ligand derivative compounds, all the data can be sorted out, so as to represent the affinity magnitudes corresponding to different substituent group structures at different target positions on the ligand compound. Based on the sorted data described above, for each target position, a series of data for measuring the affinity magnitude can be obtained. The preset activity range condition can be that the affinity magnitude corresponding to the dominant substituent group structure at this target position exceeds the average value of the affinity magnitudes corresponding to all the substituent group structures at this target position. Optionally, the preset activity range condition can also be that the affinity magnitude corresponding to the dominant substituent group structure is the maximum value of the affinity magnitudes corresponding to all the substituent group structures at this target position.

[0086] It can be seen that by implementing the method for improving ligand activity based on molecular structure analysis in the embodiments of the present invention, the interaction relationship between each ligand-derived compound and the receptor compound can be calculated. Then, according to the information of the substituents on the ligand compound, the affinity magnitudes corresponding to different substituent structures at each target position can be determined, and the dominant substituents can be further determined, thereby optimizing the molecular structure of the ligand compound with higher efficiency and reliability to improve the biological activity of the ligand compound. In addition, based on the molecular structure information of the receptor compound, the ligand compound, and the ligand-derived compound, the present invention automatically analyzes the intermolecular interaction and matching relationship by using the corresponding algorithms and models, and further calculates the activity optimization parameters representing the molecular structure optimization strategy, reducing the dependence on the working experience of researchers and the participation of manual operations.

[0087] In an alternative embodiment, to determine the activity optimization parameters corresponding to the ligand compound according to the dominant substituent structure corresponding to each target position, it may specifically include:

[0088] For each dominant substituent structure corresponding to each target position, determine the characteristic ligand-derived compound and the comparative ligand-derived compound corresponding to the dominant substituent structure from multiple ligand-derived compounds;

[0089] For each dominant substituent structure corresponding to each target position, respectively combine the characteristic ligand-derived compound and the comparative ligand-derived compound corresponding to the dominant substituent structure with the receptor compound to obtain the characteristic complex and the comparative complex corresponding to the dominant substituent structure; according to the molecular structure information of the characteristic complex and the molecular structure information of the comparative complex corresponding to the dominant substituent structure, calculate the RMSD values of the characteristic complex and the comparative complex corresponding to the dominant substituent structure relative to the receptor compound respectively, calculate the RMSF values of the characteristic complex and the comparative complex corresponding to the dominant substituent structure respectively, calculate the hydrogen bond data of the characteristic complex and the comparative complex corresponding to the dominant substituent structure respectively; calculate the binding free energy data of the characteristic complex and the comparative complex corresponding to the dominant substituent structure respectively;

[0090] Input the RMSD values, RMSF values, hydrogen bond data, and binding free energy data of the characteristic complexes and the comparative complexes corresponding to all the dominant substituent structures at all target positions into a pre-trained activity optimization model; obtain the target position-substituent structure activity relationship result output by the activity optimization model, and determine the activity optimization parameters corresponding to the ligand compound based on the target position-substituent structure activity relationship result.

[0091] In this alternative embodiment, specifically, the characteristic ligand-derived compound is provided with the advantageous substituent group structure at the target position, while the comparative ligand-derived compound is not provided with the advantageous substituent group structure at the target position and has the same target substituent group structure as the characteristic ligand-derived compound at other target positions. Moreover, the hydrogen bond data includes the hydrogen bond acceptor, the hydrogen bond length, the hydrogen bond angle, and the hydrogen bond occurrence frequency, and the binding free energy data includes the polar binding energy data and the non-polar binding energy data. The activity optimization model is trained using multiple sets of activity training data sets. Each set of activity training data sets includes the RMSD value, the RMSF value, the hydrogen bond data, the binding free energy data, and the target position-substituent group structure activity relationship corresponding to the training data set. The activity optimization model is used to output the target position-substituent group structure activity relationship result after inputting the RMSD value, the RMSF value, the hydrogen bond data, and the binding free energy data.

[0092] In this alternative embodiment, the RMSD value (Root Mean Square Deviation) is a measure of the difference between two structures. In the fields of bioinformatics and molecular simulation, the RMSD is often used to compare the similarities between the three-dimensional structures of proteins, nucleic acids, or other molecules. The calculation of the RMSD value is based on the square root of the mean of the squares of the differences in the corresponding atomic coordinates in the two structures. The RMSF value (Root Mean Square Fluctuation) is an index that measures the amplitude of the position fluctuations of each amino acid residue in a protein during a molecular dynamics simulation. The calculation formula of the RMSF takes into account the simulation time period. By calculating the average position of each atom during the simulation time period and then calculating the mean square deviation of the position deviation of the atom from the average position during the simulation time period. In a molecular dynamics simulation, the RMSF is used to describe the fluctuations of various regions in a protein structure over time, thereby inferring the flexible and rigid regions of the protein. The hydrogen bond data and the free energy data can be calculated through molecular dynamics simulation, where molecular dynamics simulation is a comprehensive technology that combines physics, mathematics, and chemistry. Molecular dynamics mainly relies on Newtonian mechanics to simulate the motion of a molecular system, extract samples from a system composed of different states of the molecular system, calculate the configurational integral of the system, and further calculate the thermodynamic quantities and other macroscopic properties of the system based on the results of the configurational integral. It solves the equations of motion for a many-body system composed of atomic nuclei and electrons and is a computational method that can solve the dynamic problems of systems composed of a large number of atoms. It can not only directly simulate the macroscopic evolution characteristics of substances and obtain calculation results that are consistent with or close to the experimental results, but also provide a clear picture of the microscopic structure, particle motion, and their relationship with macroscopic properties.

[0093] In this alternative embodiment, the characteristic ligand-derived compound is provided with the advantageous substituent group structure at the target position, while the comparative ligand-derived compound is not provided with the advantageous substituent group structure at the target position and has the same target substituent group structure as the characteristic ligand-derived compound at other target positions, that is, a control group with controlled variables is adopted. For example, if the advantageous substituent group is NH-Cyclopentyl set at the R1 position, the 3rd 4-substituted N-phenylpyrimidin-2-amine-derived compound in Table 1 can be selected as the characteristic ligand-derived compound, and the 57th 4-substituted N-phenylpyrimidin-2-amine-derived compound as the comparative ligand-derived compound. Because only the substituent group structure at the R1 position is different between the 3rd 4-substituted N-phenylpyrimidin-2-amine-derived compound and the 57th 4-substituted N-phenylpyrimidin-2-amine-derived compound, while the substituent group structures at other positions are the same. In this way, in the presence of a control group, the activity optimization parameters representing the molecular structure optimization strategy can be calculated more accurately.

[0094] In this alternative embodiment, the target position-substituent group structure activity relationship represents the change relationship of the biological activity of the ligand-derived compound brought about by the change of the substituent group structure at any target position. Based on this target position-substituent group structure activity relationship, the target advantageous substituent group structure corresponding to the target position that can truly improve the biological activity of the ligand-derived compound can be screened out.

[0095] It can be seen that implementing the ligand activity improvement implementation method based on molecular structure analysis in this alternative embodiment can generate two sets of ligand-derived compounds for comparison according to the advantageous substituent group structure corresponding to each target position, and then calculate the RMSD value, RMSF value, hydrogen bond data, and binding free energy data corresponding to the two sets of ligand-derived compounds respectively. Finally, after analyzing the above data through the activity optimization model, the results of more reliable activity parameters are obtained.

[0096] In another alternative embodiment, the method may further include:

[0097] Determine several optimized ligand-derived compounds based on the activity optimization parameters;

[0098] Obtain the proportion information of the receptor compound in the solvent, the type information of the solvent substance, and the proportion information of the solvent substance;

[0099] For each optimized ligand-derived compound, input the molecular structure information of the optimized ligand-derived compound, the molecular structure information of the receptor compound, and the molecular structure information of each solvent substance into the molecular docking model to obtain the second molecular docking result output by the molecular docking model. Based on the second molecular docking result, determine the interaction relationship between the optimized ligand-derived compound and the receptor compound, and the interaction relationship between the optimized ligand-derived compound and each solvent substance;

[0100] For each optimized ligand-derived compound, according to the proportion information of the receptor compound, the proportion information of the solvent substance, the interaction relationship between the optimized ligand-derived compound and the receptor compound, and the interaction relationship between the optimized ligand-derived compound and each solvent substance, determine the binding ability parameter corresponding to the optimized ligand-derived compound; judge whether the binding ability parameter is greater than a preset parameter threshold. If the binding ability parameter is greater than the preset parameter threshold, determine that the optimized ligand-derived compound is a high-activity ligand-derived compound.

[0101] In this optional embodiment, specifically, the optimized ligand-derived compound includes a compound obtained by setting a corresponding dominant substituent group structure at the target position of the ligand compound according to the indication of the activity optimization parameter; and the solvent is composed of a first proportion of the receptor compound and a second proportion of the solvent substance, and the binding ability parameter is used to measure the ability of the corresponding optimized ligand-derived compound to bind to the receptor compound in the solvent to form a complex.

[0102] For ligand compounds, whether they are lead drug molecules or various inhibitors, their ultimate application scenarios are mainly in the human body. Therefore, the level of their biological activity ultimately needs to be proven in the human body. In this optional embodiment, optionally, the solvent is human body fluid, and the solvent substance is various substances in the body fluid that do not include the receptor compound.

[0103] It can be seen that implementing the method for improving ligand activity based on molecular structure analysis in this optional embodiment can simulate the human body environment through the solvent, analyze the binding ability of the optimized ligand-derived compound to the receptor compound in the simulated human body environment through the molecular docking model, and finally screen out high-activity ligand-derived compounds that can also exhibit high biological activity in the simulated human body environment, thereby further optimizing the molecular structure of the ligand compound with higher efficiency and reliability to maximize the biological activity of the ligand compound.

[0104] In another optional embodiment, obtaining multiple ligand-derived compounds corresponding to the ligand compound may specifically include:

[0105] Obtain multiple known ligand-derived compounds corresponding to the ligand compound;

[0106] According to the set of known ligand-derived compounds and the positions of substitution groups, determine the structures of multiple potential substitution groups corresponding to each target position of the ligand compound;

[0107] Generate multiple combined ligand-derived compounds based on the structures of all potential substitution groups corresponding to all target positions;

[0108] Input the molecular structure information of all combined ligand-derived compounds into a pre-trained verification model; obtain the result data output by the verification model, and determine the potential ligand-derived compounds that simultaneously meet the property verification condition and the stability verification condition from all combined ligand-derived compounds according to the result data;

[0109] Determine multiple ligand-derived compounds corresponding to the ligand compound according to the known ligand-derived compounds and the potential ligand-derived compounds.

[0110] In this alternative embodiment, specifically, the known ligand-derived compounds include derivatives of existing ligand compounds, and at least one of the same target positions in any two combined ligand-derived compounds is provided with different potential substitution group structures; the verification model is trained using multiple sets of verification training data sets. Each set of verification training data sets includes the molecular structure information of the training ligand-derived compounds with preset training substitution group structures set at each target position, and the label data corresponding to the training ligand-derived compounds. The label data is used to indicate whether the training ligand-derived compounds meet the property verification condition and the stability verification condition. Among them, the property verification condition is that the training ligand-derived compound can bind to the receptor compound to form a complex, and the stability verification condition is that the training ligand-derived compound can exist stably.

[0111] It can be seen that implementing the method for improving ligand activity based on molecular structure analysis in this alternative embodiment not only takes the known ligand-derived compounds as the research objects, but also can generate those potential ligand-derived compounds that theoretically exist based on the known ligand-derived compounds, thereby providing a more comprehensive and reasonable data source for determining the activity optimization parameters, and then optimizing the molecular structure of the ligand compound with higher efficiency and reliability to improve the biological activity of the ligand compound.

[0112] Embodiment Two

[0113] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of another method for improving ligand activity based on molecular structure analysis disclosed in the embodiments of the present invention. Among them, Figure 4The described method for realizing the improvement of ligand activity based on molecular structure analysis can be applied to a system for realizing the improvement of ligand activity based on molecular structure analysis. The system for realizing the improvement of ligand activity based on molecular structure analysis can be integrated in a cloud server or a local server, which is not limited in the embodiments of the present invention. As Figure 4 shown, the method for realizing the improvement of ligand activity based on molecular structure analysis may include the following operations:

[0114] 201. Based on the molecular structure information of the ligand compound, determine the set of positions of the substituent groups of the ligand compound.

[0115] 202. Obtain a plurality of ligand-derived compounds corresponding to the ligand compound, and determine the substituent group parameters corresponding to the ligand-derived compounds according to the molecular structure information of the ligand-derived compounds and the set of positions of the substituent groups.

[0116] 203. Obtain the molecular structure information of the receptor compound. For each ligand-derived compound, input the molecular structure information of the ligand-derived compound and the molecular structure information of the receptor compound into a pre-trained molecular docking model to obtain a first molecular docking result output by the molecular docking model, and determine the interaction relationship between the ligand-derived compound and the receptor compound based on the first molecular docking result.

[0117] 204. For each ligand-derived compound, determine the substituent group activity parameter corresponding to the ligand-derived compound according to the substituent group parameter corresponding to the ligand-derived compound and the interaction relationship between the ligand-derived compound and the receptor compound.

[0118] In the embodiments of the present invention, for other descriptions of steps 201-204, please refer to the detailed descriptions of steps 101-104 in Embodiment 1, and the embodiments of the present invention will not be elaborated herein.

[0119] 205. According to the substituent group activity parameters corresponding to all the ligand-derived compounds, determine the affinity index data corresponding to different target substituent group structures at each target position on the ligand compound.

[0120] 206. For each target position on the ligand compound, determine all the dominant substituent group structures corresponding to the target position from different target substituent group structures according to the affinity index data corresponding to different target substituent group structures at the target position.

[0121] In the embodiments of the present invention, specifically, the affinity between the dominant substituent group structure and the receptor compound meets a preset activity condition. Optionally, the affinity between the dominant substituent group structure and the receptor compound meets a preset activity condition, which may specifically include:

[0122] For each target position on the ligand compound, the average value of the affinity index data corresponding to the preferred substitution group structure is greater than the average value of the affinity index data corresponding to all target substitution group structures at this target position.

[0123] 207. Determine the activity optimization parameters corresponding to the ligand compound according to the preferred substitution group structure corresponding to each target position.

[0124] In the embodiments of the present invention, for other descriptions of step 207, please refer to the detailed description of step 105 in Embodiment 1, and the embodiments of the present invention will not be elaborated herein.

[0125] It can be seen that implementing the method for improving ligand activity based on molecular structure analysis in the embodiments of the present invention can determine all the preferred substitution group structures corresponding to a target position by judging whether the average value of the affinity index data corresponding to the preferred substitution group structure is greater than the average value of the affinity index data corresponding to all target substitution group structures at this target position, so as to more reasonably and accurately determine the preferred substitution group structure corresponding to each target position, and further provide more accurate data support for the determination of activity optimization parameters.

[0126] In an alternative embodiment, the method may further include:

[0127] Determine the affinity index data corresponding to each target substitution group structure on different ligand-derived compounds according to the substitution group activity parameters corresponding to all ligand-derived compounds;

[0128] For each target substitution group structure, determine the preferred ligand-derived compound and the inferior ligand-derived compound corresponding to this target substitution group structure according to the affinity index data corresponding to this target substitution group structure on all ligand-derived compounds;

[0129] Determine the preferred substitution group combination according to all target substitution group structures on the preferred ligand-derived compound; determine the inferior substitution group combination according to all target substitution group structures on the inferior ligand-derived compound;

[0130] Determine the activity combination optimization parameters corresponding to the ligand compound according to the preferred substitution group combination and the inferior substitution group combination.

[0131] In this alternative embodiment, specifically, the average value of the affinity index data corresponding to the target substitution group structure on the advantageous ligand-derived compound is greater than the average value of the affinity index data corresponding to the target substitution group structure on all ligand-derived compounds; the average value of the affinity index data corresponding to the target substitution group structure on the disadvantageous ligand-derived compound is less than or equal to the average value of the affinity index data corresponding to the target substitution group structure on all ligand-derived compounds; the activity combination optimization parameter is used to indicate that corresponding target substitution group structures are respectively set at at least two target positions of the ligand compound.

[0132] Since the change in the molecular structure of the ligand compound often has various restrictive effects, it is also necessary to examine the influence between different target substitution group structures set at different target positions. This influence may be a mutual promotion to further improve the biological activity of the ligand-derived compound, or may be a mutual cancellation that reduces the biological activity of the ligand-derived compound instead. Therefore, in this alternative embodiment, the advantageous substitution group combination on the advantageous ligand-derived compound indicates that setting corresponding substitution group structures at different target positions can mutually promote and further improve the biological activity of the ligand-derived compound, and the disadvantageous substitution group combination on the disadvantageous ligand-derived compound indicates that setting corresponding substitution group structures at different target positions can mutually cancel and reduce the biological activity of the ligand-derived compound instead. Finally, based on the advantageous substitution group combination and the disadvantageous substitution group combination, several optimal combination methods are determined.

[0133] It can be seen that implementing the method for improving ligand activity based on molecular structure analysis in this alternative embodiment also takes into account the influence between different advantageous substitution group structures after setting different advantageous substitution group structures at different target positions. Among them, the advantageous substitution group combination indicates that setting corresponding substitution group structures at different target positions can mutually promote and further improve the biological activity of the ligand-derived compound; the disadvantageous substitution group combination indicates that setting corresponding substitution group structures at different target positions can mutually cancel and reduce the biological activity of the ligand-derived compound instead. Based on the comparison of the advantageous substitution group combination and the disadvantageous substitution group combination, several optimal combination methods are determined, so as to provide a more comprehensive molecular structure optimization scheme for the ligand compound.

[0134] In another alternative embodiment, the method may further include:

[0135] For each target position on the ligand compound, based on all the preferred substituent group structures corresponding to the target position, determine the common structural features corresponding to all the preferred substituent group structures corresponding to the target position. According to the molecular structure information of the common structural features, calculate the RMSF value of the common structure, and determine whether the RMSF value of the common structure is greater than a preset fluctuation threshold. If it is determined that the RMSF value of the common structure is greater than the preset fluctuation threshold, then determine that the common structure is a preferred common structure;

[0136] Based on all the preferred common structures corresponding to each target position, update the activity optimization parameters corresponding to the ligand compound, so that the activity optimization parameters are also used to indicate that the corresponding target preferred common structure is set at the target position of the ligand compound.

[0137] For different preferred substituent group structures determined for the same target position, there are often some common structures or common properties among these preferred substituent group structures. For example, they are all electronegative groups, all medium-sized positively charged groups, all have benzene rings, etc. Therefore, by analyzing the above rules, deeper rules for improving the biological activity of ligand-derived compounds can be discovered. This alternative embodiment uses the RMSF value of the common structure as whether the common structure is a deeper rule for improving the biological activity of ligand-derived compounds, which can better verify whether the common structure meets the requirements.

[0138] It can be seen that in this alternative embodiment, based on different preferred substituent group structures determined for the same target position, a more general common structure can be determined. Through this common structure, deeper rules for improving the biological activity of ligand-derived compounds can be discovered, so as to optimize the molecular structure of the ligand compound with higher efficiency and reliability to improve the biological activity of the ligand compound.

[0139] Embodiment III

[0140] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a ligand activity improvement implementation system based on molecular structure analysis disclosed in an embodiment of the present invention. As Figure 5 shown, the ligand activity improvement implementation system based on molecular structure analysis may include:

[0141] A position determination module 301, configured to determine a set of substituent group positions of the ligand compound based on the molecular structure information of the ligand compound, and the set of substituent group positions includes at least one target position where the corresponding substituent group structure can be substituted;

[0142] A parameter determination module 302, configured to obtain a plurality of ligand-derived compounds corresponding to a ligand compound, and determine substitution group parameters corresponding to the ligand-derived compounds according to the molecular structure information and the substitution group position set of the ligand-derived compounds, where the substitution group parameters are used to represent the target substitution group structure at each target position of the corresponding ligand-derived compound;

[0143] A first molecular docking module 303, configured to obtain the molecular structure information of a receptor compound. For each ligand-derived compound, input the molecular structure information of the ligand-derived compound and the molecular structure information of the receptor compound into a pre-trained molecular docking model, obtain a first molecular docking result output by the molecular docking model, and determine the interaction relationship between the ligand-derived compound and the receptor compound based on the first molecular docking result;

[0144] An activity analysis module 304, configured to, for each ligand-derived compound, determine substitution group activity parameters corresponding to the ligand-derived compound according to the substitution group parameters corresponding to the ligand-derived compound and the interaction relationship between the ligand-derived compound and the receptor compound, where the substitution group activity parameters are used to represent the affinity magnitude between each target substitution group structure at each target position of the ligand-derived compound and the receptor compound;

[0145] An activity optimization module 305, configured to determine a dominant substitution group structure corresponding to each target position according to the substitution group activity parameters corresponding to all ligand-derived compounds, where the affinity between the dominant substitution group structure and the receptor compound meets a preset activity condition; and determine activity optimization parameters corresponding to the ligand compound according to the dominant substitution group structure corresponding to each target position, where the activity optimization parameters are used to indicate to set a corresponding target dominant substitution group structure at the target position of the ligand compound.

[0146] It can be seen that implementing Figure 5 the described system can calculate the interaction relationship between each ligand-derived compound and the receptor compound, then determine the affinity magnitude corresponding to different substitution group structures at each target position according to the information of the substitution groups on the ligand compound, and further determine the dominant substitution groups, so as to optimize the molecular structure of the ligand compound with higher efficiency and reliability to improve the biological activity of the ligand compound. In addition, based on the molecular structure information of the receptor compound, the ligand compound and the ligand-derived compounds, the present invention automatically analyzes the intermolecular interaction and matching relationship by using corresponding algorithms and models, and further calculates the activity optimization parameters used to represent the molecular structure optimization strategy, reducing the dependence on the work experience of researchers and the participation of manual operations.

[0147] In an alternative embodiment, the specific manner in which the activity optimization module 305 determines the activity optimization parameters corresponding to the ligand compound according to the dominant substitution group structure corresponding to each target position includes:

[0148] For each dominant substitution group structure corresponding to each target position, a characteristic ligand derivative compound and a comparative ligand derivative compound corresponding to the dominant substitution group structure are determined from a plurality of ligand derivative compounds. Among them, the characteristic ligand derivative compound is provided with the dominant substitution group structure at the target position, and the comparative ligand derivative compound is not provided with the dominant substitution group structure at the target position and has the same target substitution group structure as the characteristic ligand derivative compound at other target positions;

[0149] For each dominant substitution group structure corresponding to each target position, the characteristic ligand derivative compound and the comparative ligand derivative compound corresponding to the dominant substitution group structure are respectively combined with the receptor compound to obtain a characteristic conjugate and a comparative conjugate corresponding to the dominant substitution group structure; according to the molecular structure information of the characteristic conjugate and the comparative conjugate corresponding to the dominant substitution group structure, the RMSD values of the characteristic conjugate and the comparative conjugate corresponding to the dominant substitution group relative to the receptor compound are respectively calculated, the RMSF values of the characteristic conjugate and the comparative conjugate corresponding to the dominant substitution group are respectively calculated, and the hydrogen bond data of the characteristic conjugate and the comparative conjugate corresponding to the dominant substitution group are respectively calculated. The hydrogen bond data includes the acceptor of the hydrogen bond, the bond length of the hydrogen bond, the bond angle of the hydrogen bond, and the occurrence frequency of the hydrogen bond; the binding free energy data of the characteristic conjugate and the comparative conjugate corresponding to the dominant substitution group are respectively calculated, and the binding free energy data includes polar binding energy data and non-polar binding energy data;

[0150] The RMSD values, RMSF values, hydrogen bond data, and binding free energy data of the characteristic conjugates and comparative conjugates corresponding to all the dominant substitution group structures corresponding to all target positions are input into a pre-trained activity optimization model. The activity optimization model is trained using multiple sets of activity training data sets. Each set of activity training data sets includes RMSD values, RMSF values, hydrogen bond data, binding free energy data, and the target position-substitution group structure activity relationship corresponding to the training data set; the target position-substitution group structure activity relationship result output by the activity optimization model is obtained, and the activity optimization parameters corresponding to the ligand compound are determined based on the target position-substitution group structure activity relationship result.

[0151] It can be seen that by implementing this system, two sets of ligand-derived compounds can be generated for comparison based on the corresponding dominant substituent group structures at each target position, and then the RMSD values, RMSF values, hydrogen bond data, and binding free energy data of the compounds corresponding to the two sets of ligand-derived compounds can be calculated respectively. Finally, after analyzing the above data through the activity optimization model, results with higher reliability of activity parameters can be obtained.

[0152] In another alternative embodiment, as Figure 6 shown, the system may further include:

[0153] A ligand derivation module 306, configured to determine a number of optimized ligand-derived compounds based on activity optimization parameters. The optimized ligand-derived compounds include compounds obtained by setting corresponding dominant substituent group structures at the target positions of ligand compounds according to the indications of the activity optimization parameters;

[0154] A solvent analysis module 307, configured to obtain the proportion information of receptor compounds, the type information of solvent substances, and the proportion information of solvent substances in the solvent. The solvent consists of a first proportion of receptor compounds and a second proportion of solvent substances;

[0155] A second molecular docking module 308, configured to input the molecular structure information of each optimized ligand-derived compound, the molecular structure information of the receptor compound, and the molecular structure information of each solvent substance into the molecular docking model for each optimized ligand-derived compound, obtain the second molecular docking result output by the molecular docking model, and determine the interaction relationship between the optimized ligand-derived compound and the receptor compound, and the interaction relationship between the optimized ligand-derived compound and each solvent substance based on the second molecular docking result;

[0156] An activity screening module 309, configured to determine the binding ability parameter corresponding to each optimized ligand-derived compound according to the proportion information of receptor compounds, the proportion information of solvent substances, the interaction relationship between the optimized ligand-derived compound and the receptor compound, and the interaction relationship between the optimized ligand-derived compound and each solvent substance. The binding ability parameter is used to measure the binding ability of the corresponding optimized ligand-derived compound to combine with the receptor compound to form a complex in the solvent; determine whether the binding ability parameter is greater than a preset parameter threshold. If the binding ability parameter is greater than the preset parameter threshold, determine that the optimized ligand-derived compound is a high-activity ligand-derived compound.

[0157] It can be seen that implementing the system described in this alternative embodiment can simulate the human environment through a solvent, analyze the binding ability of the optimized ligand-derived compound with the receptor compound in the simulated human environment through a molecular docking model, and finally screen out highly active ligand-derived compounds that can also exhibit high biological activity in the simulated human environment, thereby further optimizing the molecular structure of the ligand compound with higher efficiency and reliability to maximize the biological activity of the ligand compound.

[0158] In another alternative embodiment, the specific manner in which the parameter determination module 302 obtains a plurality of ligand-derived compounds corresponding to a ligand compound includes:

[0159] Obtain a plurality of known ligand-derived compounds corresponding to the ligand compound, where the known ligand-derived compounds include derivatives of existing ligand compounds;

[0160] According to the known ligand-derived compounds and the set of substitution group positions, determine a plurality of potential substitution group structures corresponding to each target position of the ligand compound;

[0161] Generate a plurality of combined ligand-derived compounds according to all the potential substitution group structures corresponding to all the target positions, where at least one of the same target positions in any two combined ligand-derived compounds is provided with different potential substitution group structures;

[0162] Input the molecular structure information of all the combined ligand-derived compounds into a pre-trained verification model. The verification model is trained using multiple sets of verification training data sets. Each set of verification training data sets includes the molecular structure information of a training ligand-derived compound with a preset training substitution group structure set at each target position, and the label data corresponding to the training ligand-derived compound. The label data is used to indicate whether the training ligand-derived compound meets the property verification condition and the stability verification condition. The property verification condition is that the training ligand-derived compound can bind to the receptor compound to form a complex, and the stability verification condition is that the training ligand-derived compound can exist stably; obtain the result data output by the verification model, and determine the potential ligand-derived compounds that simultaneously meet the property verification condition and the stability verification condition from all the combined ligand-derived compounds according to the result data.

[0163] According to the known ligand-derived compounds and the potential ligand-derived compounds, determine a plurality of ligand-derived compounds corresponding to the ligand compound.

[0164] It can be seen that implementing the system described in this alternative embodiment not only takes known ligand-derived compounds as the objects of study, but also can generate those potentially existing ligand-derived compounds based on the known ligand-derived compounds, thereby providing a more comprehensive and reasonable data source for determining activity optimization parameters, and then optimizing the molecular structure of ligand compounds with higher efficiency and reliability to improve the biological activity of ligand compounds.

[0165] In yet another alternative embodiment, the specific manner in which the activity optimization module 305 determines the dominant substituent group structure corresponding to each target position according to the substituent group activity parameters corresponding to all ligand-derived compounds includes:

[0166] According to the substituent group activity parameters corresponding to all ligand-derived compounds, determine the affinity index data corresponding to different target substituent group structures at each target position on the ligand compound;

[0167] For each target position on the ligand compound, according to the affinity index data corresponding to different target substituent group structures at this target position, determine all the dominant substituent group structures corresponding to this target position from different target substituent group structures;

[0168] Moreover, the affinity between the dominant substituent group structure and the receptor compound satisfies a preset activity condition, which specifically includes:

[0169] For each target position on the ligand compound, the average value of the affinity index data corresponding to the dominant substituent group structure is greater than the average value of the affinity index data corresponding to all the target substituent group structures at this target position.

[0170] It can be seen that implementing the system described in this alternative embodiment can determine all the dominant substituent group structures corresponding to a target position by judging whether the average value of the affinity index data corresponding to the dominant substituent group structure is greater than the average value of the affinity index data corresponding to all the target substituent group structures at this target position, thereby more reasonably and accurately determining the dominant substituent group structure corresponding to each target position, and then providing more accurate data support for the determination of activity optimization parameters.

[0171] In yet another alternative embodiment, as Figure 6 shown, the system may further include:

[0172] An index determination module 310, configured to determine the affinity index data corresponding to each type of target substituent group structure on different ligand-derived compounds according to the substituent group activity parameters corresponding to all ligand-derived compounds;

[0173] A ligand screening module 311, which is used to determine, for each target substituent group structure, a preferred ligand-derived compound and a non-preferred ligand-derived compound according to the affinity index data corresponding to the target substituent group structure on all ligand-derived compounds, where the average value of the affinity index data corresponding to the target substituent group structure on the preferred ligand-derived compound is greater than the average value of the affinity index data corresponding to the target substituent group structure on all ligand-derived compounds; the average value of the affinity index data corresponding to the target substituent group structure on the non-preferred ligand-derived compound is less than or equal to the average value of the affinity index data corresponding to the target substituent group structure on all ligand-derived compounds.

[0174] A combination analysis module 312, which is used to determine a preferred substituent group combination according to all target substituent group structures on the preferred ligand-derived compounds; and determine a non-preferred substituent group combination according to all target substituent group structures on the non-preferred ligand-derived compounds.

[0175] A combination optimization module 313, which is used to determine an active combination optimization parameter corresponding to the ligand compound according to the preferred substituent group combination and the non-preferred substituent group combination, and the active combination optimization parameter is used to indicate that corresponding target substituent group structures are respectively set at at least two target positions of the ligand compound.

[0176] It can be seen that when implementing Figure 6 the described system, the influence between these preferred substituent group structures after setting different preferred substituent group structures at different target positions is also considered. Among them, the preferred substituent group combination indicates that setting corresponding substituent group structures at different target positions can promote each other to further improve the biological activity of the ligand-derived compound; the non-preferred substituent group combination indicates that setting corresponding substituent group structures at different target positions can offset each other and instead reduce the biological activity of the ligand-derived compound. Based on the comparison between the preferred substituent group combination and the non-preferred substituent group combination, several optimal combination methods are determined, so as to provide a more comprehensive molecular structure optimization scheme for the ligand compound.

[0177] In another alternative embodiment, as Figure 6 shown, the system may further include:

[0178] A common structure determination module 314 is configured to, for each target position on the ligand compound, determine the common structural features corresponding to all the dominant substituent group structures corresponding to the target position according to all the dominant substituent group structures corresponding to the target position, calculate the RMSF value of the common structure according to the molecular structure information of the common structural features, and determine whether the RMSF value of the common structure is greater than a preset fluctuation threshold. If it is determined that the RMSF value of the common structure is greater than the preset fluctuation threshold, then the common structure is determined as the dominant common structure;

[0179] An optimization and update module 315 is configured to update the activity optimization parameters corresponding to the ligand compound according to all the dominant common structures corresponding to each target position, so that the activity optimization parameters are further used to indicate that the corresponding target dominant common structure is set at the target position of the ligand compound.

[0180] It can be seen that implementing Figure 6 the described system can determine a more general common structure according to different dominant substituent group structures determined for the same target position. Through this common structure, deeper rules for improving the biological activity of ligand-derived compounds can be discovered, so as to optimize the molecular structure of the ligand compound with higher efficiency and reliability to improve the biological activity of the ligand compound.

[0181] Example 4

[0182] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of another ligand activity improvement implementation system based on molecular structure analysis disclosed in the embodiments of the present invention. As Figure 7 shown, the ligand activity improvement implementation system based on molecular structure analysis may include:

[0183] A memory 401 storing executable program code;

[0184] A processor 402 coupled to the memory 401;

[0185] The processor 402 calls the executable program code stored in the memory 401 and executes the steps in the ligand activity improvement implementation method based on molecular structure analysis described in Embodiment 1 or Embodiment 2 of the present invention.

[0186] Example 5

[0187] The embodiments of the present invention disclose a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute the steps in the ligand activity improvement implementation method based on molecular structure analysis described in Embodiment 1 or Embodiment 2 of the present invention.

[0188] Example 6

[0189] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps in the method for improving ligand activity based on molecular structure analysis described in Embodiment 1 or Embodiment 2.

[0190] The system embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0191] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, a magnetic disk memory, a tape memory, or any other computer-readable medium capable of carrying or storing data.

Claims

1. A method for improving ligand activity based on molecular structure analysis, characterized in that: The method comprises: Based on the molecular structure information of the ligand compound, determining a set of substitution group positions of the ligand compound, wherein the set of substitution group positions includes at least one target position where the corresponding substitution group structure can be substituted; Acquire multiple ligand derivative compounds corresponding to the ligand compound, and determine the substitution group parameters corresponding to the ligand derivative compound according to the molecular structure information of the ligand derivative compound and the substitution group position set, wherein the substitution group parameters are used to represent the target substitution group structure at each target position of the corresponding ligand derivative compound; Acquire the molecular structure information of the receptor compound, and for each of the ligand-derived compounds, input the molecular structure information of the ligand-derived compound and the molecular structure information of the receptor compound into a pre-trained molecular docking model to obtain a first molecular docking result output by the molecular docking model, and determine the interaction relationship between the ligand-derived compound and the receptor compound based on the first molecular docking result; For each of the ligand-derived compounds, according to the substituent group parameters corresponding to the ligand-derived compound and the interaction relationship between the ligand-derived compound and the receptor compound, determine the substituent group activity parameters corresponding to the ligand-derived compound, wherein the substituent group activity parameters are used to represent the affinity between each target substituent group structure at each target position of the ligand-derived compound and the receptor compound; Determine the dominant substituent group structure corresponding to each of the target positions according to the substituent group activity parameters corresponding to all the ligand-derived compounds, wherein the affinity between the dominant substituent group structure and the receptor compound satisfies the preset activity condition; For each of the dominant substituent group structures corresponding to each of the target positions, determine a characteristic ligand derivative compound and a comparative ligand derivative compound corresponding to the dominant substituent group structure from the plurality of ligand derivative compounds, wherein the characteristic ligand derivative compound is provided with the dominant substituent group structure at the target position, the comparative ligand derivative compound is not provided with the dominant substituent group structure at the target position, and the same target substituent group structure as the characteristic ligand derivative compound is provided at other target positions; For each of the dominant substituent group structures corresponding to each of the target positions, the characteristic ligand derivative compound and the comparative ligand derivative compound corresponding to the dominant substituent group structure are respectively combined with the receptor compound to obtain the characteristic binder and the comparative binder corresponding to the dominant substituent group structure; according to the molecular structure information of the characteristic binder and the molecular structure information of the comparative binder corresponding to the dominant substituent group structure, the RMSD values ​​of the characteristic binder and the comparative binder corresponding to the dominant substituent group relative to the receptor compound are respectively calculated, the RMSF values ​​of the characteristic binder and the comparative binder corresponding to the dominant substituent group are respectively calculated, and the hydrogen bond data of the characteristic binder and the comparative binder corresponding to the dominant substituent group are respectively calculated, wherein the hydrogen bond data include the hydrogen bond receptor, the bond length of the hydrogen bond, the bond angle of the hydrogen bond, and the frequency of occurrence of the hydrogen bond; the binding free energy data of the characteristic binder and the comparative binder corresponding to the dominant substituent group are respectively calculated, and the binding free energy data include polar binding energy data and non-polar binding energy data; Inputting the RMSD values, RMSF values, hydrogen bond data and binding free energy data of the characteristic binders and the comparative binders corresponding to all the advantageous substituent group structures corresponding to all the target positions into a pre-trained activity optimization model, wherein the activity optimization model is trained using multiple groups of activity training data sets, each group of the activity training data sets including RMSD values, RMSF values, hydrogen bond data, binding free energy data and the target position-substituent group structure activity relationship corresponding to the training data set; obtaining the target position-substituent group structure activity relationship result output by the activity optimization model, and determining the activity optimization parameters corresponding to the ligand compound based on the target position-substituent group structure activity relationship result, wherein the activity optimization parameters are used to indicate setting the corresponding target advantageous substituent group structure at the target position of the ligand compound; Based on the activity optimization parameters, several optimized ligand derivative compounds are determined, and the optimized ligand derivative compounds include compounds obtained by setting the corresponding dominant substitution group structure at the target position of the ligand compound according to the indication of the activity optimization parameters.

2. The method for improving ligand activity based on molecular structure analysis according to claim 1, characterized in that: The method further comprises: Acquiring information on the proportion of the receptor compound in the solvent, information on the type of solvent substance, and information on the proportion of the solvent substance, wherein the solvent is composed of a first proportion of the receptor compound and a second proportion of the solvent substance; For each of the optimized ligand-derived compounds, the molecular structure information of the optimized ligand-derived compound, the molecular structure information of the receptor compound and the molecular structure information of each of the solvent substances are input into the molecular docking model to obtain a second molecular docking result output by the molecular docking model, and based on the second molecular docking result, the interaction relationship between the optimized ligand-derived compound and the receptor compound, and the interaction relationship between the optimized ligand-derived compound and each of the solvent substances are determined; For each of the optimized ligand-derived compounds, the binding capacity parameter corresponding to the optimized ligand-derived compound is determined according to the proportion information of the receptor compound, the proportion information of the solvent substance, the interaction relationship between the optimized ligand-derived compound and the receptor compound, and the interaction relationship between the optimized ligand-derived compound and each of the solvent substances. The binding capacity parameter is used to measure the ability of the corresponding optimized ligand-derived compound to bind to the receptor compound in the solvent to produce a complex; it is judged whether the binding capacity parameter is greater than a preset parameter threshold value. If the binding capacity parameter is greater than the preset parameter threshold value, the optimized ligand-derived compound is determined to be a high-activity ligand-derived compound.

3. The method for improving ligand activity based on molecular structure analysis according to claim 2, characterized in that: The step of obtaining a plurality of ligand-derived compounds corresponding to the ligand compound comprises: Acquire a plurality of known ligand derivative compounds corresponding to the ligand compound, wherein the known ligand derivative compounds include derivatives of the existing ligand compound; Determining a plurality of potential substitution group structures corresponding to each target position of the ligand compound according to the known ligand derivative compound and the substitution group position set; According to all the potential substituent group structures corresponding to all the target positions, a plurality of combined ligand derivative compounds are generated, wherein at least one of any two combined ligand derivative compounds has a different potential substituent group structure arranged at the same target position; Inputting the molecular structure information of all the combined ligand derivative compounds into a pre-trained verification model, wherein the verification model is trained using multiple sets of verification training data sets, each set of the verification training data sets including the molecular structure information of the training ligand derivative compounds having a preset training substitution group structure at each target position, and label data corresponding to the training ligand derivative compounds, wherein the label data is used to indicate whether the training ligand derivative compounds meet the property verification conditions and the stability verification conditions, wherein the property verification conditions are that the training ligand derivative compounds can combine with the receptor compound to form a complex, and the stability verification conditions are that the training ligand derivative compounds can exist stably; obtaining the result data output by the verification model, and determining the potential ligand derivative compounds that meet both the property verification conditions and the stability verification conditions from all the combined ligand derivative compounds according to the result data; Based on the known ligand derivative compounds and the potential ligand derivative compounds, a plurality of ligand derivative compounds corresponding to the ligand compound are determined.

4. The method for improving ligand activity based on molecular structure analysis according to claim 1, characterized in that: The step of determining the dominant substituent group structure corresponding to each target position according to the substituent group activity parameters corresponding to all the ligand-derived compounds comprises: Determining affinity index data corresponding to different target substituent group structures at each target position on the ligand compound according to the substituent group activity parameters corresponding to all the ligand-derived compounds; For each of the target positions on the ligand compound, according to the affinity index data corresponding to the different target substituent group structures at the target position, all the dominant substituent group structures corresponding to the target position are determined from the different target substituent group structures; Furthermore, the affinity between the dominant substituent group structure and the receptor compound satisfies the preset activity conditions, specifically including: For each of the target positions on the ligand compound, the average value of the affinity index data corresponding to the dominant substitution group structure is greater than the average value of the affinity index data corresponding to all the target substitution group structures at the target position.

5. The method for improving ligand activity based on molecular structure analysis according to claim 4, characterized in that: The method further comprises: Determining affinity index data corresponding to each target substituent group structure on different ligand derivative compounds according to the substituent group activity parameters corresponding to all the ligand derivative compounds; For each of the target substituent group structures, the superior ligand derivative compounds and inferior ligand derivative compounds corresponding to the target substituent group structure are determined according to the affinity index data corresponding to the target substituent group structure on all the ligand derivative compounds, wherein the average value of the affinity index data corresponding to the target substituent group structure on the superior ligand derivative compounds is greater than the average value of the affinity index data corresponding to the target substituent group structure on all the ligand derivative compounds; the average value of the affinity index data corresponding to the target substituent group structure on the inferior ligand derivative compounds is less than or equal to the average value of the affinity index data corresponding to the target substituent group structure on all the ligand derivative compounds; Determine the combination of superior substituent groups according to the structures of all the target substituent groups on the superior ligand derivative compound; determine the combination of inferior substituent groups according to the structures of all the target substituent groups on the inferior ligand derivative compound; According to the advantageous substitution group combination and the inferior substitution group combination, the activity combination optimization parameters corresponding to the ligand compound are determined, and the activity combination optimization parameters are used to indicate that corresponding target substitution group structures are respectively set at at least two target positions of the ligand compound.

6. The method for improving ligand activity based on molecular structure analysis according to claim 4, characterized in that: The method further comprises: For each of the target positions on the ligand compound, according to all the dominant substituent group structures corresponding to the target position, determine the common structural features corresponding to all the dominant substituent group structures corresponding to the target position, calculate the RMSF value of the common structure according to the molecular structure information of the common structural features, and judge whether the RMSF value of the common structure is greater than a preset fluctuation threshold value, if it is judged that the RMSF value of the common structure is greater than the preset fluctuation threshold value, then determine that the common structure is a dominant common structure; According to all the dominant common structures corresponding to each of the target positions, the activity optimization parameters corresponding to the ligand compound are updated, so that the activity optimization parameters are also used to indicate setting the corresponding target dominant common structure at the target position of the ligand compound.

7. A system for improving ligand activity based on molecular structure analysis, characterized in that: The system comprises: A position determination module, for determining a substitution group position set of the ligand compound based on the molecular structure information of the ligand compound, wherein the substitution group position set includes at least one target position where the corresponding substitution group structure can be substituted; A parameter determination module, used for obtaining a plurality of ligand derivative compounds corresponding to the ligand compound, and determining the substitution group parameters corresponding to the ligand derivative compound according to the molecular structure information of the ligand derivative compound and the substitution group position set, wherein the substitution group parameters are used for representing the target substitution group structure at each target position of the corresponding ligand derivative compound; A first molecular docking module, used for obtaining molecular structure information of the receptor compound, for each of the ligand-derived compounds, inputting the molecular structure information of the ligand-derived compound and the molecular structure information of the receptor compound into a pre-trained molecular docking model, obtaining a first molecular docking result output by the molecular docking model, and determining the interaction relationship between the ligand-derived compound and the receptor compound based on the first molecular docking result; An activity analysis module, for determining, for each of the ligand-derived compounds, a substituent group activity parameter corresponding to the ligand-derived compound according to the substituent group parameter corresponding to the ligand-derived compound and the interaction relationship between the ligand-derived compound and the receptor compound, wherein the substituent group activity parameter is used to indicate the affinity between each target substituent group structure at each target position of the ligand-derived compound and the receptor compound; An activity optimization module, for determining the dominant substituent group structure corresponding to each of the target positions according to the substituent group activity parameters corresponding to all of the ligand-derived compounds, wherein the affinity between the dominant substituent group structure and the receptor compound satisfies a preset activity condition; and for determining the activity optimization parameters corresponding to the ligand compound according to the dominant substituent group structure corresponding to each of the target positions, wherein the activity optimization parameters are used to indicate setting the corresponding target dominant substituent group structure at the target position of the ligand compound; A ligand derivatization module, for determining a plurality of optimized ligand derivatization compounds based on the activity optimization parameters, wherein the optimized ligand derivatization compounds include compounds obtained by setting the corresponding dominant substitution group structure at the target position of the ligand compound according to the indication of the activity optimization parameters; The specific method of determining the activity optimization parameters corresponding to the ligand compound according to the dominant substitution group structure corresponding to each target position of the activity optimization module includes: For each of the dominant substituent group structures corresponding to each of the target positions, determine a characteristic ligand derivative compound and a comparative ligand derivative compound corresponding to the dominant substituent group structure from the plurality of ligand derivative compounds, wherein the characteristic ligand derivative compound is provided with the dominant substituent group structure at the target position, the comparative ligand derivative compound is not provided with the dominant substituent group structure at the target position, and the same target substituent group structure as the characteristic ligand derivative compound is provided at other target positions; For each of the dominant substituent group structures corresponding to each of the target positions, the characteristic ligand derivative compound and the comparative ligand derivative compound corresponding to the dominant substituent group structure are respectively combined with the receptor compound to obtain the characteristic binder and the comparative binder corresponding to the dominant substituent group structure; according to the molecular structure information of the characteristic binder and the molecular structure information of the comparative binder corresponding to the dominant substituent group structure, the RMSD values ​​of the characteristic binder and the comparative binder corresponding to the dominant substituent group relative to the receptor compound are respectively calculated, the RMSF values ​​of the characteristic binder and the comparative binder corresponding to the dominant substituent group are respectively calculated, and the hydrogen bond data of the characteristic binder and the comparative binder corresponding to the dominant substituent group are respectively calculated, wherein the hydrogen bond data include the hydrogen bond receptor, the bond length of the hydrogen bond, the bond angle of the hydrogen bond, and the frequency of occurrence of the hydrogen bond; the binding free energy data of the characteristic binder and the comparative binder corresponding to the dominant substituent group are respectively calculated, and the binding free energy data include polar binding energy data and non-polar binding energy data; The RMSD values, RMSF values, hydrogen bond data and binding free energy data of the characteristic binders and the comparative binders corresponding to all the advantageous substitution group structures corresponding to all the target positions are input into a pre-trained activity optimization model, wherein the activity optimization model is trained using multiple groups of activity training data sets, each group of the activity training data sets including RMSD values, RMSF values, hydrogen bond data, binding free energy data and the target position-substitution group structure activity relationship corresponding to the training data set; the target position-substitution group structure activity relationship result output by the activity optimization model is obtained, and the activity optimization parameters corresponding to the ligand compound are determined based on the target position-substitution group structure activity relationship result.

8. A system for improving ligand activity based on molecular structure analysis, characterized in that: The system includes: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the method for improving ligand activity based on molecular structure analysis as described in any one of claims 1-6.

9. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the method for improving ligand activity based on molecular structure analysis as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Organic ER affinity quick screening and forecast method based on receptor binding mode

    CN101059520A

  • Ligand molecule massive characteristic screening method in drug design

    CN106778032A