Molecular design method and system based on multi-parameter optimization

By adopting a multi-parameter optimization method in drug molecular design, combining multiple scoring functions and weight coefficients, the problem of fixing and high computational complexity of scoring functions in the existing technology is solved, and the comprehensiveness and flexibility of molecular design is improved, and the efficiency and success rate of drug discovery is improved.

CN120072109APending Publication Date: 2025-05-30BEIJING ANGOPRO TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510400803.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art has problems such as fixed scoring function, insufficient flexibility, high computational complexity and synthesisability in drug molecular design, and it is difficult to meet the complex and changeable drug design needs.

Method used

A molecular design method based on multi-parameter optimization is proposed. Multi-parameter optimization of molecules is achieved by constructing a molecular database, selecting multiple scoring functions (such as drug similarity, molecular weight, hydrogen bond donor and acceptor number, etc.) and weighted sum based on preset weight coefficients.

Benefits of technology

It has achieved comprehensive and flexible improvement in molecular design, improved the efficiency and success rate of drug discovery, solved the problems of fixing scoring functions and high computational complexity in the existing technology, and enhanced the synthesisability of molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072109A_ABST
    Figure CN120072109A_ABST
Patent Text Reader

Abstract

The invention discloses a molecule design method and system based on multi-parameter optimization, and the method comprises the steps: constructing a molecule database which comprises a plurality of candidate molecules; selecting a plurality of scoring functions, wherein the scoring functions comprise a drug similarity scoring function, a molecular weight scoring function, a hydrogen bond donor quantity scoring function and a hydrogen bond receptor quantity scoring function; for each candidate molecule, respectively calculating a score value of the candidate molecule under each scoring function; according to a preset weight coefficient, performing weighted summation on the score values of the scoring functions to obtain a comprehensive score of each candidate molecule; and sorting the candidate molecules according to the comprehensive scores, and selecting the candidate molecule with the highest comprehensive score as an optimized molecule. According to the method, the drug properties of molecules are optimized by combining multiple scoring functions; custom combination of various scoring functions is supported, an optimization strategy can be flexibly adjusted according to different drug design targets, and the comprehensiveness and flexibility of molecular design are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of new drug research and development, and particularly relates to a molecular design method and system based on multi-parameter optimization. Background Art

[0002] Traditional molecular design methods mainly rely on single or limited parameters for optimization, such as only considering single indicators such as drug likeness, molecular weight, or biological activity of molecules. In recent years, with the development of artificial intelligence technology, molecular design methods based on multi-parameter optimization have gradually emerged. For example, Patent CN116994673A proposes a drug design method based on two-stage evolutionary multi-task optimization. By determining the objective function of each sub-task and performing intra-task population evolution and optimization in the early and late stages of evolution respectively, molecules that meet multi-parameter requirements are finally screened out. In addition, there are also patents adopting drug molecule intelligent generation methods based on reinforcement learning and docking. Through the actor-critic reinforcement learning model and docking simulation method, combined with technologies such as knowledge distillation and conditional Transformer, the complex chemical space is efficiently traversed to find new compounds that meet various property constraints.

[0003] However, there are still many deficiencies in the existing technologies. On the one hand, as in Patent CN116994673A, its scoring function is fixed and cannot be custom-combined according to specific needs, which limits its application scope. On the other hand, although the method based on reinforcement learning and docking has improved flexibility, it still fails to achieve the free combination of multiple scoring functions, and has a high computational complexity, making it difficult to be efficiently applied to large-scale molecular screening. In addition, the existing methods also face challenges in the synthesizability of the generated molecules. For example, although some models based on combinatorial chemistry technology can generate molecules with high synthesizability, they often do not consider factors such as side reactions and reaction conditions, resulting in great difficulty in actual synthesis. These limitations make it difficult for the existing technologies to meet the complex and changeable drug design requirements. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a molecular design method and system based on multi-parameter optimization. By combining multiple scoring functions, such as drug likeness, molecular weight, hydrogen bond donors, and the number of receptors, the drug properties of molecules are optimized. Its core lies in supporting the custom combination of multiple scoring functions, and the optimization strategy can be flexibly adjusted according to different drug design goals. The main structural features include a scoring function library, a parameter optimization module, and a result output module, which achieve data interaction and processing through electrical connection. The present invention is mainly used to improve the design efficiency and screening accuracy of drug molecules to solve the problems existing in the above-mentioned existing technologies.

[0005] To achieve the above object, in the first aspect, the present invention provides a molecular design method based on multi-parameter optimization, including:

[0006] Construct a molecular database that contains multiple candidate molecules;

[0007] Select a variety of scoring functions, which include a drug similarity scoring function, a molecular weight scoring function, a hydrogen bond donor number scoring function, and a hydrogen bond acceptor number scoring function;

[0008] For each candidate molecule, calculate its score value under each scoring function respectively;

[0009] According to the preset weight coefficients, sum up the score values of each scoring function after weighting to obtain the comprehensive score of each candidate molecule;

[0010] Rank the candidate molecules according to the comprehensive score, and select the candidate molecule with the highest comprehensive score as the optimized molecule.

[0011] Preferably, the molecular database includes a known drug molecular database and a virtual molecular library generated by computer-aided design.

[0012] Preferably, according to the drug design goal, customize and combine the scoring functions.

[0013] Preferably, the drug similarity scoring function is used to evaluate the structural similarity between the candidate molecule and the known drug molecule;

[0014] The molecular weight scoring function is used to evaluate whether the molecular weight of the candidate molecule is within the preset range;

[0015] The hydrogen bond donor number scoring function is used to evaluate whether the number of hydrogen bond donors in the candidate molecule meets the preset conditions;

[0016] The hydrogen bond acceptor number scoring function is used to evaluate whether the number of hydrogen bond acceptors in the candidate molecule meets the preset conditions.

[0017] Preferably, the calculation formula for the comprehensive score is:

[0018]

[0019] Among them, S i is the comprehensive score of the i-th candidate molecule, w j is the weight coefficient of the j-th scoring function, f j (i) is the score value of the i-th candidate molecule under the j-th scoring function, and n is the number of scoring functions.

[0020] Preferably, it also includes protecting the molecular database and calculation results through data encryption technology.

[0021] Preferably, the multi-parameters include drug likeness, molecular weight, number of hydrogen bond donors and number of hydrogen bond acceptors.

[0022] In a second aspect, the present invention provides a molecular design system based on multi-parameter optimization, comprising:

[0023] A database construction module for constructing a molecular database, the molecular database containing a plurality of candidate molecules;

[0024] A scoring function selection module for selecting a variety of scoring functions, the scoring functions including a drug likeness scoring function, a molecular weight scoring function, a number of hydrogen bond donors scoring function and a number of hydrogen bond acceptors scoring function;

[0025] A parameter optimization module that calculates, for each candidate molecule, its score value under each scoring function; and according to a preset weight coefficient, sums the score values of each scoring function to obtain a comprehensive score for each candidate molecule;

[0026] A result output module for sorting the candidate molecules according to the comprehensive score and selecting the candidate molecule with the highest comprehensive score as the optimized molecule.

[0027] In a third aspect, the present invention also discloses a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0028] In a fourth aspect, the present invention also discloses a computer program product comprising a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0029] Compared with the prior art, the present invention has the following advantages and technical effects:

[0030] The present invention provides a molecular design method based on multi-parameter optimization. First, a molecular database is constructed, the molecular database containing a plurality of candidate molecules; second, a variety of scoring functions are selected, the scoring functions including a drug likeness scoring function, a molecular weight scoring function, a number of hydrogen bond donors scoring function and a number of hydrogen bond acceptors scoring function; then, for each candidate molecule, its score value under each scoring function is calculated respectively; further, according to a preset weight coefficient, the score values of each scoring function are summed up to obtain a comprehensive score for each candidate molecule; finally, the candidate molecules are sorted according to the comprehensive score, and the candidate molecule with the highest comprehensive score is selected as the optimized molecule.

[0031] The present invention optimizes the drug properties of molecules by combining multiple scoring functions, supports the custom combination of multiple scoring functions, can flexibly adjust the optimization strategy according to different drug design goals, and significantly improves the comprehensiveness and flexibility of molecular design. This method can not only comprehensively improve the drug properties of molecules, but also greatly improve the efficiency and success rate of drug discovery, bringing significant economic and social benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0033] Figure 1 It is the molecular design architecture diagram based on drug similarity and molecular weight according to the embodiment of the present invention;

[0034] Figure 2 It is the molecular design architecture diagram based on the number of hydrogen bond donors / acceptors according to the embodiment of the present invention;

[0035] Figure 3 It is the molecular design architecture diagram integrating multiple parameters according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0037] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0038] Embodiment 1

[0039] In this embodiment, a molecular design method based on multi-parameter optimization is provided, including:

[0040] S1. Construct a molecular database, and the molecular database contains multiple candidate molecules;

[0041] Construct a molecular database containing a large number of candidate molecules. The data sources of the molecular database include known drug molecule databases and virtual molecule libraries generated by computer-aided design.

[0042] The molecules adopt standard chemical structure representation methods, such as SMILES (Simplified Molecular Input Line Entry System) or molecular graphic representation.

[0043] S2. Select multiple scoring functions, which include a drug similarity scoring function, a molecular weight scoring function, a hydrogen bond donor number scoring function, and a hydrogen bond acceptor number scoring function;

[0044] Specifically, select multiple scoring functions, including but not limited to a drug similarity scoring function, a molecular weight scoring function, a hydrogen bond donor / acceptor number scoring function, etc. Users can customize the combination of these scoring functions according to specific drug design goals.

[0045] The drug similarity scoring function is used to evaluate the structural similarity between a candidate molecule and a known drug molecule. The principle of drug similarity is that molecules with similar structures often have similar biological activities.

[0046] The molecular weight scoring function is used to evaluate whether the molecular weight of a candidate molecule is within a preset range. The relationship between molecular weight and drug properties: Molecular weight directly affects the absorption, distribution, metabolism, and excretion of drugs.

[0047] The hydrogen bond donor number scoring function is used to evaluate whether the number of hydrogen bond donors in a candidate molecule meets the preset conditions.

[0048] The hydrogen bond acceptor number scoring function is used to evaluate whether the number of hydrogen bond acceptors in a candidate molecule meets the preset conditions. The number of hydrogen bond donors and acceptors affects the water solubility and biological activity of the molecule.

[0049] S3. For each candidate molecule, calculate its scoring value under each scoring function respectively;

[0050] As an innovative implementation method, use the selected scoring functions to score each candidate molecule in the molecular library, and optimize the molecules according to the scoring results. The optimization process includes but not limited to adjusting the molecular structure, replacing functional groups, etc. Use intelligent optimization algorithms such as genetic algorithms and particle swarm optimization algorithms to optimize the molecular structure.

[0051] S4. According to the preset weight coefficients, perform weighted summation on the scoring values of each scoring function to obtain the comprehensive score of each candidate molecule;

[0052] The weight coefficients can be customized and adjusted according to different drug design goals.

[0053] The calculation formula for the comprehensive score is:

[0054]

[0055] where S i is the comprehensive score of the i-th candidate molecule, w j is the weight coefficient of the j-th scoring function, and f j$s_{ij}$ is the scoring value of the $i$-th candidate molecule under the $j$-th scoring function, and $n$ is the number of scoring functions.

[0056] S5. Sort the candidate molecules according to the comprehensive score, and select the candidate molecule with the highest comprehensive score as the optimized molecule;

[0057] According to different drug design objectives, dynamically adjust the weights of each scoring function. For example, for drug design requiring high water solubility, the weight of the water solubility scoring function can be increased.

[0058] This embodiment provides a variety of optimization strategies for users to choose, such as global optimization, local optimization, etc.

[0059] According to the comprehensive score, screen the optimized molecules, and select the molecules with higher scores for further in vitro and in vivo experiments to confirm their drug properties.

[0060] The optimized molecule can verify its drug activity through further experiments.

[0061] S6. Use the optimized molecule for drug discovery.

[0062] The technical solution of this embodiment is applicable to the design of small molecule drugs and large molecule drugs.

[0063] This embodiment can be applied to the research and development of various drugs such as anti-tumor drugs, anti-infection drugs, and cardiovascular and cerebrovascular disease drugs.

[0064] This embodiment can be implemented by computer software, and the software includes a molecular database management module, a scoring function calculation module, a comprehensive score calculation module, and a result output module. The computer software can run on a general computer or a high-performance computing platform.

[0065] The method described in this embodiment can be used in combination with other drug design methods to further improve the efficiency and success rate of drug discovery.

[0066] The method described in this embodiment can achieve distributed computing through cloud computing technology to improve the computing efficiency and data processing ability.

[0067] This embodiment can optimize the scoring function and weight coefficient through machine learning algorithms to improve the accuracy and efficiency of molecular design.

[0068] This embodiment can display the structure and scoring results of candidate molecules through visualization technology to facilitate user analysis and selection.

[0069] This embodiment can realize the full-automatic operation from molecular screening to drug discovery through an automated process.

[0070] This embodiment can improve the computing speed through multi-threading technology to meet the processing requirements of large-scale molecular databases.

[0071] This embodiment can protect the security of molecular databases and calculation results through data encryption technology.

[0072] Example 1: The molecular design architecture diagram based on drug similarity and molecular weight is as Figure 1 shown.

[0073] 1. Input module:

[0074] The user inputs the structural information of the target drug molecule and the desired drug property parameters, such as the drug similarity threshold and the molecular weight range, through the input module.

[0075] 2. Scoring function selection module:

[0076] The user selects the required scoring functions in the scoring function selection module. In this embodiment, the drug similarity scoring function and the molecular weight scoring function are selected.

[0077] The drug similarity scoring function is used to evaluate the similarity between the generated molecule and the target drug molecule;

[0078] The molecular weight scoring function is used to evaluate whether the molecular weight of the generated molecule is within the preset range.

[0079] 3. Molecular generation module:

[0080] The molecular generation module generates a series of candidate molecules according to the input structural information of the target drug molecule. The generation method can adopt known molecular generation algorithms, such as deep learning models or rule-based generation methods.

[0081] 4. Optimization module:

[0082] The optimization module performs multi-parameter optimization on the generated candidate molecules. First, calculate the drug similarity score and the molecular weight score of each candidate molecule.

[0083] The drug similarity score is calculated by comparing the structural features (such as fingerprint maps) of the candidate molecule and the target drug molecule;

[0084] The molecular weight score is directly calculated according to the compliance degree of the molecular weight of the candidate molecule with the preset range.

[0085] The optimization module selects the candidate molecule with the highest comprehensive score by integrating the drug similarity score and the molecular weight score according to the set optimization strategy (such as the weighted average method).

[0086] 5. Output module:

[0087] The output module outputs the optimized candidate molecules and their scoring results to the user for further analysis and experimental verification.

[0088] Example 2: Architecture diagram of molecular design based on the number of hydrogen bond donors / acceptors, as Figure 2 shown.

[0089] 1. Input module:

[0090] The user inputs the structural information of the target drug molecule and the desired range of the number of hydrogen bond donors and acceptors through the input module.

[0091] 2. Scoring function selection module:

[0092] The user selects the hydrogen bond donor number scoring function and the hydrogen bond acceptor number scoring function in the scoring function selection module.

[0093] The hydrogen bond donor number scoring function is used to evaluate whether the number of hydrogen bond donors in the generated molecule is within the preset range;

[0094] The hydrogen bond acceptor number scoring function is used to evaluate whether the number of hydrogen bond acceptors in the generated molecule is within the preset range.

[0095] 3. Molecular generation module:

[0096] The molecular generation module generates a series of candidate molecules according to the input structural information of the target drug molecule. The generation method can adopt known molecular generation algorithms, such as deep learning models or rule-based generation methods.

[0097] 4. Optimization module:

[0098] The optimization module performs multi-parameter optimization on the generated candidate molecules. First, calculate the hydrogen bond donor number score and the hydrogen bond acceptor number score for each candidate molecule.

[0099] The hydrogen bond donor number score is calculated by counting the number of hydrogen bond donors in the candidate molecule and comparing it with the preset range, and the hydrogen bond acceptor number score is calculated by counting the number of hydrogen bond acceptors in the candidate molecule and comparing it with the preset range.

[0100] The optimization module selects the candidate molecule with the highest comprehensive score by integrating the hydrogen bond donor number score and the hydrogen bond acceptor number score according to the set optimization strategy (such as the weighted average method).

[0101] 5. Output module:

[0102] The output module outputs the optimized candidate molecules and their scoring results to the user for further analysis and experimental verification.

[0103] Example 3: Architecture diagram of molecular design integrating multiple parameters, as Figure 3As shown

[0104] 1. Input module:

[0105] The user inputs the structural information of the target drug molecule and various desired drug property parameters through the input module, such as drug similarity, molecular weight, number of hydrogen bond donors / acceptors, etc.

[0106] 2. Scoring function selection module:

[0107] The user selects various scoring functions in the scoring function selection module, including drug similarity scoring function, molecular weight scoring function, hydrogen bond donor number scoring function, and hydrogen bond acceptor number scoring function.

[0108] Each scoring function is used to evaluate the performance of the generated molecules on different drug property parameters.

[0109] 3. Molecular generation module:

[0110] The molecular generation module generates a series of candidate molecules based on the input structural information of the target drug molecule. The generation method can adopt known molecular generation algorithms, such as deep learning models or rule-based generation methods.

[0111] 4. Optimization module:

[0112] The optimization module performs multi-parameter optimization on the generated candidate molecules. First, calculate the scores of each candidate molecule on each scoring function.

[0113] The drug similarity score is calculated by comparing the structural features of the candidate molecule with the target drug molecule. The molecular weight score is calculated directly based on the degree of compliance of the molecular weight of the candidate molecule with the preset range. The hydrogen bond donor number score and the hydrogen bond acceptor number score are calculated by counting the number of hydrogen bond donors and acceptors in the candidate molecule and comparing them with the preset range respectively.

[0114] The optimization module selects the candidate molecule with the highest comprehensive score by integrating the scores of each scoring function according to the set optimization strategy (such as weighted average method or multi-objective optimization algorithm).

[0115] 5. Output module:

[0116] The output module outputs the optimized candidate molecule and its scoring results to the user for further analysis and experimental verification.

[0117] Through the detailed description of the above three examples, the multi-parameter optimization molecular design method proposed in this embodiment can flexibly combine various scoring functions, optimize according to different drug design goals, thereby effectively improving the design efficiency and success rate of drug molecules. The method of this embodiment has broad application prospects in the field of drug discovery.

[0118] Example Two

[0119] Based on the same inventive concept, this embodiment also provides a molecular design system based on multi-parameter optimization, including:

[0120] A database construction module for constructing a molecular database, which contains multiple candidate molecules;

[0121] A scoring function selection module for selecting multiple scoring functions, including a drug similarity scoring function, a molecular weight scoring function, a hydrogen bond donor number scoring function, and a hydrogen bond acceptor number scoring function;

[0122] A parameter optimization module that calculates the scoring values of each candidate molecule under each scoring function respectively; according to the preset weight coefficients, the scoring values of each scoring function are weighted and summed to obtain the comprehensive score of each candidate molecule;

[0123] A result output module for sorting the candidate molecules according to the comprehensive score and selecting the candidate molecule with the highest comprehensive score as the optimized molecule.

[0124] The molecular design system based on multi-parameter optimization provided in this embodiment has all the advantages of the molecular design method based on multi-parameter optimization provided in Example One.

[0125] Example Three

[0126] This embodiment also discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in Example One are implemented.

[0127] Example Four

[0128] This embodiment also discloses a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in Example One are implemented.

[0129] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A molecular design method based on multi-parameter optimization, characterized in that: The following steps are involved: constructing a molecular database, wherein the molecular database comprises a plurality of candidate molecules; selecting a plurality of scoring functions, wherein the scoring functions include a drug similarity scoring function, a molecular weight scoring function, a hydrogen bond donor quantity scoring function, and a hydrogen bond acceptor quantity scoring function; For each candidate molecule, calculate its score under each scoring function; According to the preset weight coefficient, the score values ​​of each scoring function are weighted and summed to obtain the comprehensive score of each candidate molecule; The candidate molecules were ranked according to the comprehensive scores, and the candidate molecules with the highest comprehensive scores were selected as the optimized molecules.

2. The method according to claim 1, characterized in that: The molecular database includes a known drug molecule database and a virtual molecule library generated by computer-aided design.

3. The method according to claim 1, characterized in that The scoring functions are custom combined according to the drug design goals.

4. The method according to claim 1, characterized in that The drug similarity scoring function is used to evaluate the structural similarity between candidate molecules and known drug molecules; The molecular weight scoring function is used to evaluate whether the molecular weight of the candidate molecule is within a preset range; The hydrogen bond donor quantity scoring function is used to evaluate whether the quantity of hydrogen bond donors in the candidate molecule meets the preset conditions; The hydrogen bond acceptor quantity scoring function is used to evaluate whether the quantity of hydrogen bond acceptors in the candidate molecule meets a preset condition.

5. The method according to claim 1, characterized in that The calculation formula of the comprehensive score is: Among them, S i is the comprehensive score of the i-th candidate molecule, w j is the weight coefficient of the jth scoring function, f j (i) is the score of the i-th candidate molecule under the j-th scoring function, and n is the number of scoring functions.

6. The method according to claim 1, characterized in that It also includes protecting molecular databases and calculation results through data encryption technology.

7. The method according to claim 1, characterized in that The multiple parameters include drug similarity, molecular weight, and the number of hydrogen bond donors and acceptors.

8. A molecular design system based on multi-parameter optimization, characterized in that: include: A database construction module, used to construct a molecular database, wherein the molecular database contains a plurality of candidate molecules; A scoring function selection module, used to select a plurality of scoring functions, wherein the scoring functions include a drug similarity scoring function, a molecular weight scoring function, a hydrogen bond donor quantity scoring function, and a hydrogen bond acceptor quantity scoring function; The parameter optimization module calculates the score of each candidate molecule under each scoring function; according to the preset weight coefficient, the score of each scoring function is weighted and summed to obtain the comprehensive score of each candidate molecule; The result output module is used to sort the candidate molecules according to the comprehensive scores and select the candidate molecules with the highest comprehensive scores as the optimized molecules.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • AI-assisted drug delivery molecule design method and system

    CN121366665A