Molecular descriptor generation method, device, equipment, medium and product

By using multiple preset conformation generation algorithms in the conformation algorithm sequence, and trying to successfully generate molecular conformation data sequentially until the molecular conformation data is successfully generated, the problem of poor applicability of a single algorithm is solved, and the generation efficiency and success rate of molecular descriptors are improved.

CN120340658APending Publication Date: 2025-07-18SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510408266.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

A single conformational generation algorithm is difficult to apply to the conformational space of different molecules, resulting in low molecular conformational generation rate and descriptor generation number and success rate.

Method used

The target conformational algorithm is used to use a preset conformation generation algorithm that contains at least two sequential arrangements in the conformational algorithm sequence, and the target conformational algorithm is attempted in sequence until successful, and molecular conformational data is generated and molecular descriptor data is finally generated.

Benefits of technology

The generation rate of molecular conformation and the generation number and success rate of molecular descriptors are improved, the complexity of molecular conformation space is adapted to the complexity of molecular conformation space, and the generation dimension of molecular descriptors is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340658A_ABST
    Figure CN120340658A_ABST
Patent Text Reader

Abstract

The invention discloses a molecular descriptor generation method, device and equipment, a medium and a product. The method comprises the following steps: acquiring a chemical representation object and a conformation algorithm sequence of a molecule; taking a first preset conformation generation algorithm in the conformation algorithm sequence as a target conformation generation algorithm, and performing conformation generation according to the chemical representation object by adopting the target conformation generation algorithm to obtain molecular conformation data; if the generation of the target conformation generation algorithm fails, taking the next preset conformation generation algorithm in the conformation algorithm sequence as the target conformation generation algorithm, and returning to execute the step of adopting the target conformation generation algorithm and performing conformation generation according to the chemical representation object to obtain molecular conformation data; and when the target conformation generation algorithm is successfully generated, the molecular descriptor data of the molecules are generated according to the molecular conformation data, so that the generation rate of the molecular conformation is improved, and the generation quantity and success rate of the molecular descriptors are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chemoinformatics, and particularly to a method, apparatus, device, medium and product for generating molecular descriptors. Background Art

[0002] A molecular descriptor is a parameter information that quantitatively expresses information such as the structure and properties of a molecule, and plays a key promoting role in fields such as chemical property prediction, structure-activity relationship exploration, and molecular structure design.

[0003] Molecular conformation is the basis for generating molecular descriptors, representing different arrangements of a molecule in three-dimensional space generated by different rotational states of rotatable chemical bonds in the molecule. As the number of atoms in the molecule increases, the number of molecular conformations grows exponentially.

[0004] Therefore, the conformational spaces of different molecules have great differences, and a single conformational generation algorithm is difficult to apply to the conformational spaces of different molecules, greatly reducing the generation rate of molecular conformations, and thus affecting the generation quantity and success rate of molecular descriptors. Summary of the Invention

[0005] Embodiments of the present invention provide a method, apparatus, device, medium and product for generating molecular descriptors to solve the problem that a single conformational generation algorithm has poor applicability to the conformational space of molecules, improve the generation rate of molecular conformations, and thus improve the generation quantity and success rate of molecular descriptors.

[0006] According to an embodiment of the present invention, a method for generating a molecular descriptor is provided, the method comprising:

[0007] Obtaining a chemical representation object of a molecule and a conformational algorithm sequence; wherein, the conformational algorithm sequence includes at least two preset conformational generation algorithms arranged in sequence;

[0008] Taking the first preset conformational generation algorithm in the conformational algorithm sequence as a target conformational generation algorithm, and using the target conformational generation algorithm to generate conformational data of the molecule according to the chemical representation object;

[0009] If the target conformational generation algorithm fails to generate, taking the next preset conformational generation algorithm in the conformational algorithm sequence as the target conformational generation algorithm, and returning to execute the step of using the target conformational generation algorithm to generate conformational data of the molecule according to the chemical representation object;

[0010] Until the target conformational generation algorithm generates successfully, generating molecular descriptor data of the molecule according to the conformational data.

[0011] According to another embodiment of the present invention, there is provided a device for generating molecular descriptors, the device comprising:

[0012] A conformational algorithm sequence acquisition module for acquiring a chemical representation object of a molecule and a conformational algorithm sequence; wherein, the conformational algorithm sequence includes at least two preset conformational generation algorithms arranged in sequence;

[0013] A molecular conformation data determination module for using the first preset conformational generation algorithm in the conformational algorithm sequence as a target conformational generation algorithm, and using the target conformational generation algorithm to generate molecular conformation data according to the chemical representation object;

[0014] A return execution module for, if the target conformational generation algorithm fails to generate, using the next preset conformational generation algorithm in the conformational algorithm sequence as the target conformational generation algorithm, and returning to execute the step of generating molecular conformation data according to the chemical representation object using the target conformational generation algorithm;

[0015] A molecular descriptor data generation module for generating molecular descriptor data of the molecule according to the molecular conformation data until the target conformational generation algorithm generates successfully.

[0016] According to another embodiment of the present invention, there is provided an electronic device, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for generating molecular descriptors according to any embodiment of the present invention.

[0020] According to another embodiment of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the method for generating molecular descriptors according to any embodiment of the present invention when executed.

[0021] According to another embodiment of the present invention, there is provided a computer program product including a computer program which implements the method for generating molecular descriptors according to any embodiment of the present invention when executed by a processor.

[0022] In the technical solution of this embodiment, by setting that the conformational algorithm sequence of the molecule includes at least two preset conformational generation algorithms arranged in sequence, repeatedly taking the preset conformational generation algorithms in the conformational algorithm sequence of the molecule as the target conformational generation algorithms in turn, and using the target conformational generation algorithms to generate molecular conformational data according to the chemical representation object of the molecule. Until the molecular conformational data is not empty, according to the molecular conformational data, generate the molecular descriptor data of the molecule. Considering the complexity of the conformational space of the molecule, it solves the problem that a single conformational generation algorithm has poor applicability to the conformational space of the molecule, improves the generation rate of molecular conformations, and improves the generation quantity and success rate of molecular descriptors from the generation dimension of molecular conformations.

[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0025] Figure 1 It is a flowchart of a method for generating a molecular descriptor provided by an embodiment of the present invention;

[0026] Figure 2 It is a flowchart of another method for generating a molecular descriptor provided by an embodiment of the present invention;

[0027] Figure 3 It is a flowchart of a specific example of a method for generating a molecular descriptor provided by an embodiment of the present invention;

[0028] Figure 4 It is a comparison diagram of the generation results of a molecular descriptor provided by an embodiment of the present invention;

[0029] Figure 5 It is a schematic structural diagram of a device for generating a molecular descriptor provided by an embodiment of the present invention;

[0030] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0032] It should be noted that the terms "initial", "target", "preset", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0033] Figure 1 The flowchart of a method for generating a molecular descriptor provided by an embodiment of the present invention is applicable to the situation of performing feature quantification analysis on molecules. This method can be executed by a molecular descriptor generation device, which can be implemented in the form of hardware and / or software, and the molecular descriptor generation device can be configured in a terminal device. As Figure 1 shown, the method includes:

[0034] S110, obtain the chemical representation object and the conformational algorithm sequence of the molecule.

[0035] Exemplarily, the type to which the molecule belongs can be an organic small molecule, a biological macromolecule, a metal-organic compound, or a polymer. Among them, an organic small molecule refers to an organic compound composed of several to dozens of atoms, with a relative molecular mass usually below 1000, such as hydrocarbons, alcohols, aldehydes, ketones, carboxylic acids, esters, or amines, etc. A biological macromolecule refers to a macromolecule present in the cells of an organism, formed by connecting many repeating monomers through specific chemical bonds, with a relatively large relative molecular mass, ranging from several thousand to millions or even higher, such as proteins, nucleic acids, or polysaccharides, etc. A metal-organic compound is a type of compound containing a metal-carbon bond (M-C bond), where the metal atom is combined with an organic group through a chemical bond, such as metal carbonyl compounds or metal alkyl compounds, etc. A polymer is a macromolecular compound formed by connecting many repeating structural units through covalent bonds, and these repeating structural units are interconnected through a polymerization reaction to form a polymer chain, such as polyethylene, polypropylene, or polystyrene, etc.

[0036] The type to which the molecule belongs is not limited here, and it can be specifically customized according to actual needs.

[0037] Specifically, the molecular representation object is used to represent the spatial structure information of the molecule. In an optional embodiment, the molecular representation object is a SMILES (Simplified Molecular Input Line Entry System) object or a mol object (Molecular object).

[0038] Among them, a SMILES object is a linear format, using a string composed of ASCII (American Standard Code for Information Interchange) characters to represent the molecular structure, that is, through specific characters and rules, information such as atoms, chemical bonds, and the topological structure of the molecule in the molecule is encoded into a string. For example, the SMILES object of ethanol is cco, where c represents a carbon atom and o represents an oxygen atom.

[0039] Among them, a mol object is a structured format, usually containing rich three-dimensional structure information such as the coordinates of atoms in the molecule, atom types, chemical bond types, bond lengths, and bond angles.

[0040] In a specific embodiment, when the type to which the molecule belongs is an organic small molecule, perform mol format conversion processing on the SMILES object of the molecule to obtain a format conversion result; according to the format conversion result, determine the number of chemical bonds of non-hydrogen atoms; according to the valence state and the number of chemical bonds of non-hydrogen atoms, determine the amount of hydrogen atom supplementation for non-hydrogen atoms; according to the amount of hydrogen atom supplementation, update the format conversion result to obtain the mol object of the molecule.

[0041] In this embodiment, the conformational algorithm sequence includes at least two preset conformational generation algorithms arranged in sequence. Specifically, the conformational generation algorithm aims to systematically explore various molecular conformations that a molecule may have through certain mathematical models and calculation methods to find molecular conformations with lower energy or that meet specific conditions. Among them, the arrangement order of the preset conformational generation algorithms in the conformational algorithm sequence represents the execution order of the preset conformational generation algorithms.

[0042] In an alternative embodiment, each preset conformational generation algorithm in the conformational algorithm sequence represents a conformational generation algorithm adapted to a type of molecule, and the types of molecules corresponding to at least two preset conformational generation algorithms are different.

[0043] Exemplarily, the preset conformational generation algorithm corresponding to small organic molecules is the MMFF (Merck Molecular Force Field) algorithm, the preset conformational generation algorithm corresponding to biological macromolecules is the molecular dynamics simulation algorithm, and the preset conformational generation algorithm corresponding to polymers is the coarse-grained molecular dynamics algorithm, but it is not limited to the exemplary situation.

[0044] In another alternative embodiment, at least two preset conformational generation algorithms in the conformational algorithm sequence respectively include force field optimization algorithms, and at least two force field optimization algorithms are arranged in descending order according to the force field accuracy level corresponding to the type to which the molecule belongs.

[0045] Among them, the force field is a potential energy function that describes the interaction between atoms in a molecule. The force field optimization algorithm calculates the force on the molecule in the current molecular conformation, such as the direction and magnitude of the force, and adjusts the positions of the atoms according to the force situation to move the molecule in the direction of energy reduction and finally reach the conformation with the minimum energy.

[0046] Exemplarily, the force field optimization algorithm can be the AMBER (Assisted Model Building with Energy Refinement) algorithm, the CHARMM (Chemistry at HARvard Macromolecular Mechanics) algorithm, the MMFF algorithm, or the UFF (Universal Force Field) algorithm, etc.

[0047] In the above example, when the type of the molecule is an organic small molecule, the force field accuracy level of the MMFF algorithm or the AMBER algorithm is higher than that of the CHARMM algorithm, and the force field accuracy level of the CHARMM algorithm is higher than that of the UFF algorithm. When the type of the molecule is a biomacromolecule, the force field accuracy level of the CHARMM algorithm or the AMBER algorithm is higher than that of the MMFF algorithm, and the force field accuracy level of the MMFF algorithm is higher than that of the UFF algorithm. When the type of the molecule is a metal-organic compound, the force field accuracy level of the UFF algorithm is higher than that of the AMBER algorithm, the CHARMM algorithm or the MMFF algorithm.

[0048] In another alternative embodiment, the first preset conformation generation algorithm in the conformation algorithm sequence is a knowledge-based conformation generation algorithm or a sampling-based conformation generation algorithm. At least one preset conformation generation algorithm other than the first preset conformation generation algorithm respectively includes a force field optimization algorithm. When the number of force field optimization algorithms is at least two, the at least two force field optimization algorithms are arranged in descending order according to the force field accuracy level corresponding to the type of the molecule.

[0049] Among them, the knowledge-based conformation generation algorithm aims to generate molecular conformations by using existing knowledge and experience. The specific knowledge can come from experimental data, theoretical calculation results, and the understanding of molecular structures and properties. By analyzing and summarizing the chemical representation objects of a large number of known molecular conformations, a relevant knowledge model is established. Then, according to the chemical representation objects of other molecules, information is obtained from the knowledge model to predict or generate the molecular conformations of other molecules. Exemplarily, the knowledge-based conformation generation algorithm can be the ETKDG (Extended-Tongue Kinetically Driven Growth) algorithm, the GROMOS (Groningen Molecular Simulation) algorithm, or the SCHNAPS (Stochastic Chemical Network Algorithm for Protein Structure) algorithm, etc.

[0050] Among them, the sampling-based conformation generation algorithm aims to perform random or regular sampling in the conformation space of the molecule to find the molecular conformation. Exemplarily, the sampling-based conformation generation algorithm is the Monte Carlo algorithm, the simulated annealing algorithm, the genetic algorithm, or the molecular dynamics simulation algorithm, etc.

[0051] No limitation is imposed on the first preset conformation generation algorithm here, and it can be specifically customized according to actual requirements.

[0052] Based on the above embodiments, optionally, the preset conformation generation algorithms other than the first preset conformation generation algorithm further include a random coordinate algorithm. Specifically, the preset conformation generation algorithms other than the first preset conformation generation algorithm represent a combination form of the random coordinate algorithm and the force field optimization algorithm, where the random coordinate algorithm aims to randomly change the three-dimensional coordinate values of some or all of the atoms in the molecule to generate one or more atomic coordinate data.

[0053] The independent force field optimization algorithm is restricted by the initial molecular conformation and is likely to fall into a local optimal solution and fail to find the global optimal conformation or other important low-energy conformations, or fail to converge. In this embodiment, by adopting the random coordinate algorithm, the possibility of the force field optimization algorithm finding the global optimal conformation or other important low-energy conformations is increased, thereby improving the accuracy of the molecular conformation data and reducing the dependence of multiple force field optimization algorithms on a single molecular conformation, thus improving the generation rate of the molecular conformation data.

[0054] S120. Take the first preset conformation generation algorithm in the conformation algorithm sequence as the target conformation generation algorithm.

[0055] S130. Adopt the target conformation generation algorithm to generate molecular conformation data according to the chemical representation object.

[0056] In an alternative embodiment, adopting the target conformation generation algorithm to generate molecular conformation data according to the chemical representation object includes: when the target conformation generation algorithm is a knowledge-based conformation generation algorithm, a sampling-based conformation generation algorithm, or a force field optimization algorithm, taking the chemical representation object as the input parameter of the target conformation generation algorithm to obtain the molecular conformation data.

[0057] In another alternative embodiment, adopting the target conformation generation algorithm to generate molecular conformation data according to the chemical representation object includes: when the target conformation generation algorithm includes a random coordinate algorithm and a force field optimization algorithm, adopting the random coordinate algorithm to generate atomic coordinate data according to the chemical representation object, and adopting the force field optimization algorithm to perform force field optimization on the atomic coordinate data to obtain the molecular conformation data.

[0058] S140. Determine whether the target conformation generation algorithm fails to generate. If so, execute S150; if not, execute S160.

[0059] In an alternative embodiment, specifically, if the performance metrics of the target conformation generation algorithm meet the performance anomaly condition, or the molecular conformation data meets the data anomaly condition, then set the target conformation generation algorithm to generate unsuccessfully; if the performance metrics of the target conformation generation algorithm do not meet the performance anomaly condition, and the molecular conformation data does not meet the data anomaly condition, then set the target conformation generation algorithm to generate successfully.

[0060] Specifically, the performance anomaly condition is used to describe a situation where the target conformation generation algorithm exhibits performance inconsistent with normal performance. Exemplarily, the performance anomaly condition includes, but is not limited to, at least one of the target conformation generation algorithm being unable to converge, the number of iterations exceeding the number threshold, and the running duration exceeding the duration threshold, etc.

[0061] Specifically, the data anomaly condition is used to describe a situation where the molecular conformation data does not conform to the expected normal situation. Exemplarily, the data anomaly condition includes, but is not limited to, at least one of the molecular conformation data being empty, the molecular energy corresponding to the molecular conformation being greater than the energy threshold, atomic overlap existing in the molecular conformation, the bond length not satisfying the bond length range, the bond angle not satisfying the bond angle range, and the molecular conformation not satisfying the stereochemical rules, etc.

[0062] S150. Take the next preset conformation generation algorithm in the conformation algorithm sequence as the target conformation generation algorithm, and execute S130.

[0063] In an alternative embodiment, before taking the next preset conformation generation algorithm in the conformation algorithm sequence as the target conformation generation algorithm, the method further includes: if the target conformation generation algorithm is the last preset conformation generation algorithm in the conformation algorithm sequence, then end the generation process of the molecular descriptor.

[0064] S160. Generate the molecular descriptor data of the molecule according to the molecular conformation data.

[0065] In an alternative embodiment, generating the molecular descriptor data of the molecule according to the molecular conformation data includes: using the molecular conformation data as the input parameter of the preset descriptor generation algorithm to obtain the molecular descriptor data of the molecule.

[0066] Exemplarily, the preset descriptor generation algorithm can be the Omx (Orthogonal Matching Pursuit witheXtended search) algorithm, the AIQM1 (Artificial Intelligence-Quantum Mechanics 1) algorithm, the ANI-2x (Artificial Neural Network-2x) algorithm, or the GFN2xTB algorithm, etc., but is not limited to the example cases.

[0067] Exemplarily, when the type of the molecule is an organic small molecule, the molecular descriptor data includes, but is not limited to, volume, surface area, dipole moment, charge distribution, charge separation degree, and orbital energy, etc. When the type of the molecule is a biological macromolecule, the molecular descriptor data includes, but is not limited to, amino acid composition, amino acid sequence, secondary structure content, enzyme active site, and binding site, etc. When the type of the molecule is a metal-organic compound, the molecular descriptor data includes, but is not limited to, the radius of the metal ion, the steric hindrance of the ligand, the covalent and ionic degrees of the chemical bond, etc. When the type of the molecule is a polymer, the molecular descriptor data includes, but is not limited to, number-average molecular weight, weight-average molecular weight, repeating unit structure, chain configuration, degree of branching, and crystallinity, etc.

[0068] Specifically, molecular descriptors are used in application scenarios such as molecular property prediction, chemical reaction effect, or molecular classification. Taking the property prediction scenario as an example, the number of generated molecular descriptors and the success rate can effectively improve the accuracy and execution efficiency of molecular property prediction.

[0069] In the technical solution of this embodiment, by setting that at least two sequentially arranged preset conformation generation algorithms are included in the conformation algorithm sequence of the molecule, and repeatedly using the preset conformation generation algorithms in the conformation algorithm sequence of the molecule as the target conformation generation algorithms in turn, and adopting the target conformation generation algorithms to generate molecular conformation data according to the chemical representation object of the molecule, until the molecular conformation data is not empty, and then generating molecular descriptor data of the molecule according to the molecular conformation data. This solution takes into account the complexity of the conformational space of the molecule, solves the problem of poor applicability of a single conformation generation algorithm to the conformational space of the molecule, improves the generation rate of molecular conformations, and increases the number of generated molecular descriptors and the success rate from the dimension of generating molecular conformations.

[0070] Figure 2The flowchart of another method for generating molecular descriptors provided by an embodiment of the present invention further refines the step of "generating molecular descriptor data of a molecule according to molecular conformation data" in the above embodiment. In this embodiment, generating molecular descriptor data of a molecule according to molecular conformation data includes: obtaining descriptor identification data and a descriptor algorithm sequence corresponding to the molecule; taking the first preset descriptor generation algorithm in the descriptor algorithm sequence as the target descriptor generation algorithm, and using the target descriptor generation algorithm to generate initial descriptor data according to the molecular conformation data; if the descriptor identification data is not included in the initial descriptor data, taking the next preset descriptor generation algorithm in the descriptor algorithm sequence as the target descriptor generation algorithm, and returning to execute the step of using the target descriptor generation algorithm to generate initial descriptor data according to the molecular conformation data; until the descriptor identification data is included in the initial descriptor data, determining the molecular descriptor data of the molecule according to the descriptor identification data and the initial descriptor data. As Figure 2 shown, the method includes:

[0071] S210. Obtain the chemical representation object and the conformation algorithm sequence of the molecule.

[0072] S220. Take the first preset conformation generation algorithm in the conformation algorithm sequence as the target conformation generation algorithm.

[0073] S230. Use the target conformation generation algorithm to generate molecular conformation data according to the chemical representation object.

[0074] S240. Determine whether the target conformation generation algorithm fails to generate. If so, execute S250; if not, execute S260.

[0075] S250. Take the next preset conformation generation algorithm in the conformation algorithm sequence as the target conformation generation algorithm, and execute S230.

[0076] S210 - S250 in this embodiment correspond to and are the same as or similar to S110 - S150 Figure 1 shown above, and will not be elaborated herein.

[0077] S260. Obtain the descriptor identification data and the descriptor algorithm sequence corresponding to the molecule.

[0078] Specifically, the descriptor identification data contains at least one preset descriptor identification for uniquely identifying the molecular descriptor.

[0079] In this embodiment, the descriptor algorithm sequence includes at least two preset descriptor generation algorithms arranged in sequence. Specifically, the preset descriptor generation algorithm represents a method of converting molecular structure information into quantified numerical features, i.e., molecular descriptors. Among them, the arrangement order of the preset descriptor generation algorithms in the descriptor algorithm sequence represents the execution order of the preset descriptor generation algorithms.

[0080] In an alternative embodiment, at least two preset descriptor generation algorithms in the descriptor algorithm sequence are arranged in descending order of algorithm accuracy.

[0081] Exemplarily, the preset descriptor generation algorithm can be the GFN2xTB algorithm, the GFN1xTB algorithm, the GFN0xTB algorithm, the density functional tight binding algorithm, the PM7 (Parameterized Model 7) algorithm, the OM3 algorithm, the OM2 algorithm, the OM1 algorithm, the MNDO (Modified Neglect of Diatomic Overlap) algorithm, or the ANI-2x algorithm.

[0082] In the above examples, the algorithm accuracies of the GFN2xTB algorithm, the GFN1xTB algorithm, and the GFN0xTB algorithm decrease in sequence, the algorithm accuracies of the OM3 algorithm, the OM2 algorithm, and the OM1 algorithm decrease in sequence, the algorithm accuracy of the PM7 algorithm is less than that of the GFN1xTB algorithm and greater than that of the GFN0xTB algorithm, and the algorithm accuracy of the MNDO algorithm is less than that of the OM2 algorithm.

[0083] In another alternative embodiment, each preset descriptor generation algorithm in the descriptor algorithm sequence represents a conformation generation algorithm adapted to a type of molecule, and the molecular types corresponding to at least two preset descriptor generation algorithms are different.

[0084] Exemplarily, the preset descriptor generation algorithm corresponding to organic small molecules is the GFN2xTB algorithm or the ANI-2x algorithm, the preset descriptor generation algorithm corresponding to biological macromolecules is the density functional tight binding algorithm, and the preset descriptor generation algorithm corresponding to polymers is the GFN0xTB algorithm, but it is not limited to the example situation.

[0085] S270. Use the first preset descriptor generation algorithm in the descriptor algorithm sequence as the target descriptor generation algorithm.

[0086] S280. Adopt the target descriptor generation algorithm to generate descriptors based on the molecular conformation data to obtain initial descriptor data.

[0087] Specifically, the molecular conformation data is used as the input parameter of the target descriptor generation algorithm to obtain the initial descriptor data. The initial descriptor data represents the calculation result of the target descriptor generation algorithm, and the initial descriptor data contains at least one descriptor identifier and the parameter value of each descriptor identifier.

[0088] S290. Determine whether the descriptor identifier data is included in the initial descriptor data. If so, execute S292; if not, execute S291.

[0089] Specifically, if each descriptor identifier in the descriptor identifier data is included in the initial descriptor data, it is set that the descriptor identifier data is included in the initial descriptor data; if at least one descriptor identifier in the descriptor identifier data is not included in the initial descriptor data, it is set that the descriptor identifier data is not included in the initial descriptor data.

[0090] Exemplarily, assume that the descriptor identifier data includes molecular descriptor 1, molecular descriptor 2, and molecular descriptor 3, and the initial descriptor data includes molecular descriptor 1 and its parameter value, molecular descriptor 3 and its parameter value, and molecular descriptor 4 and its parameter value. Since molecular descriptor 2 is not included in the initial descriptor data, the descriptor identifier data is not included in the initial descriptor data.

[0091] S291. Use the next preset descriptor generation algorithm in the descriptor algorithm sequence as the target descriptor generation algorithm, and execute S280.

[0092] In a specific embodiment, before using the next preset descriptor generation algorithm in the descriptor algorithm sequence as the target descriptor generation algorithm, the method further includes: if the target descriptor generation algorithm is the last preset descriptor generation algorithm in the descriptor algorithm sequence, end the generation process of the molecular descriptor.

[0093] S292. Determine the molecular descriptor data of the molecule according to the descriptor identifier data and the initial descriptor data.

[0094] Specifically, the molecular descriptor data contains the parameter value corresponding to each descriptor identifier in the descriptor identifier data.

[0095] Figure 3 This is a flowchart of a specific example of a method for generating a molecular descriptor provided by an embodiment of the present invention. Figure 3Taking the conformational algorithm sequence including the ETKDGv3 algorithm, the random coordinate + MMFF algorithm, and the random coordinate + UFF algorithm, and the descriptor algorithm sequence including the GFN2xTB algorithm, the GFN1xTB algorithm, and the GFN0xTB algorithm as examples, where the ETKDGv3 algorithm represents the 3rd version of the ETKDG algorithm. Specifically, convert the SMILES object of the molecule into a mol object, and use the ETKDGv3 algorithm to generate molecular conformation data based on the mol object. If the generation by the ETKDGv3 algorithm fails, then continue to use the random coordinate + MMFF algorithm to generate molecular conformation data based on the mol object. If the generation by the MMFF algorithm fails, then finally use the random coordinate + UFF algorithm to generate molecular conformation data based on the mol object. If the generation by the UFF algorithm fails, then end the generation process of the molecular descriptor.

[0096] If the generation by the ETKDGv3 algorithm, the MMFF algorithm, or the UFF algorithm is successful, then use the GFN2xTB algorithm to generate initial descriptor data based on the molecular conformation data, and screen the initial descriptor data according to the descriptor identification data. If the screening is successful, obtain the molecular descriptor data. If the screening fails, then continue to use the GFN1xTB algorithm to perform geometric optimization on the molecular conformation data, and generate initial descriptor data based on the optimized molecular conformation data. Screen the initial descriptor data according to the descriptor identification data. If the screening is successful, obtain the molecular descriptor data. If the continuous screening fails, then finally use the GFN0xTB algorithm to continue performing geometric optimization on the molecular conformation data, and generate initial descriptor data based on the optimized molecular conformation data. Screen the initial descriptor data according to the descriptor identification data. If the screening is successful, obtain the molecular descriptor data. If the final screening fails, then end the generation process of the molecular descriptor.

[0097] Figure 4 This is a comparison graph of the generation results of a molecular descriptor provided by an embodiment of the present invention. Specifically, Figure 4 the left figure in [ ] represents the generation ratio graph A of the molecular descriptor obtained by using a single conformational generation algorithm and a single descriptor generation algorithm, Figure 4 the right figure in [ ] represents the generation ratio graph B of the molecular descriptor obtained by using the molecular descriptor generation method provided by the embodiment of the present invention.

[0098] From Figure 4 it can be seen that in the prior art, the successful generation ratio of the molecular descriptor is 53%, while the successful generation ratio of the molecular descriptor using the embodiment of the present invention is 81%, which is significantly higher than the prior art level.

[0099] Based on the above embodiments, optionally, the method further includes: obtaining a trained property prediction model; inputting the chemical representation object and the molecular descriptor data into the property prediction model to obtain the molecular property attributes of the output molecule.

[0100] Specifically, the property prediction model represents a deep learning model for predicting the molecular property attributes of molecules. Exemplarily, the property prediction model can be a deep multitask neural network, a convolutional neural network, a recurrent neural network, a graph neural network, etc., but is not limited to the exemplary situation.

[0101] Exemplarily, when the type of the molecule is an organic small molecule, the molecular property attributes include but are not limited to melting point, boiling point, density, solubility, refractive index, optical band gap, and spectral properties, etc. When the type of the molecule is a biological macromolecule, the molecular property attributes include but are not limited to sedimentation coefficient, electrophoretic mobility, viscosity, hydrolyzability, and enzyme activity. When the type of the molecule is a metal-organic compound, the molecular property attributes include but are not limited to catalytic activity, nuclear magnetic resonance signal, and optical band gap, etc. When the type of the molecule is a polymer, the molecular property attributes include but are not limited to melting point, tensile strength, optical band gap, thermal stability, and corrosion resistance, etc.

[0102] The advantage of such a setting is that the accuracy of the molecular property attributes is improved.

[0103] The technical solution of this embodiment attempts to adopt a multi-level descriptor generation algorithm, generates initial descriptor data according to the molecular conformation data, and screens the initial descriptor data according to the descriptor identification data to determine whether the screening is successful. If the screening fails, continue to try the next descriptor generation algorithm. If the screening is successful, use the screening result as the molecular descriptor data of the molecule, which solves the problem that the calculation accuracy setting of a single descriptor generation algorithm is unreasonable, and further improves the generation quantity and success rate of the molecular descriptor from the calculation dimension of the molecular descriptor.

[0104] The following is an embodiment of the molecular descriptor generation device provided by the embodiments of the present invention. The device and the molecular descriptor generation method of the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiment of the molecular descriptor generation device, reference can be made to the content of the above embodiment regarding the molecular descriptor generation method.

[0105] Figure 5 It is a schematic structural diagram of a molecular descriptor generation device provided by an embodiment of the present invention. As Figure 5 shown, the device includes: a conformation algorithm sequence acquisition module 310, a molecular conformation data determination module 320, a return execution module 330, and a molecular descriptor data generation module 340.

[0106] Among them, the conformational algorithm sequence acquisition module 310 is used to acquire the chemical representation object of the molecule and the conformational algorithm sequence; wherein, the conformational algorithm sequence includes at least two preset conformational generation algorithms arranged in sequence;

[0107] The molecular conformation data determination module 320 is used to take the first preset conformational generation algorithm in the conformational algorithm sequence as the target conformational generation algorithm, and use the target conformational generation algorithm to generate molecular conformation data according to the chemical representation object;

[0108] The return execution module 330 is used to, if the target conformational generation algorithm fails to generate, take the next preset conformational generation algorithm in the conformational algorithm sequence as the target conformational generation algorithm, and return to execute the step of using the target conformational generation algorithm to generate molecular conformation data according to the chemical representation object;

[0109] The molecular descriptor data generation module 340 is used to, until the target conformational generation algorithm generates successfully, generate the molecular descriptor data of the molecule according to the molecular conformation data.

[0110] The technical solution of this embodiment, by setting that the conformational algorithm sequence of the molecule includes at least two preset conformational generation algorithms arranged in sequence, repeatedly taking the preset conformational generation algorithms in the conformational algorithm sequence of the molecule as the target conformational generation algorithms in turn, using the target conformational generation algorithm to generate molecular conformation data according to the chemical representation object of the molecule, until the molecular conformation data is not empty, generating the molecular descriptor data of the molecule according to the molecular conformation data, considers the complexity of the conformational space of the molecule, solves the problem that a single conformational generation algorithm has poor applicability to the conformational space of the molecule, improves the generation rate of molecular conformations, and improves the generation quantity and success rate of molecular descriptors from the generation dimension of molecular conformations.

[0111] In an alternative embodiment, the first preset conformational generation algorithm in the conformational algorithm sequence is a conformational generation algorithm based on specific knowledge or a conformational generation algorithm based on sampling, and at least one preset conformational generation algorithm other than the first preset conformational generation algorithm respectively includes a force field optimization algorithm;

[0112] When the number of force field optimization algorithms is at least two, the at least two force field optimization algorithms are arranged in descending order according to the force field accuracy level corresponding to the type of the molecule.

[0113] In an alternative embodiment, the preset conformational generation algorithms other than the first preset conformational generation algorithm further include a random coordinate algorithm; the molecular conformation data determination module 320 is specifically used for:

[0114] When the target conformation generation algorithm includes a random coordinate algorithm and a force field optimization algorithm, the random coordinate algorithm is adopted to generate atomic coordinate data according to the chemical representation object, and the force field optimization algorithm is adopted to perform force field optimization on the atomic coordinate data to obtain molecular conformation data.

[0115] In an alternative embodiment, the molecular descriptor data generation module 340 is specifically configured to:

[0116] Obtain the descriptor identification data and the descriptor algorithm sequence corresponding to the molecule; wherein, the descriptor algorithm sequence includes at least two preset descriptor generation algorithms arranged in sequence;

[0117] Take the first preset descriptor generation algorithm in the descriptor algorithm sequence as the target descriptor generation algorithm, and use the target descriptor generation algorithm to generate initial descriptor data according to the molecular conformation data;

[0118] If the descriptor identification data is not included in the initial descriptor data, take the next preset descriptor generation algorithm in the descriptor algorithm sequence as the target descriptor generation algorithm, and return to execute the step of using the target descriptor generation algorithm to generate initial descriptor data according to the molecular conformation data;

[0119] Until the descriptor identification data is included in the initial descriptor data, determine the molecular descriptor data of the molecule according to the descriptor identification data and the initial descriptor data.

[0120] In an alternative embodiment, at least two preset descriptor generation algorithms in the descriptor algorithm sequence are arranged in descending order of algorithm accuracy.

[0121] In an alternative embodiment, the apparatus further includes:

[0122] A molecular property attribute determination module, configured to obtain a trained property prediction model;

[0123] Input the chemical representation object and the molecular descriptor data into the property prediction model to obtain the output molecular property attributes of the molecule.

[0124] The molecular descriptor generation apparatus provided by the embodiments of the present invention can execute the molecular descriptor generation method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0125] Figure 6A schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0126] As Figure 6 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor 11. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0127] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information or data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0128] The processor 11 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for generating molecular descriptors provided in the above embodiments.

[0129] In some embodiments, the method for generating molecular descriptors provided in the above embodiments may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps in the method for generating molecular descriptors described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the method for generating molecular descriptors by any other suitable means (e.g., by means of firmware).

[0130] In particular, according to the embodiments of the present invention, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present invention include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network via the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above functions defined in the methods of the embodiments of the present invention are executed.

[0131] The various embodiments of the systems and techniques described above in this specification can be implemented in the following systems or combinations thereof: digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip (SOC), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0132] The computer programs for implementing the method for generating molecular descriptors of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0133] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable storage medium. Examples of machine-readable storage media would include electrical connections based on at least one wire, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0134] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a terminal device having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the terminal device. Other kinds of devices can also provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0135] The systems and techniques described herein can be implemented in a computing system that includes backend components (such as, for example, a data server), or a computing system that includes middleware components (such as, for example, an application server), or a computing system that includes frontend components (such as, for example, a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (such as, for example, a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0136] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.

[0137] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0138] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for generating molecular descriptors, characterized in that, Including: Obtaining a chemical representation object of a molecule and a conformational algorithm sequence; wherein, the conformational algorithm sequence includes at least two preset conformational generation algorithms arranged in sequence; Taking the first preset conformational generation algorithm in the conformational algorithm sequence as the target conformational generation algorithm, and using the target conformational generation algorithm to generate conformational data of the molecule according to the chemical representation object; If the generation by the target conformational generation algorithm fails, taking the next preset conformational generation algorithm in the conformational algorithm sequence as the target conformational generation algorithm, and returning to execute the step of using the target conformational generation algorithm to generate conformational data of the molecule according to the chemical representation object; Until the generation by the target conformational generation algorithm is successful, generating molecular descriptor data of the molecule according to the conformational data of the molecule.

2. The method according to claim 1, wherein The first preset conformational generation algorithm in the conformational algorithm sequence is a conformational generation algorithm based on specific knowledge or a conformational generation algorithm based on sampling, and at least one preset conformational generation algorithm other than the first preset conformational generation algorithm respectively includes a force field optimization algorithm; When the number of the force field optimization algorithms is at least two, at least two force field optimization algorithms are arranged in descending order according to the force field precision level corresponding to the type to which the molecule belongs.

3. The method according to claim 2, wherein The preset conformational generation algorithm other than the first preset conformational generation algorithm further includes a random coordinate algorithm; The step of using the target conformational generation algorithm to generate conformational data of the molecule according to the chemical representation object includes: When the target conformational generation algorithm includes a random coordinate algorithm and a force field optimization algorithm, using the random coordinate algorithm to generate atomic coordinate data according to the chemical representation object, and using the force field optimization algorithm to perform force field optimization on the atomic coordinate data to obtain conformational data of the molecule.

4. The method according to claim 1, wherein The step of generating molecular descriptor data of the molecule according to the conformational data of the molecule includes: Obtaining descriptor identification data corresponding to the molecule and a descriptor algorithm sequence; wherein, the descriptor algorithm sequence includes at least two preset descriptor generation algorithms arranged in sequence; Taking the first preset descriptor generation algorithm in the descriptor algorithm sequence as the target descriptor generation algorithm, and using the target descriptor generation algorithm to generate initial descriptor data according to the conformational data of the molecule; If the descriptor identification data is not included in the initial descriptor data, taking the next preset descriptor generation algorithm in the descriptor algorithm sequence as the target descriptor generation algorithm, and returning to execute the step of using the target descriptor generation algorithm to generate initial descriptor data according to the conformational data of the molecule; Until the descriptor identification data is included in the initial descriptor data, determining the molecular descriptor data of the molecule according to the descriptor identification data and the initial descriptor data.

5. The method according to claim 4, characterized in that, At least two preset descriptor generation algorithms in the descriptor algorithm sequence are arranged in descending order according to the algorithm precision.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Obtaining a trained property prediction model; Input the chemical representation object and the molecular descriptor data into the property prediction model to obtain the molecular property attributes of the output molecule.

7. A generating device for molecular descriptors, characterized in that, It includes: A conformational algorithm sequence acquisition module for acquiring a chemical representation object of a molecule and a conformational algorithm sequence; wherein, the conformational algorithm sequence includes at least two preset conformational generation algorithms arranged in sequence; A molecular conformation data determination module for using the first preset conformational generation algorithm in the conformational algorithm sequence as the target conformational generation algorithm, and using the target conformational generation algorithm to generate molecular conformation data according to the chemical representation object; A return execution module for, if the target conformational generation algorithm fails to generate, using the next preset conformational generation algorithm in the conformational algorithm sequence as the target conformational generation algorithm, and returning to execute the step of using the target conformational generation algorithm to generate molecular conformation data according to the chemical representation object; A molecular descriptor data generation module for generating the molecular descriptor data of the molecule according to the molecular conformation data until the target conformational generation algorithm generates successfully.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method for generating a molecular descriptor according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to execute the method for generating a molecular descriptor according to any one of claims 1-6 when executed.

10. A computer program product, including a computer program which, when executed by a processor, implements the method for generating a molecular descriptor according to any one of claims 1-6.