Molecular structure data generation method, device, equipment, medium and program product

By generating and converting the molecular structure data of the electrolyte, the problem of the lack of rapid generation and update of the molecular structure database of the electrolyte in the prior art is solved, and efficient electrolyte research and development is achieved, supporting high-throughput computing and machine learning applications.

CN119993321AActive Publication Date: 2025-05-13TSINGHUA UNIVERSITY
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202510080602.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The existing technology lacks methods to quickly generate and update the database of electrolyte molecular structures, limiting the application of high-throughput calculation, high-throughput screening and machine learning technologies in electrolyte research and development.

Method used

Provides a method for generating molecular structure data, by acquiring initial molecules and adding atomic and chemical bonds according to preset molecular generation rules, multiple target molecules are generated and converted to SMILES format to support high-throughput computing and machine learning applications.

Benefits of technology

It realizes efficient batch generation of molecular structure data, supports high-throughput computing, machine learning and high-throughput screening, significantly shortens the electrolyte R&D cycle, reduces costs, and improves design and optimization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993321A_ABST
    Figure CN119993321A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a molecular structure data generation method and device, equipment, a medium and a program product. The method comprises the steps that an initial molecule is obtained, the initial molecule is represented by a self-loop-free undirected graph, and the undirected graph comprises nodes used for representing atoms and edges used for representing chemical bonds; according to a preset molecule generation rule, preset atoms and / or preset chemical bonds are / is added into the initial molecules, a plurality of target molecules are obtained, the molecule generation rule comprises an atom increase rule, a chemical bond increase rule, a valence rule and a molecule stability rule, and the target molecules are represented by a self-loop-free undirected graph; and converting the plurality of target molecules into a simplified molecular linear input specification SMILES format to obtain a plurality of target molecular structure data. According to the embodiment of the invention, the molecular structure data can be efficiently generated in batches, so that application scenes such as machine learning, high-throughput calculation and high-throughput screening can be supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of chemical informatics, and in particular, relates to a method, device, equipment, medium and program product for generating molecular structure data. Background Art

[0002] With the development of battery technology, batteries are widely used in portable electronic devices, electric vehicles and energy storage systems. Among the components of batteries, electrolyte, as a key electrochemical medium, directly affects the energy density, cycle life and safety performance of batteries. Therefore, the research and development of high-performance electrolytes has become an important direction for the development of battery technology.

[0003] Traditional electrolyte research and development mainly relies on experience and trial and error, testing electrolyte molecules one by one in the laboratory to find the electrolyte with the best performance. This method is not only time-consuming and costly, but also due to the limited test samples, it can only find the local optimal solution and it is difficult to achieve comprehensive optimization. In this context, there is an urgent need for an efficient method to accelerate the development of electrolytes. The development of big data and artificial intelligence technology has provided new possibilities for the rapid development of electrolytes. High-throughput computing and high-throughput screening technologies can evaluate the performance of a large number of candidate molecules in a short period of time, greatly shortening the research and development cycle; machine learning algorithms can predict the performance of new materials and guide experimental design by analyzing a large amount of experimental data. The application of these technologies requires a large amount of molecular structure data as a basic support.

[0004] However, there is currently a lack of a method that can quickly generate and update electrolyte molecular structure databases, which limits the application of high-throughput computing, high-throughput screening, and machine learning technologies in electrolyte research and development. Summary of the invention

[0005] The embodiments of the present application provide a molecular structure data generation method, apparatus, device, computer storage medium and computer program product, which can efficiently generate molecular structure data in batches, thereby supporting application scenarios such as machine learning, high-throughput computing and high-throughput screening.

[0006] In a first aspect, an embodiment of the present application provides a method for generating molecular structure data, comprising:

[0007] Obtain an initial molecule, where the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds;

[0008] According to a preset molecule generation rule, preset atoms and / or preset chemical bonds are added to the initial molecule to obtain a plurality of target molecules, wherein the molecule generation rule includes an atom addition rule, a chemical bond addition rule, a valence rule, and a molecule stability rule, and the target molecule is represented by an undirected graph without self-loops;

[0009] The multiple target molecules are converted into the simplified molecular linear input specification SMILES format to obtain the structural data of the multiple target molecules.

[0010] In an optional embodiment, the preset molecule generation rule includes executing the following steps for each initial molecule:

[0011] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of first molecules;

[0012] For each first molecule of the plurality of first molecules, according to a chemical bond addition rule, a valence rule, and a molecular stability rule, a preset type of chemical bond is added to the first molecule to obtain a plurality of second molecules;

[0013] The plurality of first molecules and the plurality of second molecules are stored as a plurality of primary target molecules in a preset target molecule set;

[0014] For each of the first-level target molecules among the plurality of first-level target molecules, a preset number of preset atoms are added to the first-level target molecule according to the atom addition rule, the valence rule and the molecular stability rule to obtain a plurality of third molecules; for each of the third molecules among the plurality of third molecules, a preset type of chemical bond is added to the third molecule according to the chemical bond addition rule, the valence rule and the molecular stability rule to obtain a plurality of fourth molecules; the plurality of third molecules and the plurality of fourth molecules are taken as a plurality of second-level target molecules and stored in a preset target molecule set;

[0015] Taking the multiple secondary target molecules as new multiple primary target molecules, returning to each of the multiple primary target molecules, respectively adding a preset number of preset atoms to the primary target molecule according to the atom addition rule, valence rule and molecular stability rule, to obtain multiple third molecules until the preset iteration stop condition is met.

[0016] In an alternative embodiment, the molecular stability rule includes a molecular structure diagram that does not comply with the molecular stability;

[0017] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of first molecules, including:

[0018] Adding a preset number of preset atoms to the initial molecule according to the atom addition rule and the valence rule to obtain a plurality of initial first molecules;

[0019] searching for a molecular structure diagram among a plurality of initial first molecules;

[0020] The initial first molecule including the molecular structure graph is deleted to obtain a plurality of first molecules.

[0021] In an alternative embodiment, the initial molecule includes at least one non-hydrogen atom;

[0022] According to the atomic addition rule and the valence rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of initial first molecules, including performing the following steps on the initial molecule for each preset atom:

[0023] Get the number of preset atoms in the initial molecule;

[0024] When the number of preset atoms is less than a preset threshold, for each non-hydrogen atom in the initial molecule, the following steps are performed respectively: the atomic degree of the non-hydrogen atom is calculated, the atomic degree includes the sum of the weighted values ​​of the edges connected to the corresponding nodes of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attributes, and the edge attributes are used to represent the attributes of the chemical bond corresponding to the edge; according to the corresponding relationship between the preset atomic degree and the chemical bond type, the first chemical bond type corresponding to the atomic degree of the non-hydrogen atom is determined; the non-hydrogen atom and the preset atom are connected according to the first chemical bond type to obtain the initial first molecule.

[0025] In an optional embodiment, after the non-hydrogen atom is connected to the preset atom according to the first chemical bond type to obtain the initial first molecule, the method further comprises:

[0026] Perform graph isomorphism judgment on the molecular graph level between the initial first molecule and other initial first molecules;

[0027] If there are other initial first molecules that are isomorphic to the initial first molecule graph, the initial first molecule is deleted.

[0028] In an optional embodiment, for each first molecule of the plurality of first molecules, according to the chemical bond addition rule, the valence rule and the molecular stability rule, a preset type of chemical bond is added to the first molecule to obtain a plurality of second molecules, including:

[0029] For each first molecule among the plurality of first molecules, adding a preset type of chemical bond in the first molecule according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules;

[0030] searching for a molecular structure diagram among a plurality of initial second molecules;

[0031] The initial second molecule including the molecular structure graph is deleted to obtain a plurality of second molecules.

[0032] In an alternative embodiment, the first molecule comprises at least two non-hydrogen atoms;

[0033] For each first molecule among the plurality of first molecules, a preset type of chemical bond is added to the first molecule according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules, including performing the following steps for each first molecule among the plurality of first molecules:

[0034] For any two non-hydrogen atoms in the first molecule, obtain the target atomic degrees and target edge attributes of the two non-hydrogen atoms respectively;

[0035] According to the preset correspondence between the atomic degrees of the two non-hydrogen atoms and the edge attributes and the chemical bond type, determine the second chemical bond type corresponding to the target atomic degrees and the target edge attributes of the two non-hydrogen atoms;

[0036] The two non-hydrogen atoms are connected according to the second chemical bond type to obtain an initial second molecule.

[0037] In an optional embodiment, after connecting two non-hydrogen atoms according to the second chemical bond type to obtain an initial second molecule, the method further comprises:

[0038] Performing graph isomorphism judgment on the initial second molecule, the plurality of initial first molecules and other initial second molecules at the molecular graph level;

[0039] When there is an initial first molecule or other initial second molecules that are graph-isomorphic to the initial second molecule, the initial second molecule is deleted.

[0040] In an optional embodiment, the preset molecule generation rule includes executing the following steps for each initial molecule:

[0041] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of fifth molecules;

[0042] storing a plurality of fifth molecules into a preset target molecule set;

[0043] Determine a target fifth molecule from a plurality of fifth molecules, where the target fifth molecule is a fifth molecule having the least number of edges among the plurality of fifth molecules;

[0044] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the target fifth molecule to obtain a plurality of sixth molecules;

[0045] storing a plurality of sixth molecules into a preset target molecule set;

[0046] Taking the plurality of sixth molecules as new plurality of fifth molecules, and returning to determine the target fifth molecule from the plurality of fifth molecules until a preset iteration stop condition is satisfied;

[0047] According to the chemical bond addition rule, the valence rule, and the molecular stability rule, respectively adding chemical bonds of a preset type to a plurality of fifth molecules and a plurality of sixth molecules in the preset target molecule set to obtain a plurality of seventh molecules;

[0048] storing a plurality of seventh molecules into a preset target molecule set;

[0049] According to the preset graph isomorphism merging rule, the molecules with graph isomorphism in the preset target molecule set are merged.

[0050] In an optional implementation, the graph isomorphism merging rule includes a graph isomorphism merging rule based on a step-by-step merging algorithm.

[0051] In an optional implementation, the graph isomorphism merging rule includes a graph isomorphism merging rule based on a multi-grouping algorithm.

[0052] In an alternative embodiment, the molecular stability rule includes a molecular structure diagram that does not comply with the molecular stability;

[0053] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain multiple fifth molecules, including:

[0054] Adding a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain a plurality of initial fifth molecules;

[0055] Finding a molecular structure diagram among a plurality of initial fifth molecules;

[0056] The initial fifth molecule including the molecular structure diagram is deleted to obtain a plurality of fifth molecules.

[0057] In an alternative embodiment, the initial molecule includes at least one non-hydrogen atom;

[0058] According to the atomic addition rule and the valence rule, a preset number of preset atoms are added to the initial molecule to obtain multiple initial fifth molecules, including:

[0059] Get the number of preset atoms in the initial molecule;

[0060] When the number of preset atoms is less than a preset threshold, for each non-hydrogen atom in the initial molecule, the following steps are performed respectively: the atomic degree of the non-hydrogen atom is calculated, the atomic degree includes the sum of the weighted values ​​of the edges connected to the corresponding nodes of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attributes, and the edge attributes are used to represent the properties of the chemical bond corresponding to the edge; according to the corresponding relationship between the preset atomic degree and the chemical bond type, the third chemical bond type corresponding to the atomic degree of the non-hydrogen atom is determined; the non-hydrogen atom and the preset atom are connected according to the third chemical bond type to obtain the initial fifth molecule.

[0061] In an optional embodiment, according to the chemical bond addition rule, the valence rule and the molecular stability rule, a preset type of chemical bond is added to a plurality of fifth molecules and a plurality of sixth molecules in the preset target molecule set to obtain a plurality of seventh molecules, including:

[0062] According to the chemical bond addition rule and the valence rule, chemical bonds of preset types are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set respectively, so as to obtain the plurality of initial seventh molecules and search the molecular structure diagram in the plurality of initial seventh molecules;

[0063] The initial seventh molecule including the molecular structure diagram is deleted to obtain a plurality of seventh molecules.

[0064] In an alternative embodiment, the fifth molecule and the sixth molecule each include at least two non-hydrogen atoms;

[0065] According to the chemical bond addition rule and the valence rule, a preset type of chemical bond is added to a plurality of fifth molecules and a plurality of sixth molecules in the preset target molecule set to obtain a plurality of initial seventh molecules, including performing the following steps for each molecule in the plurality of fifth molecules and the plurality of sixth molecules:

[0066] For any two non-hydrogen atoms in the molecule, obtain the target atomic degrees and target edge attributes of the two non-hydrogen atoms respectively;

[0067] According to the preset correspondence between the atomic degrees of the two non-hydrogen atoms and the edge attributes and the chemical bond types, determine the fourth chemical bond type corresponding to the target atomic degrees and the target edge attributes of the two non-hydrogen atoms;

[0068] Two non-hydrogen atoms are connected according to the fourth chemical bond type to obtain a plurality of initial seventh molecules.

[0069] In an alternative embodiment, the molecular stability rule includes a molecular structure diagram that does not comply with the molecular stability;

[0070] Before adding preset atoms and / or preset chemical bonds to the initial molecule according to the preset molecule generation rule, the method further includes:

[0071] Application environment for obtaining structural data of multiple target molecules;

[0072] According to the chemical stability rules and the application environment, the molecular structure diagram that does not meet the molecular stability in the application environment is determined.

[0073] In an optional embodiment, the preset atom includes at least one of carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl).

[0074] In an optional embodiment, multiple target molecules are converted into SMILES format to obtain multiple target molecule structure data, including:

[0075] Convert multiple target molecules into an adjacency matrix or adjacency table format to obtain intermediate data of target molecule structures;

[0076] The target molecular structure intermediate data is converted into SMILES format to obtain multiple target molecular structure data.

[0077] In an optional embodiment, after converting the multiple target molecules into SMILES format to obtain the multiple target molecule structure data, the method further includes:

[0078] Based on multiple target molecular structure data, a target molecular structure database is created.

[0079] In a second aspect, an embodiment of the present application provides a molecular structure data generating device, comprising:

[0080] An acquisition module, used for acquiring an initial molecule, wherein the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds;

[0081] An adding module, used for adding preset atoms and / or preset chemical bonds to the initial molecules according to preset molecular generation rules to obtain multiple target molecules, wherein the molecular generation rules include atom addition rules, chemical bond addition rules, valence rules and molecular stability rules, and the target molecules are represented by an undirected graph without self-loops;

[0082] The conversion module is used to convert multiple target molecules into SMILES format to obtain multiple target molecular structure data.

[0083] In a third aspect, an embodiment of the present application provides an electronic device, the device comprising: a processor and a memory storing computer program instructions;

[0084] When the processor executes the computer program instructions, it implements the molecular structure data generating method as described in any optional embodiment of the first aspect of the present application.

[0085] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, a method for generating molecular structure data according to any optional implementation of the first aspect of the present application is implemented.

[0086] In a fifth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes a method for generating molecular structure data as described in any optional implementation of the first aspect of the present application.

[0087] The molecular structure data generation method, device, equipment, computer storage medium and computer program product of the embodiment of the present application can obtain the initial molecule, which is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds. In this way, by using the graph theory method to represent the molecule, it is possible to ensure the structuring and standardization of the molecular data in the molecular generation process, thereby improving the efficiency of data processing. Then, according to the preset molecular generation rules, preset atoms and / or preset chemical bonds can be added to the initial molecule to obtain multiple target molecules, wherein the molecular generation rules include atomic increase rules, chemical bond increase rules, valence rules and molecular stability rules, and the target molecule is represented by an undirected graph without self-loops. On the one hand, the molecular generation rules can make the structure of the generated target molecule have chemical rationality and stability, reducing the risk of generating unreasonable molecules; on the other hand, by specifying atomic increase rules and chemical bond increase rules, batches of target molecules that meet user needs can be generated, thereby improving the efficiency and flexibility of molecular generation. In this way, the generated multiple target molecules can adapt to the needs of different research scenarios. Subsequently, multiple target molecules can also be converted into SMILES format to obtain multiple target molecular structure data. SMILES format is a more standardized molecular representation, which is convenient for retrieval and storage, and there is no ambiguity. By converting multiple target molecules into SMILES format, the consistency and accuracy of the target molecule structure data can be guaranteed. In this way, a solid data foundation can be provided for high-throughput computing, machine learning and high-throughput screening. The target molecule structure data generated in batches in the application embodiment are used to construct an electrolyte molecular structure database, which can effectively support the rapid development of new electrolyte systems, significantly shorten the development cycle, reduce research and development costs, and improve the efficiency of electrolyte design and optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0089] Figure 1 It is a flowchart of a method for generating molecular structure data provided by an embodiment of the present application;

[0090] Figure 2 is a schematic structural diagram of a molecular structure data generating device provided in yet another embodiment of the present application;

[0091] Figure 3 A schematic structural diagram of an electronic device provided in yet another embodiment of the present application. DETAILED DESCRIPTION

[0092] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.

[0093] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0094] As described in the background technology, there is currently a lack of a method that can quickly generate and update the electrolyte molecular structure database, which limits the application of high-throughput computing, high-throughput screening and machine learning technologies in electrolyte research and development.

[0095] In the field of chemistry, a molecule is a whole structure composed of its constituent atoms according to specific valence bond rules and spatial arrangements. Within a molecule, each atom interacts with each other, and when the attractive and repulsive forces of these interactions reach a balance, the molecule is in a stable state.

[0096] The bonding structure and arrangement of atoms in a molecule is called molecular structure. Methods for describing molecular structure usually include structural formula, bond-line formula, ball-and-stick formula, and ratio formula, which can accurately describe the structural information of molecules in chemistry. However, in computer languages, relying solely on these representation methods often cannot accurately describe molecular structure, and there are many ambiguous problems (for example, there may be a one-to-many situation when using structural formulas to build a database).

[0097] In view of this, the inventor, after in-depth thinking, cleverly proposed a molecular structure data generating method, information receiving method, device, equipment, computer storage medium and computer program product.

[0098] In the embodiment of the present application, the molecule can be regarded as an undirected graph without self-loops in graph theory, denoted by G=<V,E> , where each atom is regarded as a set of points V = {v1,v2,...,v n}, each key can be regarded as a set of edges E = {e1,e2,...,e m}, the atomic information is stored in the attribute set A as the attribute of the point set V, denoted as A = {a1, a2, ..., a n}, the key information is stored as the attribute of the edge set E in the attribute set B, denoted as B = {b1, b2, ..., b m}.

[0099] In the undirected graph corresponding to the molecule, the atomic degree of a node can be expressed as the sum of the weighted values ​​of the edges connected to the node and the attributes on the edges. As an example, the atomic degree of node V1 can be recorded as d1=b1+b2+...+b p , where the number of edges connected by V1 is p, and the number of edges from b1 to b p They can respectively represent the weighted values ​​of the attributes on the first to pth edges and the attributes on the edges. In some embodiments, when the edge is a single bond, the weighted value of the edge and the attribute on the edge can be 1; when the edge is a double bond, the weighted value of the edge and the attribute on the edge can be 2; when the edge is a triple bond, the weighted value of the edge and the attribute on the edge can be 3. The node degree can be represented as not including the set of attributes on the edge, which can be equal to the number of edges connected to the node. Generally, the atomic degree of a node is usually less than or equal to 4.

[0100] In conjunction with the accompanying drawings, the molecular structure data generation method provided by the embodiment of the present application is introduced through specific embodiments and their application scenarios. The molecular structure data generation method provided by the embodiment of the present application, the device for executing the sending can be a molecular structure data generation device, or a partial module in the molecular structure data generation device for executing the molecular structure data generation method. In the embodiment of the present application, the molecular structure data generation method provided by the embodiment of the present application is described in detail by taking the molecular structure data generation device executing the molecular structure data generation method as an example.

[0101] The following is combined with Figure 1 The molecular structure data generation method provided in the embodiments of the present application is described in detail.

[0102] Figure 1 FIG. 1 is a flow chart of a method for generating molecular structure data provided by an embodiment of the present application. Figure 1 As shown, the molecular structure data generating method may specifically include the following steps S110 to S130.

[0103] S110, obtaining an initial molecule, where the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds.

[0104] In step S110, the initial molecule may be a basic molecule for molecular amplification. The initial molecule may be selected according to the needs of the user and is not limited here. The number of initial molecules may be one or more. It is understood that when the number of initial molecules is more than one, the amplification of atoms and chemical bonds may be performed for each of the multiple initial molecules.

[0105] S120, according to the preset molecular generation rules, adding preset atoms and / or preset chemical bonds to the initial molecules to obtain multiple target molecules, wherein the molecular generation rules include atom addition rules, chemical bond addition rules, valence rules and molecular stability rules, and the target molecules are represented by an undirected graph without self-loops.

[0106] In step S120, the preset atoms can be selected according to the needs of the user, which is not limited here. As an example, the preset atoms can include non-hydrogen atoms. For example, when the user needs to construct molecular structure data containing only carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl) elements, the initial molecule can be a molecule composed of at least one of carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl), and the preset atoms can include carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl). The atom increase rule can be used to specify the position, conditions, and method of adding new atoms to the molecule and adding chemical bonds accordingly, and the chemical bond increase rule can be used to specify the position, conditions, and method of adding chemical bonds to the molecule. The valence rule can be used to specify the basic chemical rules that an atom needs to meet to form a bond with other atoms, so that the number of other atoms that the atom can connect to is saturated. The molecular stability rule can be used to ensure that the generated target molecule meets the chemical stability requirements. It is understandable that when adding atoms and / or chemical bonds to a molecule, the atoms connected to the newly added atoms or chemical bonds can correspondingly lose a corresponding number of hydrogen atoms according to the valence rules.

[0107] In step S120, according to the molecular generation rules, different types or quantities of preset atoms and / or preset chemical bonds can be added to the initial molecule in a variety of ways, and the initial molecule can also be iteratively amplified with atoms and / or chemical bonds, thereby efficiently and automatically generating a large number of target molecules in batches.

[0108] S130, converting the multiple target molecules into SMILES format to obtain multiple target molecule structure data.

[0109] The molecular structure data generation method of the embodiment of the present application can obtain the initial molecule, and the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds. In this way, by using the graph theory method to represent the molecule, it is possible to ensure the structuring and standardization of the molecular data in the molecular generation process, thereby improving the efficiency of data processing. Then, according to the preset molecular generation rules, preset atoms and / or preset chemical bonds can be added to the initial molecule to obtain multiple target molecules, wherein the molecular generation rules include atomic increase rules, chemical bond increase rules, valence rules and molecular stability rules, and the target molecule is represented by an undirected graph without self-loops. On the one hand, the molecular generation rules can make the structure of the generated target molecule have chemical rationality and stability, reducing the risk of generating unreasonable molecules; on the other hand, by specifying atomic increase rules and chemical bond increase rules, batch generation of target molecules that meet user needs can improve the efficiency and flexibility of molecular generation. In this way, the generated multiple target molecules can adapt to the needs of different research scenarios. Subsequently, multiple target molecules can also be converted into SMILES format to obtain multiple target molecular structure data. SMILES format is a more standardized molecular representation, which is convenient for retrieval and storage, and there is no ambiguity. By converting multiple target molecules into SMILES format, the consistency and accuracy of the target molecule structure data can be guaranteed. In this way, a solid data foundation can be provided for high-throughput computing, machine learning and high-throughput screening. The target molecule structure data generated in batches in the application embodiment are used to construct an electrolyte molecular structure database, which can effectively support the rapid development of new electrolyte systems, significantly shorten the development cycle, reduce research and development costs, and improve the efficiency of electrolyte design and optimization.

[0110] In one embodiment, the molecular structure data can be generated based on a serial molecular generation algorithm. This solution can be implemented based on a serial generation algorithm system. For ease of understanding, the serial generation algorithm system is briefly introduced below.

[0111] The serial generation algorithm system can include an initial molecule generation module, a graph constraint module, an atom augmentation module, a bond augmentation module, a format conversion module and an auxiliary operation module. In the serial generation algorithm system, the initial molecule can be used as the original input, and new nodes are added according to possible structural combinations of atoms and bonds, and isomorphic pruning is judged at the same time. Single bonds, double bonds and triple bonds are added according to the breadth-first search combination. With the help of the auxiliary queue, the graph with the new node added is used as the root node, and it is queued. After the queue is removed, the edge is added, and the existing graph is searched to judge the isomorphism. The graph with the current number of nodes and the current number of edges is not generated or is not in the same structure as the existing graph, and it is queued and looped.

[0112] The various modules described in the serial generation algorithm are described in detail below.

[0113] In the initial molecule generation module, the initial molecule can be represented according to the graph theory rules, can be encapsulated by defining classes or functional modules, or can be reasonably stored using data structures such as adjacency matrices or adjacency lists. During the construction process, in order to ensure the consistency of the generated database, all elements in the database should be molecules and should not contain isolated atoms. In other words, the initial molecule and each output result generated by its iteration should be a complete and independent molecule. In order to simplify the database construction algorithm and from the perspective of reducing space complexity and time complexity, the representation of hydrogen atoms is omitted in the molecule. That is, if it is found that an atom in the molecule does not meet the valence rule, the missing valence is automatically filled by the hydrogen atom. This representation method can use functional functions to automatically complete the molecule when querying the database to avoid ambiguity that may occur in database search. Therefore, in the process of database construction and generation, the molecule representation method can be both concise and reasonable.

[0114] For example, if you want to build a database that only contains carbon, hydrogen, nitrogen, oxygen, and fluorine elements, then at the beginning of the construction, you only need to define the strings "C", "N", "O", and "F" and store them in a reasonable form. It should be noted that although the initial molecules look like single atoms, they actually represent "CH4", "NH3", "H2O", and "HF" molecules respectively.

[0115] The form of the initial molecule can be determined according to the system being studied and the actual situation, and the embodiments of the present application do not limit this. For example, if you want to explore the relationship between the structure and properties of lithium battery electrolyte solvent molecules, and want to build an initial database for applications such as machine learning, you need to consider the atoms that may exist in the electrolyte molecules. Specifically, most electrolyte molecules usually contain functional groups such as ethers and esters, and the introduction of fluorine atoms helps to improve the stability of the electrolyte. Therefore, when constructing the initial data set, water molecules and hydrogen fluoride molecules should be introduced to provide the required oxygen and fluorine elements.

[0116] The graph constraint module can be used to delete molecules that do not meet the molecular stability rules. In the process of generating the electrolyte molecule database, chemically unstable structures will inevitably be generated. It is necessary to find these unstable structures and efficiently delete the molecules containing unstable structures.

[0117] In mathematics, if the nodes and edges between nodes in a graph S can be mapped to another graph G, then it is considered that the graph S and the graph G satisfy the subgraph isomorphism relationship. In the embodiment of the present application, subgraph isomorphism plays an important role in searching for functional groups in molecules and finding related topological structures. Using the principle of subgraph isomorphism, molecular graphs containing unstable substructures can be efficiently deleted. Taking the structure containing carbon atoms and oxygen atoms as an example, it is recommended to delete the following molecular graphs:

[0118] (1) Molecular diagrams containing double bonds or triple bonds in small rings (three-membered rings or four-membered rings);

[0119] (2) Molecular diagrams containing bridgehead atoms in small rings;

[0120] (3) A molecular diagram containing a spirocycle with two small rings on both sides.

[0121] The purpose of eliminating these molecular graphs is mainly to remove stress-concentrated structures. These structures with large ring tension or topological distortion are difficult to synthesize in practice. Even if they can be constructed in a computer, they have no practical significance for research. In some embodiments, molecular graphs that are unstable in the application environment can also be deleted. The constraints of such molecular graphs can be determined based on prior knowledge of the molecular structure. For example, in a metal lithium battery system, it is usually required that the molecular structure does not contain active hydrogen (i.e., hydrogen atoms on groups such as hydroxyl and carboxyl groups), so molecular graphs containing active hydrogen can be deleted.

[0122] The atomic augmentation module is used to add a preset number of new atoms to each molecule and add corresponding bonds. The organic small molecule database generation algorithm needs to call the atomic augmentation module when iterating and adding points. The preset number can be set according to actual needs. For example, 1 new atom, 2 new atoms, 3 new atoms, etc. can be added to each molecule in each round, which is not limited here. Exemplarily, the preset number can be 1, and the atomic augmentation module can be called only once during one round. For example, in each round of iteration, an atom can be added to the molecule of the current round through the atomic augmentation module.

[0123] Taking the preset number of 1 as an example, the atomic augmentation module can take the number of atoms in the molecule in the current round as input, and output the set of molecular graphs after adding an atom and the corresponding chemical bond. When the atomic augmentation module starts running, it first determines whether the number of atoms in the current round is equal to the set upper limit of the number of atoms. If they are equal, it returns directly, indicating that the iteration is completed. If they are not equal, continue to perform subsequent operations. After judgment, enter the loop operation, extract all molecular graphs generated in the previous round, if the previous round is the initial molecule, extract all the initial molecular graphs, and then consider whether each atom in each molecule meets the following conditions (only common atoms are analyzed as examples here).

[0124] 1. Carbon Atom

[0125] If the atom currently traversed is a carbon atom, first check whether the number of atoms connected to the carbon atom in the current molecule has reached the preset upper limit. If so, skip the atom, but do not skip the current molecule graph, and continue to traverse other atoms. If not, perform subsequent operations.

[0126] (1) Check the degree of the current carbon atom to determine whether it is less than 4. The degree here refers to the weighted number considering the chemical bond. If the degree is greater than or equal to 4, skip this judgment; if the degree is less than 4, continue to perform subsequent operations.

[0127] Next, consider the candidate elements (excluding hydrogen) contained in the current round of molecules. If these elements can form single bonds with carbon atoms, then all candidate atoms corresponding to the candidate elements are connected to the carbon atom with single bonds in turn, and then store the molecular graphs with one node and one edge added. For example, if the molecule of the current round is the initial molecule, and the elements it contains are carbon, oxygen, nitrogen, and fluorine, then the candidate elements can be carbon, oxygen, nitrogen, and fluorine. At this time, the atoms that can form single bonds with carbon atoms are carbon atoms, oxygen atoms, nitrogen atoms, and fluorine atoms. Therefore, the atoms and bonds that can be added are "C-C", "C-O", "C-N", and "C-F" respectively. These substructures will be saved in the storage set of the graph as part of the newly generated molecular graph.

[0128] It is recommended to copy the original molecular graph before performing atomic augmentation, rather than directly operating on the graph in the loop process, because in the loop process, the base graph needs to add the candidate atoms in turn to generate a new molecular graph. Due to the limitations of different programming languages, not copying may lead to unexpected errors, which in turn lead to the overwriting and addition of atoms.

[0129] (2) Check the degree of the carbon atom. The carbon atom here refers to the atom without adding new atoms and new bonds, that is, the carbon atom in the cyclic base graph, and determine whether its degree is less than 3, and check whether a double bond already exists. If the degree of the carbon atom is greater than or equal to 3 or there is a double bond, skip this judgment; if the degree is less than 3 and the carbon atom is not connected to a double bond, continue to perform subsequent operations. It should be noted that the purpose of judging whether a carbon atom has a double bond here is to avoid two double bonds on a carbon atom, that is, the formation of cumulative dienes. Such a structure is extremely unstable in the case of small molecular weight, and is easy to react into other substances or cannot be synthesized.

[0130] Next, you need to select an atom that can form a double bond with the carbon atom from the set of atoms to be selected, and form a new molecular graph with the carbon atom, that is, add a new molecule with an atom and a double bond, and save them in the graph set in sequence. For example, the elements contained in the initial molecule are carbon, oxygen, nitrogen and fluorine, then the atoms that can form a double bond with the carbon atom are carbon atoms, oxygen atoms and nitrogen atoms, so the atoms and bonds that can be added are "C=C", "C=O" and "C=N", respectively, and then the molecular graphs with this substructure added are stored separately.

[0131] (3) Then check the degree of the carbon atom again to determine whether it is less than 2. If the degree of the carbon atom is greater than or equal to 2, skip this judgment; if the degree is less than 2, continue to perform subsequent operations. The reason why it is not necessary to consider whether it is connected to a double bond here is that if it already has a double bond, the maximum degree can only be equal to 2, and it cannot be less than 2. Therefore, the instability problem of two double bonds on a carbon atom can be eliminated by only using the degree restriction. Next, select atoms that can form triple bonds with carbon atoms from the set of atoms to be selected, and add them to the graph set after forming a bond with them. For example, the atoms that can form triple bonds with carbon atoms at this time are carbon atoms and nitrogen atoms. Therefore, the atoms and bonds that can be added are "C≡C" and "C≡N" respectively, and then the molecular graphs with triple bonds and a new atom are stored in sequence.

[0132] 2. Nitrogen Atom

[0133] If the atom currently traversed is a nitrogen atom, similar to the way of processing carbon atoms, first check whether the number of atoms connected to the nitrogen atom exceeds the given range, which will not be repeated here. Then, check the atomic degree of the nitrogen atom. If its degree is less than 3, find an atom that can form a single bond with it from the set of atoms to be selected, and add atoms and bonds in sequence before storing them; if its degree is less than 2, find an atom that can form a double bond with it from the set of atoms to be selected, and add atoms and bonds in sequence before storing them. It should be noted that at this stage, the nitrogen-nitrogen triple bond cannot be added. This is because the traversed nitrogen atom has at least one single bond connected to it in the original molecular graph to ensure the connectivity of the original molecular graph. If another triple bond is added, the nitrogen atom will be bonded to four atoms, which does not comply with the valence rules (more complex situations such as coordination bonds are not considered here).

[0134] (3) Oxygen atom

[0135] If the atom currently traversed is an oxygen atom, first check whether the number of atoms connected to the oxygen atom exceeds the given range. Then, check whether the degree of the oxygen atom is less than 2. Since oxygen has two lone pairs of electrons, it can usually only form bonds with two atoms. Therefore, if and only if this is the case, consider selecting atoms in the candidate set to form single bonds with the oxygen atom, and save the newly generated molecular graphs to the corresponding graph sets. Similar to nitrogen atoms, oxygen atoms cannot form double bonds with other atoms when they are atoms in the traversed base graph during the atomic augmentation stage.

[0136] (4) Fluorine atom

[0137] If the current traversal is a fluorine atom, you only need to consider if this is the initial molecule, then you can add an atom and a single bond, such as the fluorine gas molecule (F2). It is more recommended to consider this situation when constructing the initial molecule, so as to directly skip these possible cases.

[0138] After adding atoms and bonds, it is necessary to determine the graph isomorphism. This process should be performed before storing the molecular graph to reduce space complexity. In graph theory, graph isomorphism refers to a set of graphs with the same number of nodes and edges, and the nodes and edges have a bijective relationship. In topology, we believe that this set of graphs is the same. Specifically, you can call ready-made toolkits to determine graph isomorphism. Common graph isomorphism toolkits include NAUTY based on C / C++ language, Networkx based on Python language, etc. Graph isomorphism determination is a very time-consuming process, so you need to find a way to optimize this process. In addition to choosing a suitable and efficient algorithm, you can group and store the graphs, because only molecules with the same nodes, the same edges, and the same degree list may have graph isomorphism problems, so these three conditions can be used to make judgments in advance to reduce the large amount of time consumed by the isomorphism process. It should also be noted in this process that since non-planar graphs may appear in our generation process, we should also remove K5 and K at this stage. 3,3 Graphs with isomorphic subgraphs.

[0139] The above four representative cases respectively indicate that the atoms in the traversed base graph have 4, 3, 2, and 1 lone pairs of electrons, that is, they can form bonds with up to 4, 3, 2, and 1 atoms. They have certain universality. If silicon atoms need to be considered in the construction process, only the steps of carbon atoms (C) need to be migrated to silicon (Si). If sulfur atoms (S) need to be considered, only the steps of oxygen atoms (O) need to be migrated. That is, considering the similarity of atoms of the same main group elements, the algorithm can be well migrated, thereby greatly increasing the reuse of the algorithm and making the program simple and efficient. Therefore, the atom augmentation algorithm is universal for the generation process of various organic small molecules, and can be used to add points to the atomic sets of different selected elements as needed, which reflects its good adaptability.

[0140] The bond augmentation module is used to add an atom and a bond to a molecule, and then continue to iteratively add bonds until saturation is reached or no further additions can be made due to given constraints. In the serial algorithm, the bond augmentation module is closely related to the atom augmentation module, that is, the input of the bond augmentation module comes from the output of the atom augmentation module. In each round, as long as the atom augmentation module provides an output, the bond augmentation module starts running immediately, which ensures the efficiency of the serial algorithm and reduces the waiting time between modules.

[0141] In general, the bond augmentation module in the serial algorithm adopts a breadth-first generation strategy and uses an auxiliary queue as its auxiliary data structure. The input of the module includes the molecular graph from the atomic augmentation module, the current number of nodes and the number of edges. For each input, first build an auxiliary queue. Initially, the queue is empty. Add the input of the module, that is, a single molecular graph, to the queue and perform a cyclic judgment. When the queue is not empty, perform the following operations: take a molecular graph from the queue (only one graph at a time), and then traverse all atoms on the graph. If there are two atoms that satisfy the valence of each atom that is not saturated, then the two can add bonds.

[0142] The following example illustrates the execution process of the key augmentation module in detail.

[0143] 1. Carbon Atom – Carbon Atom

[0144] When the two atoms traversed are two carbon atoms, the premise is that there is no chemical bond between the two atoms, and both atoms can be connected to new atoms, that is, their valence bonds are not saturated. If the degrees of the two carbon atoms are less than 4, the two carbon atoms can be connected with a single bond and the result can be saved. It should be noted that, as described in the atomic augmentation section, it is recommended to copy the base graph from the loop to avoid superposition during the bond addition process. If the degrees of the two carbon atoms are less than 3, and each carbon atom itself has no other double bonds connected, a double bond can be added between the two carbon atoms. If the degrees of the two carbon atoms are less than 2, a triple bond can be added between the two carbon atoms. It is particularly emphasized that in the process of atomic augmentation or bond augmentation, they are all serial operations. If the degree meets the requirements, it is necessary to continue to judge next, instead of directly jumping out of the current round. Specifically, when the degrees of both carbon atoms are 2 and neither of them has a double bond connected to them, during the bond augmentation process, the degree is first determined to be less than 4. At this time, the base graph is copied and a new single bond is added between the two carbon atoms. After isomorphism judgment, it is saved in the graph collection, but this does not mean that this round of the loop is over. It is also necessary to determine whether the degrees of both carbon atoms are less than 3 and there are no double bonds connected to them. At this time, the base graph will be copied again, and a double bond will be added between the two carbon atoms. After isomorphism judgment, it will also be saved. Finally, determine whether the degrees of both carbon atoms are less than 2. If it is found that the condition is not met, the loop will be jumped out and the judgment of other atom pairs will be carried out.

[0145] 2. Carbon atom – Nitrogen atom

[0146] When the two atoms traversed are carbon and nitrogen atoms, if the degree of the carbon atom is less than 4 and the degree of the nitrogen atom is less than 3, a single bond can be added between the two. If the degree of the carbon atom is less than 3 and the degree of the nitrogen atom is less than 2, a double bond can be added between the two. As mentioned above, since the two atoms are connected by at least one chemical bond in the bond augmentation process to ensure the connectivity of the molecular graph, considering the valence bond rule, it is impossible to add an additional triple bond to the nitrogen atom (not considering complex situations such as coordination bonds). When generating each new molecular graph, the graph isomorphism judgment should be performed and then stored in the graph collection.

[0147] The bonding methods between other atoms are similar to the above cases and will not be discussed one by one here.

[0148] The format conversion module converts molecules represented in graph theory into the universal SMILES representation. Such conversion makes it easier to read molecules, easier for chemical researchers to understand, and easier to add, delete, modify and query the database. In addition, many chemical toolkits such as RDKit and Openbabel can easily operate molecules based on the SMILES format.

[0149] In the execution process of the format conversion module, first, all possible atoms in the database need to be defined and stored in a sequence. Secondly, all possible bond types are defined, such as single bonds, double bonds, triple bonds, aromatic bonds, etc. Then, the molecular graph is read, all atoms in the molecular graph are extracted in sequence, and the adjacent relationship between the atoms in the molecules is converted into an adjacency matrix. According to the numbering in the form of the graph, it can be directly corresponded to the previously predefined atoms and bonds, and the atomic information and bond information are stored in the form of an atom list and an adjacency matrix. Finally, the SMILES format is standardized using the toolkit to ensure that the molecular formula obtained in the SMILES format is unique and there is no ambiguity. This is a very important step in the construction of the database, which ensures the consistency and accuracy of the data.

[0150] The auxiliary operation module may include a degree judgment function, a file reading and writing function, a counter and a progress bar, etc.

[0151] The degree judgment function can be used to judge the degree of a specified atomic node, which is slightly different from the degree judgment in traditional graph theory. In a molecular graph, the edge represents a chemical bond, so there are different types of chemical bonds such as single bonds, double bonds and triple bonds. Therefore, the edge of the graph will contain bond information or attributes. Considering the valence rule, the maximum number of molecules that an atom can connect to is determined by the number of lone pairs of electrons it carries. Therefore, when calculating the degree of an atomic node, it is necessary to consider the weighted value of the edge attribute for calculation. In the process of building an electrolyte database, since hydrogen atoms will be filled in the last stage, the influence of hydrogen atoms can be ignored. For example, a carbon atom has formed a double bond and a single bond with other atoms, so the degree of the carbon atom is 3, that is, it can also be connected to an atom in the form of a single bond. In this way, when making a degree judgment, only the bonding of other atoms needs to be considered, without taking into account the influence of hydrogen atoms, thereby simplifying the calculation process.

[0152] File read and write functions can be used to handle situations where the initial set of molecules is large or the number of atoms contained in the generated molecules is large. In this case, relying solely on memory algorithms may lead to memory overflow problems, and the generated molecules need to be placed in external storage for auxiliary storage. It should be noted that because the input / output (I / O) operation will consume a lot of time when the program calls data in external storage, if the number of molecules generated by the program can be stored in memory, the memory algorithm should be given priority instead of using file read and write functions. This method can improve the efficiency of program operation while ensuring data integrity.

[0153] The counter can be used to record the number of molecules generated. When the program is finished, the number of molecules generated in this round needs to be given. The program can hierarchically count the number of molecules with a specified number of atoms and the total number of molecules generated. This statistics not only helps to understand the scale of the generated problem, but also predicts the running time of the next round of the program based on the generated molecular data. In this way, computing resources can be planned more effectively and program running efficiency can be optimized.

[0154] The progress bar can be used to provide visual feedback when the program generates molecules with many atoms. If the program has no screen output for a long time, it is difficult to determine whether the program is running normally or an exception has occurred and caused it to stagnate. Using a progress bar can help visualize the running progress of the program, understand the execution status of the program in real time, and make a rough prediction of the program running time. This not only improves the user experience, but also can promptly discover and solve possible problems to ensure the smooth operation of the program.

[0155] Next, the molecular structure data generation method based on the serial molecular generation algorithm is introduced. In one embodiment, the preset molecular generation rule includes executing the following steps for each initial molecule:

[0156] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of first molecules.

[0157] For each first molecule among the plurality of first molecules, a preset type of chemical bond is added to the first molecule according to a chemical bond addition rule, a valence rule, and a molecular stability rule to obtain a plurality of second molecules.

[0158] The plurality of first molecules and the plurality of second molecules are taken as a plurality of primary target molecules and stored in a preset target molecule set.

[0159] For each of the first-level target molecules in the plurality of first-level target molecules, a preset number of preset atoms are added to the first-level target molecules according to the atom addition rule, the valence rule, and the molecular stability rule, to obtain a plurality of third molecules. For each of the third molecules in the plurality of third molecules, a preset type of chemical bond is added to the third molecule according to the chemical bond addition rule, the valence rule, and the molecular stability rule, to obtain a plurality of fourth molecules. The plurality of third molecules and the plurality of fourth molecules are stored as a plurality of second-level target molecules in a preset target molecule set.

[0160] Taking the multiple secondary target molecules as new multiple primary target molecules, returning to each of the multiple primary target molecules, respectively adding a preset number of preset atoms to the primary target molecule according to the atom addition rule, valence rule and molecular stability rule, to obtain multiple third molecules until the preset iteration stop condition is met.

[0161] In the above embodiment, the iteration stop condition may include that the number of atoms in the generated secondary target molecules reaches a preset number, or the number of iterations reaches a preset number, or the number of target molecules in the target molecule set reaches a preset value, etc. Those skilled in the art may select a suitable iteration stop condition according to actual needs, which is not limited here.

[0162] According to the above implementation, the molecule can be amplified by atom, and then the chemical bond can be further amplified based on the result of atom amplification, and then the target molecule of the current round can be obtained based on the result of atom amplification and the result of chemical bond amplification. In this way, after multiple rounds of iterations, a large amount of target molecule structure data can be efficiently generated.

[0163] In one embodiment, the molecular stability rules include molecular structure graphs that do not comply with the molecular stability rules.

[0164] According to the atomic addition rule, the valence rule and the molecular stability rule, adding a preset number of preset atoms to the initial molecule to obtain a plurality of first molecules may specifically include:

[0165] According to the atomic addition rule and the valence rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of initial first molecules.

[0166] Finding a molecular structure diagram among a plurality of initial first molecules.

[0167] The initial first molecule including the molecular structure graph is deleted to obtain a plurality of first molecules.

[0168] According to the above implementation, the initial first molecule containing the unstable molecular structure can be determined by subgraph isomorphism based on the graph search method, and then the initial first molecule containing the unstable molecular structure can be deleted. In this way, the unstable molecules generated during the molecular amplification process can be deleted immediately, thereby improving the quality of the target molecular structure data.

[0169] In one embodiment, the initial molecule includes at least one non-hydrogen atom.

[0170] According to the atomic addition rule and the valence rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of initial first molecules, which may specifically include performing the following steps on the initial molecule for each preset atom:

[0171] Get the number of preset atoms in the initial molecule.

[0172] When the number of preset atoms is less than the preset threshold, for each non-hydrogen atom in the initial molecule, the following steps are performed respectively: the atomic degree of the non-hydrogen atom is calculated, and the atomic degree includes the sum of the weighted values ​​of the edges connected to the corresponding nodes of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attributes, and the edge attributes are used to represent the attributes of the chemical bonds corresponding to the edges. According to the correspondence between the preset atomic degree and the chemical bond type, the first chemical bond type corresponding to the atomic degree of the non-hydrogen atom is determined. The non-hydrogen atom is connected to the preset atom according to the first chemical bond type to obtain the initial first molecule.

[0173] In the above embodiment, the correspondence between the preset atomic degree and the chemical bond type can be set according to actual needs, and can be specifically set in combination with actual needs and valence rules. Exemplarily, when the non-hydrogen atom is a C atom and the preset atoms are all C atoms, if the atomic degree of the C atom satisfies less than 4, then the corresponding first chemical bond type can be a single bond; if the atomic degree of the C atom satisfies less than 3, then the corresponding first chemical bond type can be a double bond; if the atomic degree of the C atom satisfies less than 2, then the corresponding first chemical bond type can be a triple bond. For example, when the atomic degree of the C atom is 2, the corresponding first chemical bond type can be a single bond or a double bond, then a new C atom can be connected to the C atom through a single bond and a double bond, respectively, to obtain two initial first molecules.

[0174] According to the above embodiment, a suitable bonding method can be selected according to the number of non-hydrogen atoms in the molecule and the correspondence between the preset number of atoms and the type of chemical bonds, which is conducive to improving the chemical rationality of the initial first molecule.

[0175] In one embodiment, after the non-hydrogen atom and the preset atom are connected according to the first chemical bond type to obtain an initial first molecule, the method may further include:

[0176] The graph isomorphism between the initial first molecule and other initial first molecules is judged at the molecular graph level.

[0177] If there are other initial first molecules that are isomorphic to the initial first molecule graph, the initial first molecule is deleted.

[0178] According to the above embodiment, after the initial first molecule is generated, the graph isomorphism judgment can be performed in time, so as to delete the repeatedly generated molecules. In this way, the uniqueness of the generated initial first molecule can be guaranteed, which is conducive to maintaining the simplicity and uniqueness of the molecular structure data.

[0179] In one embodiment, for each first molecule in the plurality of first molecules, according to the chemical bond addition rule, the valence rule, and the molecular stability rule, a preset type of chemical bond is added to the first molecule to obtain the plurality of second molecules, which may specifically include:

[0180] For each first molecule among the plurality of first molecules, a preset type of chemical bond is added to the first molecule according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules.

[0181] Finding a molecular structure diagram among a plurality of initial second molecules.

[0182] The initial second molecule including the molecular structure graph is deleted to obtain a plurality of second molecules.

[0183] According to the above implementation, the initial second molecule containing the unstable molecular structure can be determined by subgraph isomorphism based on the graph search method, and then the initial second molecule containing the unstable molecular structure can be deleted. In this way, the unstable molecules generated during the molecular amplification process can be deleted immediately, thereby improving the quality of the target molecular structure data.

[0184] In one embodiment, the first molecule includes at least two non-hydrogen atoms.

[0185] For each first molecule among the plurality of first molecules, adding a preset type of chemical bond in the first molecule according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules may specifically include performing the following steps for each first molecule among the plurality of first molecules:

[0186] For any two non-hydrogen atoms in the first molecule, the target atomic degrees and target edge attributes of the two non-hydrogen atoms are obtained respectively.

[0187] According to the preset correspondence between the atomic degrees of the two non-hydrogen atoms and the edge attributes and the chemical bond type, the second chemical bond type corresponding to the target atomic degrees and the target edge attributes of the two non-hydrogen atoms is determined.

[0188] The two non-hydrogen atoms are connected according to the second chemical bond type to obtain an initial second molecule.

[0189] In the above embodiment, the atomic degrees of the two non-hydrogen atoms and the corresponding relationship between the edge attributes and the chemical bond type can be set according to actual needs, and can be set in combination with actual needs and valence rules. Exemplarily, when the two non-hydrogen atoms in the first molecule are C atoms-C atoms, if there is no chemical bond between the two C atoms, and both C atoms can be connected to the new atom, that is, their valence bonds are not saturated, then a new chemical bond can be added between the two C atoms. If the atomic degrees of the two C atoms are less than 4, then the corresponding second chemical bond type can be a single bond; if the atomic degrees of the two C atoms are both less than 3, and each C atom itself has no other double bonds connected, then the corresponding second chemical bond type can be a double bond; if the atomic degrees of the two C atoms are both less than 2, then the corresponding second chemical bond type can be a triple bond. For example, when the atomic degrees of the two C atoms are both 1, the corresponding second chemical bond type can be a single bond, a double bond or a triple bond, then the two C atoms can be connected by a single bond, a double bond and a triple bond, respectively, to obtain three initial first molecules. The process of determining the type of second chemical bond that can be added between other atoms is similar to that of C atom-C atom and will not be elaborated here.

[0190] According to the above embodiment, a suitable bonding method can be selected according to the number of non-hydrogen atoms in the two molecules and the correspondence between the preset number of atoms and the chemical bond type, which is conducive to improving the chemical rationality of the initial second molecule.

[0191] In one embodiment, after the two non-hydrogen atoms are connected according to the second chemical bond type to obtain an initial second molecule, the method may further include:

[0192] The graph isomorphism between the initial second molecule and the multiple initial first molecules and other initial second molecules is judged at the molecular graph level.

[0193] When there is an initial first molecule or other initial second molecules that are graph-isomorphic to the initial second molecule, the initial second molecule is deleted.

[0194] According to the above embodiment, after the initial second molecule is generated, the graph isomorphism judgment can be performed in time, so as to delete the repeatedly generated molecules. In this way, the uniqueness of the generated initial second molecule can be guaranteed, which is conducive to maintaining the simplicity and uniqueness of the molecular structure data.

[0195] In one embodiment, the molecular structure data can be generated based on a molecular incremental generation algorithm. This solution can be implemented based on a molecular incremental generation algorithm system. For ease of understanding, the molecular incremental generation algorithm system is briefly introduced below.

[0196] The molecular incremental generation algorithm system may include an initial molecular generation module, a graph constraint module, an atomic augmentation module, a bond augmentation module, a graph isomorphism judgment module, a format conversion module and an auxiliary operation module. In the molecular incremental generation algorithm system, the functions of the initial molecular generation module, the graph constraint module, the atomic augmentation module, the bond augmentation module, the format conversion module and the auxiliary operation module are similar to the functions of similar modules in the serial generation algorithm, the only difference is that the atomic augmentation module and the bond augmentation module no longer perform the graph isomorphism judgment action when performing molecular amplification, and the graph isomorphism judgment module performs the action; and the input of the bond augmentation module is changed.

[0197] Specifically, the bond augmentation module generates the molecular graph of the current round based on the graph with the minimum number of edges generated in the previous round. The number of edges here is different from the weighted number of bonds. The bond is an attribute attached to the edge, that is, an edge can be a single bond, a double bond, or a triple bond, but its number of edges is only 1. In other words, the number of edges here is the node degree. The rationality and completeness of this generation method lies in that according to the atomic augmentation and molecular augmentation strategies proposed in this algorithm, any molecular graph can be generated from its subgraph. Therefore, when constructing an organic small molecule database, it is only necessary to find the minimum subgraph set of the pre-generated molecular graph, that is, a set of graphs with the minimum number of edges in the molecular graph generated in the previous round. In this way, the probability of generating graph isomorphic molecules during bond augmentation can be reduced, thereby greatly saving the time for graph isomorphism judgment.

[0198] The molecule incremental generation algorithm may include a step-by-step merging-based incremental generation algorithm and a multi-grouping-based incremental generation algorithm.

[0199] The biggest difference between the incremental generation algorithm based on step-by-step merging and the serial molecular graph generation algorithm is that the graph isomorphism judgment module is abstracted, and the isomorphism judgment process originally coupled in the atomic augmentation and bond augmentation modules is replaced by a step-by-step merging parallel algorithm. This abstract process is conducive to enhancing code reusability and the division of module functions. At the same time, the decoupling of the atomic augmentation and bond augmentation modules is realized. According to the idea of ​​"high cohesion and low coupling", the coupling between modules is minimized and the connection within the module is increased, which is conducive to improving the efficiency of the algorithm.

[0200] In the incremental molecular generation algorithm system based on step-by-step merging, the parallelized atomic augmentation module performs calculations based on multiple cores. The module input is a set of molecular graphs with the least number of edges in the previous round. That is, if the number of nodes in the pre-generated molecular graph in this round is n, the molecular graph input in the atomic augmentation module needs to satisfy the requirement that its number of nodes is n-1 and the number of edges is n-2. The output of the atomic augmentation module is a molecular graph with one more point (atom) and one more edge (bond, which may be a single bond, double bond, or triple bond, etc.).

[0201] The parallelized bond augmentation module adds edges to the molecular graph after atomic augmentation in sequence. Different from the serialized algorithm, the parallelized bond augmentation module realizes the decoupling of atomic augmentation and bond augmentation. In the parallelized bond augmentation module, the input of the bond augmentation module comes from the entire structure of atomic augmentation, rather than being coupled in the process of atomic augmentation. The output of the bond augmentation module is a set of molecular graphs with increased number of edges.

[0202] In the incremental molecular generation algorithm system based on step-by-step merging, the graph isomorphism judgment module adopts the "step-by-step merging" method compared to the serial algorithm, which effectively shortens the time consumption caused by isomorphism judgment. The graph isomorphism module is immediately after the atomic augmentation module and the bond augmentation module. Since each process in the parallel algorithm is independent of each other, although the graphs generated in each process are not isomorphic, the graphs between processes may be isomorphic. Therefore, the input of the graph isomorphism judgment module is the molecular graph generated by the atomic augmentation module and the bond augmentation module, and the output is a molecular graph without isomorphic graphs.

[0203] When judging isomorphism, if the number of graphs in the input module is too small, less than twice the number of computer cores, there is no need to start the gradual merging algorithm. This is because starting the process also has a certain amount of time overhead. When the number of graphs is too small, it is faster to directly judge isomorphism, and there is no need for complex algorithms and parallel processing. At this time, you only need to aggregate the collection of graphs from the atomic augmentation or bond augmentation module and judge the isomorphism together. Since the number of graphs is small, this process can be completed relatively quickly.

[0204] In the incremental generation algorithm based on step-by-step merging, the time consumed by graph isomorphism judgment accounts for the vast majority of the total algorithm time. Therefore, an effective means to accelerate the algorithm is to reduce the number of graph isomorphism judgments or reduce unnecessary graphs. Generating graphs based on the minimum number of edges can effectively avoid the generation of redundant graphs due to the graph with a large number of edges in the previous round during the parallelization process, which requires graph isomorphism judgment to delete them, thereby saving the time cost required for program running.

[0205] The molecular incremental generation algorithm system based on multiple groups is based on the discovery that multiple groups can effectively reduce the time in the graph isomorphism process. It aims to use the idea of ​​multiple groups to reconstruct the graph isomorphism judgment module in the gradual merging algorithm, and replace the graph storage with serial numbers to reduce space complexity. The main ideas of other modules are basically the same.

[0206] In the execution process of the graph isomorphism judgment module of the molecular incremental generation algorithm system based on multiple groups, first, the graphs from atomic augmentation or bond augmentation need to be grouped in pairs. This process can be regarded as a merging rate of 2, but unlike the step-by-step merging algorithm, the multiple grouping algorithm needs to match all graph sets in pairs. For example, if 24 graph sets are generated after atomic augmentation or bond augmentation, they are numbered separately. Assuming that the supercomputer has 24 cores and 24 subprocesses are created, the number of groups required in the multiple grouping algorithm is (24×23) / 2=276 groups, and each group is marked by the serial numbers of the 24 graph sets it contains, with the smaller serial number in front and the larger serial number in the back. Afterwards, the isomorphism needs to be judged for each group of graphs separately. Assuming that the system can create up to 24 subprocesses at the same time, then each round can judge the isomorphism of up to 24 graph sets, so a total of 12 rounds are required. In each round of judgment, record the position sequence of isomorphic graphs in the sets with the earlier sequence numbers instead of directly copying the generated graphs, so as to achieve in-situ elimination of isomorphic graphs. After 12 rounds of judgment, take the intersection of all isomorphic position sequence groups with the same previous sequence number, and the result is the set of all graphs in each of the 24 graph sets that may be isomorphic with other sets. Keep the graphs in the sets with smaller sequence numbers and delete the graphs in the sets with larger sequence numbers. Finally, delete the graphs with the corresponding graph isomorphism position numbers in the 24 graph sets one by one, and what remains are all non-isomorphic molecular graphs. This method can effectively reduce the number of graph isomorphism judgments and unnecessary graphs, thereby improving the overall efficiency of the algorithm.

[0207] For the database of fluorine-containing small molecules in electrolytes, the rule is that the initial atom is an oxygen atom, that is, each molecule must contain an oxygen atom. During the construction process, the atoms that can be added in the atom augmentation stage include carbon, oxygen and fluorine atoms, and the bond augmentation stage can add carbon-carbon single bonds, carbon-oxygen single bonds, carbon-carbon double bonds, carbon-oxygen double bonds and carbon-carbon triple bonds. At the same time, the following conditions need to be met: oxygen does not form bonds with oxygen or fluorine (no peroxide bonds are produced), two adjacent bonds cannot be double bonds at the same time, the total number of carbon atoms in a single molecule does not exceed 8, and the total number of oxygen atoms does not exceed 4.

[0208] In the actual molecular generation process, it can be observed that the number of graphs increases exponentially, and the time consumed also shows an exponential upward trend. For example, for the generation of a graph with 7 nodes, the incremental generation algorithm based on step-by-step merging is 2 times shorter than the serial generation algorithm, and the incremental generation algorithm based on multiple groups is 7.6 times shorter. For the generation of a graph with 8 nodes, the incremental generation algorithm based on step-by-step merging takes a long time due to the presence of more I / O processes in the design process; comparing the serial generation algorithm with the incremental generation algorithm based on multiple groups, the latter's running time is shortened by 8.8 times. It can be seen that the incremental generation algorithm based on multiple groups can significantly accelerate the database generation process and save a lot of time. A total of 402,830 electrolyte organic small molecules with a maximum of 9 atoms were generated using this algorithm, meeting the demand for small molecules.

[0209] Next, the molecular structure data generation method based on the molecular incremental generation algorithm is introduced.

[0210] In one embodiment, the preset molecule generation rule includes executing the following steps for each initial molecule:

[0211] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of fifth molecules.

[0212] The plurality of fifth molecules are stored in a preset target molecule set.

[0213] A target fifth molecule is determined from the plurality of fifth molecules, where the target fifth molecule is the fifth molecule having the least number of edges among the plurality of fifth molecules.

[0214] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms is added to the target fifth molecule to obtain a plurality of sixth molecules.

[0215] The plurality of sixth molecules are stored in a preset target molecule set.

[0216] The plurality of sixth molecules are used as new plurality of fifth molecules, and the target fifth molecule is determined from the plurality of fifth molecules, until a preset iteration stop condition is satisfied.

[0217] According to the chemical bond addition rule, the valence rule and the molecular stability rule, chemical bonds of preset types are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of seventh molecules.

[0218] The plurality of seventh molecules are stored in a preset target molecule set.

[0219] According to the preset graph isomorphism merging rule, the molecules with graph isomorphism in the preset target molecule set are merged.

[0220] In the above embodiment, the target fifth molecule is the fifth molecule with the least number of edges among the multiple fifth molecules. It is understandable that there can be multiple target fifth molecules, and when there are multiple target fifth molecules, the corresponding steps in the above embodiment can be performed on each target fifth molecule.

[0221] In the above implementation, according to the preset graph isomorphism merging rule, the preset target molecule set is merged and there are molecules with graph isomorphism, and the same molecules can be merged, thereby there are no duplicate molecules in the target molecule set, and the molecules in the target molecule set are the target molecules.

[0222] In the above embodiment, the iteration stop condition may include that the number of atoms in the generated sixth molecule reaches a preset number, or the number of iterations reaches a preset number, or the number of molecules in the target molecule set reaches a preset value, etc. Those skilled in the art may select a suitable iteration stop condition according to actual needs, which is not limited here.

[0223] According to the above implementation, firstly, the atom amplification is performed based on the molecule with the smallest number of edges generated in the previous round, and then the bond amplification is performed. Generating a molecular graph based on the minimum number of edges can effectively reduce the risk of redundant graphs generated by the graph with a large number of edges in the previous round during the parallelization process, which requires the graph isomorphism judgment to delete them, thereby saving the time cost required for program operation. In this way, the efficiency of molecular structure data generation can be improved.

[0224] In one embodiment, the graph isomorphism merging rule includes a graph isomorphism merging rule based on a step-by-step merging algorithm. The graph isomorphism merging rule based on a step-by-step merging algorithm has been described above in conjunction with the system, and will not be described in detail here.

[0225] According to the above implementation, the efficiency of graph isomorphism judgment and merging can be improved, thereby improving the efficiency of molecular structure data generation.

[0226] In one embodiment, the graph isomorphism merging rule includes a graph isomorphism merging rule based on a multi-grouping algorithm. The graph isomorphism merging rule based on a multi-grouping algorithm has been described above in conjunction with the system, and will not be repeated here.

[0227] According to the above implementation, the efficiency of graph isomorphism judgment and merging can be improved, thereby improving the efficiency of molecular structure data generation.

[0228] In one embodiment, the molecular stability rules may include molecular structure graphs that do not comply with the molecular stability rules.

[0229] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of fifth molecules, which may specifically include:

[0230] According to the atomic addition rule and the valence rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of initial fifth molecules.

[0231] Find molecular structure diagrams among multiple initial fifth molecules.

[0232] The initial fifth molecule including the molecular structure diagram is deleted to obtain a plurality of fifth molecules.

[0233] According to the above implementation, the initial fifth molecule containing the unstable molecular structure can be determined by subgraph isomorphism based on the graph search method, and then the initial fifth molecule containing the unstable molecular structure can be deleted. In this way, the unstable molecules generated during the molecular amplification process can be deleted immediately, thereby improving the quality of the target molecular structure data.

[0234] In one embodiment, the initial molecule includes at least one non-hydrogen atom.

[0235] Adding a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain a plurality of initial fifth molecules may specifically include:

[0236] Get the number of preset atoms in the initial molecule.

[0237] When the number of preset atoms is less than the preset threshold, for each non-hydrogen atom in the initial molecule, the following steps are performed respectively: the atomic degree of the non-hydrogen atom is calculated, and the atomic degree includes the sum of the weighted values ​​of the edges connected to the corresponding nodes of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attributes, and the edge attributes are used to represent the attributes of the chemical bonds corresponding to the edges. According to the correspondence between the preset atomic degree and the chemical bond type, the third chemical bond type corresponding to the atomic degree of the non-hydrogen atom is determined. The non-hydrogen atom is connected to the preset atom according to the third chemical bond type to obtain the initial fifth molecule.

[0238] In the above embodiment, the method for determining the third chemical bond type is similar to the method for determining the first to third chemical bond types, which will not be elaborated herein.

[0239] According to the above embodiment, a suitable bonding method can be selected according to the number of non-hydrogen atoms in the molecule and the correspondence between the preset number of atoms and the type of chemical bonds, which is conducive to improving the chemical rationality of the initial fifth molecule.

[0240] In one embodiment, according to the chemical bond addition rule, the valence rule, and the molecular stability rule, adding a preset type of chemical bond to a plurality of fifth molecules and a plurality of sixth molecules in the preset target molecule set to obtain a plurality of seventh molecules may specifically include:

[0241] According to the chemical bond addition rule and the valence rule, chemical bonds of preset types are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of initial seventh molecules.

[0242] Find molecular structure diagrams among multiple initial seventh molecules.

[0243] The initial seventh molecule including the molecular structure diagram is deleted to obtain a plurality of seventh molecules.

[0244] According to the above implementation, the initial seventh molecule containing the unstable molecular structure can be determined by subgraph isomorphism based on graph search, and then the initial seventh molecule containing the unstable molecular structure can be deleted. In this way, the unstable molecules generated during the molecular amplification process can be deleted immediately, thereby improving the quality of the target molecular structure data.

[0245] In one embodiment, the fifth molecule and the sixth molecule each include at least two non-hydrogen atoms.

[0246] According to the chemical bond addition rule and the valence rule, chemical bonds of preset types are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain the plurality of initial seventh molecules. Specifically, the method may include performing the following steps for each of the plurality of fifth molecules and the plurality of sixth molecules:

[0247] For any two non-hydrogen atoms in a molecule, obtain the target atomic degrees and target edge attributes of the two non-hydrogen atoms respectively.

[0248] According to the preset correspondence between the atomic degrees of the two non-hydrogen atoms and the edge attributes and the chemical bond type, the fourth chemical bond type corresponding to the target atomic degrees and the target edge attributes of the two non-hydrogen atoms is determined.

[0249] Two non-hydrogen atoms are connected according to the fourth chemical bond type to obtain a plurality of initial seventh molecules.

[0250] In the above embodiment, the method for determining the fourth chemical bond type is similar to the method for determining the second chemical bond type, which will not be described in detail herein.

[0251] According to the above embodiment, a suitable bonding method can be selected according to the correspondence between the non-hydrogen atomic number in the two molecules and the preset atomic number and chemical bond type, which is conducive to improving the chemical rationality of the initial seventh molecule.

[0252] In one embodiment, the molecular stability rules include molecular structure graphs that do not comply with the molecular stability rules.

[0253] Before adding preset atoms and / or preset chemical bonds to the initial molecule according to the preset molecule generation rule, the method may further include:

[0254] An application environment for obtaining structural data of multiple target molecules.

[0255] According to the chemical stability rules and the application environment, the molecular structure diagram that does not meet the molecular stability in the application environment is determined.

[0256] In the above embodiments, the application environment may include an environment corresponding to the research or application system. Taking electrolyte research as an example, the application environment may include the type of battery system and the type of electrolyte system.

[0257] According to the above implementation, by constructing molecular stability rules through chemical stability rules and application environment, it is helpful to screen out molecules that are unstable in the application environment, thereby retaining target molecules with higher stability. In this way, the target molecular structure data can meet the needs of practical applications, thereby helping to improve the flexibility of molecular structure data generation and the quality of molecular structure data.

[0258] In one embodiment, the preset atoms may include at least one of carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl). In this way, the needs of most electrolyte systems can be met.

[0259] In one embodiment, converting multiple target molecules into SMILES format to obtain multiple target molecule structure data may specifically include:

[0260] Convert multiple target molecules into adjacency matrix or adjacency table format to obtain intermediate data of target molecule structure.

[0261] The target molecular structure intermediate data is converted into SMILES format to obtain multiple target molecular structure data.

[0262] According to the above embodiment, the multiple target molecules are converted into an adjacency matrix or adjacency table format, which is beneficial for computer storage and convenient for format conversion. In this way, the intermediate data of the target molecule structure can be efficiently converted into the SMILES format.

[0263] In one embodiment, after converting the plurality of target molecules into SMILES format to obtain the plurality of target molecule structure data, the method may further include:

[0264] Based on multiple target molecular structure data, a target molecular structure database is created.

[0265] According to the above implementation, a target molecular structure database is created based on multiple target molecular structure data. This database, when combined with high-throughput computing, high-throughput screening and machine learning technologies, is conducive to the high-speed research and development of electrolytes.

[0266] Based on the same inventive concept as the molecular structure data generating method, the embodiment of the present application also provides a molecular structure data generating device.

[0267] like Figure 3 As shown, the molecular structure data generating device 200 may include a first acquisition module 201 , an adding module 202 and a conversion module 203 .

[0268] The first acquisition module 201 is used to acquire an initial molecule, where the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds.

[0269] The adding module 202 is used to add preset atoms and / or preset chemical bonds to the initial molecules according to preset molecular generation rules to obtain multiple target molecules, wherein the molecular generation rules include atom addition rules, chemical bond addition rules, valence rules and molecular stability rules, and the target molecules are represented by an undirected graph without self-loops.

[0270] The conversion module 203 is used to convert the multiple target molecules into SMILES format to obtain multiple target molecule structure data.

[0271] The molecular structure data generating device of the embodiment of the present application can obtain the initial molecule, and the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds. In this way, by using the graph theory method to represent the molecule, it is possible to ensure the structuring and standardization of the molecular data in the process of molecular generation, thereby improving the efficiency of data processing. Then, according to the preset molecular generation rules, preset atoms and / or preset chemical bonds can be added to the initial molecule to obtain multiple target molecules, wherein the molecular generation rules include atomic increase rules, chemical bond increase rules, valence rules and molecular stability rules, and the target molecule is represented by an undirected graph without self-loops. On the one hand, the molecular generation rules can make the structure of the generated target molecule have chemical rationality and stability, reducing the risk of generating unreasonable molecules; on the other hand, by specifying atomic increase rules and chemical bond increase rules, batch generation of target molecules that meet the needs of users can be performed, thereby improving the efficiency and flexibility of molecular generation. In this way, the generated multiple target molecules can adapt to the needs of different research scenarios. Subsequently, multiple target molecules can also be converted into SMILES format to obtain multiple target molecular structure data. SMILES format is a more standardized molecular representation, which is convenient for retrieval and storage, and there is no ambiguity. By converting multiple target molecules into SMILES format, the consistency and accuracy of the target molecule structure data can be guaranteed. In this way, a solid data foundation can be provided for high-throughput computing, machine learning and high-throughput screening. The target molecule structure data generated in batches in the application embodiment are used to construct an electrolyte molecular structure database, which can effectively support the rapid development of new electrolyte systems, significantly shorten the development cycle, reduce research and development costs, and improve the efficiency of electrolyte design and optimization.

[0272] In one embodiment, the apparatus may further include a first molecule generation rule module, and the first molecule generation rule module may include:

[0273] The first adding submodule is used to add a preset number of preset atoms to each initial molecule according to the atomic adding rule, the valence rule and the molecular stability rule to obtain a plurality of first molecules.

[0274] The first adding submodule is also used to add a preset type of chemical bond to each first molecule in the multiple first molecules according to the chemical bond adding rule, the valence rule and the molecular stability rule to obtain multiple second molecules.

[0275] The first storage submodule is used to store the plurality of first molecules and the plurality of second molecules as a plurality of primary target molecules into a preset target molecule set.

[0276] The processing submodule is used for adding a preset number of preset atoms to each of the first-level target molecules in accordance with the atom addition rule, the valence rule and the molecular stability rule to obtain a plurality of third molecules. For each of the third molecules in the plurality of third molecules, adding a preset type of chemical bond to the third molecule in accordance with the chemical bond addition rule, the valence rule and the molecular stability rule to obtain a plurality of fourth molecules. The plurality of third molecules and the plurality of fourth molecules are used as a plurality of second-level target molecules and stored in a preset target molecule set.

[0277] The first return submodule is used to take the multiple secondary target molecules as new multiple primary target molecules, return each of the multiple primary target molecules, and add a preset number of preset atoms to the primary target molecule according to the atom addition rule, valence rule and molecular stability rule to obtain multiple third molecules until the preset iteration stop condition is met.

[0278] In one embodiment, the molecular stability rules may include molecular structure graphs that do not comply with the molecular stability rules.

[0279] The first adding submodule is used to add a preset number of preset atoms to the initial molecule according to the atomic adding rule, the valence rule and the molecular stability rule to obtain a plurality of first molecules, which may specifically include:

[0280] The first adding unit is used to add a preset number of preset atoms to the initial molecule according to the atomic adding rule and the valence rule to obtain a plurality of initial first molecules.

[0281] The first searching unit is used to search for the molecular structure diagram in a plurality of initial first molecules.

[0282] The first deleting unit is used to delete the initial first molecule including the molecular structure diagram to obtain multiple first molecules.

[0283] In one embodiment, the initial molecule may include at least one non-hydrogen atom.

[0284] The first adding unit is used to add a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain a plurality of initial first molecules, which may specifically include performing the following steps on the initial molecule for each preset atom through each subunit in the first adding unit:

[0285] The first acquisition subunit is used to acquire the number of preset atoms in the initial molecule.

[0286] The first execution subunit is used to perform the following steps for each non-hydrogen atom in the initial molecule when the number of preset atoms is less than the preset threshold: calculate the atomic degree of the non-hydrogen atom, the atomic degree includes the sum of the weighted values ​​of each edge connected to the corresponding node of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attribute, and the edge attribute is used to represent the attribute of the chemical bond corresponding to the edge. According to the correspondence between the preset atomic degree and the chemical bond type, determine the first chemical bond type corresponding to the atomic degree of the non-hydrogen atom. Connect the non-hydrogen atom and the preset atom according to the first chemical bond type to obtain the initial first molecule.

[0287] In one embodiment, the apparatus may further include:

[0288] The first judgment module is used to perform graph isomorphism judgment on the initial first molecule and other initial first molecules at the molecular graph level after the initial first molecule is obtained by connecting non-hydrogen atoms and preset atoms according to the first chemical bond type.

[0289] The first deleting module is used to delete the initial first molecule when there are other initial first molecules that are isomorphic to the initial first molecule graph.

[0290] In one embodiment, the first adding submodule is used to add a preset type of chemical bond to each first molecule in the plurality of first molecules according to the chemical bond adding rule, the valence rule and the molecular stability rule to obtain the plurality of second molecules, which may specifically include:

[0291] The second adding unit is used to add a preset type of chemical bond in each first molecule of the plurality of first molecules according to a chemical bond adding rule and a valence rule to obtain a plurality of initial second molecules.

[0292] The second searching unit is used to search for the molecular structure diagram in a plurality of initial second molecules.

[0293] The second deleting unit is used to delete the initial second molecule including the molecular structure diagram to obtain a plurality of second molecules.

[0294] In one embodiment, the first molecule may include at least two non-hydrogen atoms.

[0295] The second adding unit is used for adding a preset type of chemical bond in each of the first molecules according to the chemical bond adding rule and the valence rule to obtain a plurality of initial second molecules. Specifically, the second adding unit may include performing the following steps for each of the first molecules in the plurality of first molecules through each subunit in the second adding unit:

[0296] The second acquisition subunit is used to respectively acquire the target atomic degree and target edge attribute of each of any two non-hydrogen atoms in the first molecule.

[0297] The first determination subunit is used to determine the second chemical bond type corresponding to the target atomic degree and target edge attribute of each of the two non-hydrogen atoms according to the preset correspondence between the atomic degree and edge attribute of each of the two non-hydrogen atoms and the chemical bond type.

[0298] The first linker unit is used to connect two non-hydrogen atoms according to a second chemical bond type to obtain an initial second molecule.

[0299] In one embodiment, the apparatus may further include:

[0300] The second judgment module is used to perform graph isomorphism judgment on the initial second molecule and multiple initial first molecules and other initial second molecules at the molecular graph level after connecting two non-hydrogen atoms according to the second chemical bond type to obtain the initial second molecule.

[0301] The second deleting module is used to delete the initial second molecule when there is an initial first molecule or other initial second molecules that are isomorphic to the initial second molecule graph.

[0302] In one embodiment, the apparatus may further include a second molecule generation rule module, and the second molecule generation rule module may include:

[0303] The second adding submodule is used to add a preset number of preset atoms to the initial molecule according to the atomic adding rule, the valence rule and the molecular stability rule to obtain a plurality of fifth molecules.

[0304] The second storage submodule is used to store the plurality of fifth molecules into a preset target molecule set.

[0305] The determination submodule is used to determine a target fifth molecule from multiple fifth molecules, where the target fifth molecule is the fifth molecule with the least number of edges among the multiple fifth molecules.

[0306] The second adding submodule is further used to add a preset number of preset atoms to the target fifth molecule according to the atomic adding rule, the valence rule and the molecular stability rule to obtain a plurality of sixth molecules.

[0307] The second storage submodule is further used to store the plurality of sixth molecules into a preset target molecule set.

[0308] The second returning submodule is used to take the plurality of sixth molecules as new plurality of fifth molecules and return to determine a target fifth molecule from the plurality of fifth molecules until a preset iteration stop condition is met.

[0309] The second adding submodule is also used to add preset types of chemical bonds to multiple fifth molecules and multiple sixth molecules in the preset target molecule set according to chemical bond addition rules, valence rules and molecular stability rules to obtain multiple seventh molecules.

[0310] The second storage submodule is further used to store the plurality of seventh molecules into a preset target molecule set.

[0311] The merging submodule is used to merge the molecules with graph isomorphism in the preset target molecule set according to the preset graph isomorphism merging rules.

[0312] In one embodiment, the graph isomorphism merging rule may include a graph isomorphism merging rule based on a step-by-step merging algorithm.

[0313] In one embodiment, the graph isomorphism merging rule may include a graph isomorphism merging rule based on a multi-grouping algorithm.

[0314] In one embodiment, the molecular stability rules may include molecular structure graphs that do not comply with the molecular stability rules.

[0315] The second adding submodule is used to add a preset number of preset atoms to the initial molecule according to the atomic adding rule, the valence rule and the molecular stability rule to obtain a plurality of fifth molecules, which may specifically include:

[0316] The third adding unit is used to add a preset number of preset atoms to the initial molecule according to the atomic adding rule and the valence rule to obtain a plurality of initial fifth molecules.

[0317] The third searching unit is used to search for the molecular structure diagram in a plurality of initial fifth molecules.

[0318] The third deleting unit is used to delete the initial fifth molecule including the molecular structure diagram to obtain multiple fifth molecules.

[0319] In one embodiment, the initial molecule may include at least one non-hydrogen atom.

[0320] The third adding unit is used to add a preset number of preset atoms to the initial molecule according to the atomic adding rule and the valence rule to obtain a plurality of initial fifth molecules, which may specifically include:

[0321] The third acquisition subunit is used to acquire the number of preset atoms in the initial molecule.

[0322] The second execution subunit is used to perform the following steps for each non-hydrogen atom in the initial molecule when the number of preset atoms is less than the preset threshold: calculate the atomic degree of the non-hydrogen atom, the atomic degree includes the sum of the weighted values ​​of each edge connected to the corresponding node of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attribute, and the edge attribute is used to represent the attribute of the chemical bond corresponding to the edge. According to the correspondence between the preset atomic degree and the chemical bond type, determine the third chemical bond type corresponding to the atomic degree of the non-hydrogen atom. Connect the non-hydrogen atom and the preset atom according to the third chemical bond type to obtain the initial fifth molecule.

[0323] In one embodiment, the second adding submodule is used to add a preset type of chemical bond to a plurality of fifth molecules and a plurality of sixth molecules in the preset target molecule set according to the chemical bond adding rule, the valence rule and the molecular stability rule, respectively, to obtain a plurality of seventh molecules, which may specifically include:

[0324] The fourth adding unit is used to add preset types of chemical bonds to multiple fifth molecules and multiple sixth molecules in the preset target molecule set according to the chemical bond adding rule and the valence rule, so as to obtain multiple initial seventh molecules.

[0325] The fourth searching unit is used to search for the molecular structure diagram in a plurality of initial seventh molecules.

[0326] The fourth deleting unit is used to delete the initial seventh molecule including the molecular structure diagram to obtain multiple sixth molecules.

[0327] In one embodiment, the fifth molecule and the sixth molecule may each include at least two non-hydrogen atoms.

[0328] The fourth adding unit is used to add a preset type of chemical bond to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set according to the chemical bond adding rule and the valence rule, respectively, to obtain a plurality of initial seventh molecules, and specifically may include calling the following subunits for processing for each molecule in the plurality of fifth molecules and the plurality of sixth molecules:

[0329] The fourth acquisition subunit is used to obtain the target atomic degree and target edge attribute of each of the two non-hydrogen atoms in the molecule.

[0330] The second determination subunit is used to determine the fourth chemical bond type corresponding to the target atomic degree and target edge attribute of each of the two non-hydrogen atoms according to the preset correspondence between the atomic degree and edge attribute of each of the two non-hydrogen atoms and the chemical bond type.

[0331] The second linker unit is used to connect two non-hydrogen atoms according to the fourth chemical bond type to obtain a plurality of initial seventh molecules.

[0332] In one embodiment, the molecular stability rules may include molecular structure graphs that do not comply with the molecular stability rules.

[0333] The device may also include:

[0334] The second acquisition module is used to acquire the application environment of multiple target molecular structure data before adding preset atoms and / or preset chemical bonds to the initial molecule according to the preset molecule generation rule.

[0335] The determination module is used to determine the molecular structure diagram that does not meet the molecular stability in the application environment according to the chemical stability rule and the application environment.

[0336] In one embodiment, the predetermined atoms may include at least one of carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl).

[0337] In one embodiment, the conversion module is used to convert multiple target molecules into SMILES format to obtain multiple target molecule structure data, which may specifically include:

[0338] The conversion submodule is used to convert multiple target molecules into an adjacency matrix or adjacency table format to obtain intermediate data of the target molecule structure.

[0339] The conversion submodule is also used to convert the intermediate data of the target molecular structure into the SMILES format to obtain multiple target molecular structure data.

[0340] In one embodiment, the apparatus may further include:

[0341] The creation module is used to create a target molecule structure database based on the multiple target molecule structure data after converting the multiple target molecules into SMILES format to obtain the multiple target molecule structure data.

[0342] The molecular structure data generating device provided in the embodiment of the present application can achieve Figure 1 To avoid repetition, the various processes implemented by the method embodiment are not described here.

[0343] Figure 3 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.

[0344] The electronic device may include a processor 301 and a memory 302 storing computer program instructions.

[0345] Specifically, the processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0346] The memory 302 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 302 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid-state memory.

[0347] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0348] The processor 301 reads and executes the computer program instructions stored in the memory 302 to implement any one of the molecular structure data generation methods in the above embodiments.

[0349] As an example, the electronic device may further include a communication interface 303 and a bus 310. Figure 3 As shown, the processor 301, the memory 302, and the communication interface 303 are connected via a bus 310 and communicate with each other.

[0350] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0351] Bus 310 includes hardware, software or both, and the parts of molecular structure data generation equipment are coupled to each other.For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransmission (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In suitable cases, bus 310 may include one or more buses. Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.

[0352] The electronic device can execute the molecular structure data generation method in the embodiment of the present application, thereby realizing the combination Figure 1 and Figure 2 The described molecular structure data generation method and device.

[0353] In addition, in combination with the molecular structure data generation method in the above embodiment, the embodiment of the present application can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any one of the molecular structure data generation methods in the above embodiment is implemented.

[0354] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed, implements any one of the methods for generating molecular structure data in the above embodiments.

[0355] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.

[0356] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0357] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.

[0358] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0359] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.

Claims

1. A method for generating molecular structure data, characterized in that: include: Acquire an initial molecule, wherein the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds; According to a preset molecule generation rule, preset atoms and / or preset chemical bonds are added to the initial molecule to obtain a plurality of target molecules, wherein the molecule generation rule includes an atom addition rule, a chemical bond addition rule, a valence rule, and a molecule stability rule, and the target molecule is represented by an undirected graph without self-loops; The multiple target molecules are converted into a simplified molecular linear input specification SMILES format to obtain multiple target molecular structure data.

2. The method according to claim 1, characterized in that The preset molecule generation rule includes executing the following steps for each initial molecule: According to the atomic addition rule, the valence rule and the molecular stability rule, adding a preset number of preset atoms to the initial molecule to obtain a plurality of first molecules; For each first molecule of the plurality of first molecules, according to a chemical bond addition rule, a valence rule, and a molecular stability rule, a preset type of chemical bond is added to the first molecule to obtain a plurality of second molecules; The plurality of first molecules and the plurality of second molecules are stored as a plurality of primary target molecules in a preset target molecule set; For each of the plurality of primary target molecules, a preset number of preset atoms are added to the primary target molecule according to the atom addition rule, the valence rule and the molecular stability rule to obtain a plurality of third molecules; for each of the plurality of third molecules, a preset type of chemical bond is added to the third molecule according to the chemical bond addition rule, the valence rule and the molecular stability rule to obtain a plurality of fourth molecules; the plurality of third molecules and the plurality of fourth molecules are taken as a plurality of secondary target molecules and stored in a preset target molecule set; The multiple secondary target molecules are taken as new multiple primary target molecules, and for each primary target molecule in the multiple primary target molecules, a preset number of preset atoms are added to the primary target molecule according to the atom addition rule, the valence rule and the molecular stability rule to obtain multiple third molecules until the preset iteration stop condition is met.

3. The method according to claim 2, characterized in that The molecular stability rules include molecular structure diagrams that do not comply with molecular stability; The step of adding a preset number of preset atoms to the initial molecule according to the atomic addition rule, the valence rule and the molecular stability rule to obtain a plurality of first molecules includes: Adding a preset number of preset atoms to the initial molecule according to an atom addition rule and a valence rule to obtain a plurality of initial first molecules; searching for the molecular structure diagram in the plurality of initial first molecules; The initial first molecule including the molecular structure diagram is deleted to obtain the plurality of first molecules.

4. The method according to claim 3, characterized in that The initial molecule includes at least one non-hydrogen atom; The step of adding a preset number of preset atoms to the initial molecule according to the atom addition rule and the valence rule to obtain a plurality of initial first molecules includes performing the following steps on the initial molecule for each preset atom: Obtaining the number of the preset atoms in the initial molecule; In the case where the number of the preset atoms is less than the preset threshold, for each non-hydrogen atom in the initial molecule, the following steps are respectively performed: calculating the atomic degree of the non-hydrogen atom, the atomic degree comprising the sum of the weighted values ​​of the edges connected to the corresponding nodes of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attributes, the edge attributes being used to represent the attributes of the chemical bond corresponding to the edge; determining the first chemical bond type corresponding to the atomic degree of the non-hydrogen atom according to the correspondence between the preset atomic degree and the chemical bond type; The non-hydrogen atom and the preset atom are connected according to the first chemical bond type to obtain an initial first molecule.

5. The method according to claim 4, characterized in that After the non-hydrogen atom and the preset atom are connected according to the first chemical bond type to obtain an initial first molecule, the method further includes: Performing graph isomorphism judgment on the initial first molecule and other initial first molecules at the molecular graph level; If there are other initial first molecules that are isomorphic to the initial first molecule graph, the initial first molecule is deleted.

6. The method according to claim 3, characterized in that The step of adding a preset type of chemical bond to each of the plurality of first molecules according to a chemical bond addition rule, a valence rule, and a molecular stability rule to obtain a plurality of second molecules includes: For each first molecule of the plurality of first molecules, adding a preset type of chemical bond in the first molecule according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules; searching for the molecular structure diagram in the plurality of initial second molecules; The initial second molecule including the molecular structure diagram is deleted to obtain the plurality of second molecules.

7. The method according to claim 6, characterized in that The first molecule includes at least two non-hydrogen atoms; For each first molecule of the plurality of first molecules, adding a preset type of chemical bond in the first molecule according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules, comprising performing the following steps for each first molecule of the plurality of first molecules: For any two non-hydrogen atoms in the first molecule, respectively obtain target atomic degrees and target edge attributes of the two non-hydrogen atoms; According to the preset correspondence between the atomic degrees of the two non-hydrogen atoms and the edge attributes and the chemical bond type, determine the second chemical bond type corresponding to the target atomic degrees and the target edge attributes of the two non-hydrogen atoms; The two non-hydrogen atoms are connected according to the second chemical bond type to obtain an initial second molecule.

8. The method according to claim 7, characterized in that After the two non-hydrogen atoms are connected according to the second chemical bond type to obtain an initial second molecule, the method further comprises: Performing graph isomorphism judgment on the initial second molecule, the multiple initial first molecules and other initial second molecules at the molecular graph level; When there is an initial first molecule or other initial second molecule that is graph-isomorphic to the initial second molecule, the initial second molecule is deleted.

9. The method according to claim 1, characterized in that: The preset molecule generation rule includes executing the following steps for each initial molecule: According to the atomic addition rule, the valence rule and the molecular stability rule, adding a preset number of preset atoms to the initial molecule to obtain a plurality of fifth molecules; storing the plurality of fifth molecules into a preset target molecule set; Determine a target fifth molecule from the plurality of fifth molecules, the target fifth molecule being the fifth molecule having the least number of edges among the plurality of fifth molecules; According to the atomic addition rule, the valence rule and the molecular stability rule, adding a preset number of preset atoms to the target fifth molecule to obtain a plurality of sixth molecules; storing the plurality of sixth molecules into a preset target molecule set; Taking the plurality of sixth molecules as new plurality of fifth molecules, and returning to the step of determining a target fifth molecule from the plurality of fifth molecules until a preset iteration stop condition is satisfied; According to the chemical bond addition rule, the valence rule and the molecular stability rule, respectively adding chemical bonds of preset types to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of seventh molecules; storing the plurality of seventh molecules into a preset target molecule set; According to a preset graph isomorphism merging rule, molecules with graph isomorphism in the preset target molecule set are merged.

10. The method according to claim 9, characterized in that The graph isomorphism merging rule includes a graph isomorphism merging rule based on a step-by-step merging algorithm.

11. The method according to claim 9, characterized in that The graph isomorphism merging rule includes a graph isomorphism merging rule based on a multi-grouping algorithm.

12. The method according to claim 9, characterized in that The molecular stability rules include molecular structure diagrams that do not comply with molecular stability; The step of adding a preset number of preset atoms to the initial molecule according to the atomic addition rule, the valence rule and the molecular stability rule to obtain a plurality of fifth molecules includes: Adding a preset number of preset atoms to the initial molecule according to an atom addition rule and a valence rule to obtain a plurality of initial fifth molecules; searching for the molecular structure diagram in the plurality of initial fifth molecules; The initial fifth molecule including the molecular structure diagram is deleted to obtain the plurality of fifth molecules.

13. The method according to claim 12, characterized in that The initial molecule includes at least one non-hydrogen atom; The step of adding a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain a plurality of initial fifth molecules includes: Obtaining the number of the preset atoms in the initial molecule; When the number of the preset atoms is less than a preset threshold, the following steps are performed for each non-hydrogen atom in the initial molecule: the atomic degree of the non-hydrogen atom is calculated, the atomic degree including the sum of the weighted values ​​of the edges connected to the corresponding nodes of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attributes, the edge attributes being used to represent the attributes of the chemical bonds corresponding to the edges; the third chemical bond type corresponding to the atomic degree of the non-hydrogen atom is determined according to the correspondence between the preset atomic degree and the chemical bond type; and the non-hydrogen atom and the preset atom are connected according to the third chemical bond type to obtain an initial fifth molecule.

14. The method according to claim 12, characterized in that According to the chemical bond addition rule, the valence rule and the molecular stability rule, chemical bonds of preset types are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set respectively to obtain a plurality of seventh molecules, including: According to the chemical bond addition rule and the valence rule, respectively adding chemical bonds of preset types to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of initial seventh molecules, and searching the molecular structure diagram in the plurality of initial seventh molecules; The initial seventh molecule including the molecular structure diagram is deleted to obtain the plurality of seventh molecules.

15. The method according to claim 14, characterized in that The fifth molecule and the sixth molecule each include at least two non-hydrogen atoms; According to the chemical bond addition rule and the valence rule, chemical bonds of preset types are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of initial seventh molecules, including performing the following steps for each molecule in the plurality of fifth molecules and the plurality of sixth molecules: For any two non-hydrogen atoms in the molecule, respectively obtain the target atomic degree and target edge attribute of each of the two non-hydrogen atoms; According to the preset correspondence between the atomic degrees of the two non-hydrogen atoms and the edge attributes and the chemical bond types, determining the fourth chemical bond type corresponding to the target atomic degrees and the target edge attributes of the two non-hydrogen atoms; The two non-hydrogen atoms are connected according to the fourth chemical bond type to obtain a plurality of initial seventh molecules.

16. The method according to any one of claims 1 to 15, characterized in that The molecular stability rules include molecular structure diagrams that do not comply with molecular stability; Before adding preset atoms and / or preset chemical bonds to the initial molecule according to the preset molecule generation rule, the method further includes: Acquiring the application environment of the plurality of target molecular structure data; According to the chemical stability rule and the application environment, a molecular structure diagram that does not meet the molecular stability requirements in the application environment is determined.

17. The method according to any one of claims 1 to 15, characterized in that The predetermined atoms include at least one of carbon, nitrogen, oxygen, fluorine, silicon, sulfur, and chlorine.

18. The method according to any one of claims 1 to 15, characterized in that The multiple target molecules are converted into a simplified molecular linear input specification SMILES format to obtain multiple target molecular structure data, including: Converting the plurality of target molecules into an adjacency matrix or adjacency table format to obtain intermediate data of target molecule structures; The target molecular structure intermediate data is converted into SMILES format to obtain the multiple target molecular structure data.

19. The method according to any one of claims 1 to 15, characterized in that After converting the plurality of target molecules into a simplified molecular linear input specification SMILES format to obtain a plurality of target molecular structure data, the method further comprises: Based on multiple target molecular structure data, a target molecular structure database is created.

20. A molecular structure data generating device, characterized in that: include: An acquisition module, used for acquiring an initial molecule, wherein the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds; An adding module, used for adding preset atoms and / or preset chemical bonds to the initial molecules according to preset molecular generation rules to obtain multiple target molecules, wherein the molecular generation rules include atom addition rules, chemical bond addition rules, valence rules and molecular stability rules, and the target molecules are represented by an undirected graph without self-loops; The conversion module is used to convert the multiple target molecules into SMILES format to obtain multiple target molecule structure data.

21. An electronic device, characterized in that: The device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the method for generating molecular structure data according to any one of claims 1 to 19 is implemented.

22. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the method for generating molecular structure data according to any one of claims 1 to 19 is implemented.

23. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the molecular structure data generating method as described in any one of claims 1 to 19.

Citation Information

Patent Citations

  • Graph model drug generation method, device and medium based on reinforcement learning

    CN110459275A

  • Molecular optimization method and system, terminal equipment and readable storage medium

    CN112509644A

  • Directional molecule generation method based on graph neural network

    CN113140267A

  • Molecular map generation method and device

    CN114822721A

  • Drug data enhancement method and device, electronic equipment and storage medium

    CN115240789A