Molecular structure data generation method, device, equipment, medium and program product

By representing the molecular structure through graph theory methods, the electrolyte molecular structure data is generated and converted into SMILES format, which solves the problem of slow generation of electrolyte molecular database in the existing technology, realizes efficient electrolyte R&D support, and improves the efficiency of electrolyte design and optimization.

CN119993321BActive Publication Date: 2025-10-21TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510080602.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-10-21
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Existing technologies lack methods to quickly generate and update electrolyte molecular structure databases, which limits the application of high-throughput computing, high-throughput screening and machine learning technologies in electrolyte research and development.

Method used

The molecular structure is represented by graph theory. The initial molecules are represented by undirected graphs. Preset atoms and chemical bonds are added according to the preset molecular generation rules to generate multiple target molecules. These molecules are converted into SMILES format to construct an electrolyte molecular structure database.

Benefits of technology

It achieves efficient generation of target molecular structure data with chemical rationality and stability, supports high-throughput computing and machine learning, significantly shortens the electrolyte R&D cycle, reduces R&D costs, and improves design and optimization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993321B_ABST
    Figure CN119993321B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a molecular structure data generation method, device, equipment, medium and program product. The method comprises: obtaining an initial molecule, the initial molecule being represented by an undirected graph without self-loop, the undirected graph comprising nodes for representing atoms and edges for representing chemical bonds; adding preset atoms and / or preset chemical bonds to the initial molecule according to a preset molecule generation rule to obtain a plurality of target molecules, wherein the molecule generation rule comprises an atom addition rule, a chemical bond addition rule, a valence rule and a molecule stability rule, and the target molecules are represented by undirected graphs without self-loop; and converting the plurality of target molecules into a simplified molecular linear input specification (SMILES) format to obtain a plurality of target molecular structure data. The embodiments of the present application can efficiently generate molecular structure data in batches, thereby supporting application scenarios such as machine learning, high-throughput calculation and high-throughput screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of chemical informatics, and in particular relates to a method, apparatus, device, medium and program product for generating molecular structure data. Background Art

[0002] With the development of battery technology, batteries are widely used in portable electronic devices, electric vehicles, and energy storage systems. Among the components of a battery, the electrolyte, as a key electrochemical medium, directly affects the battery's energy density, cycle life, and safety performance. Therefore, the development of high-performance electrolytes has become a key direction in the development of battery technology.

[0003] Traditional electrolyte research and development mainly relies on experience and trial and error, testing electrolyte molecules one by one in the laboratory to find the electrolyte with the best performance. This method is not only time-consuming and costly, but also due to the limited number of test samples, it can often only find local optimal solutions, making it difficult to achieve comprehensive optimization. In this context, there is an urgent need for an efficient method to accelerate electrolyte research and development. The development of big data and artificial intelligence technologies has provided new possibilities for the rapid development of electrolytes. High-throughput computing and high-throughput screening technologies can evaluate the performance of a large number of candidate molecules in a short period of time, greatly shortening the R&D cycle; machine learning algorithms can predict the performance of new materials and guide experimental design by analyzing large amounts of experimental data. The application of these technologies requires a large amount of molecular structure data as a basic support.

[0004] However, there is currently a lack of a method that can quickly generate and update electrolyte molecular structure databases, which limits the application of high-throughput computing, high-throughput screening and machine learning technologies in electrolyte research and development. Summary of the Invention

[0005] The embodiments of the present application provide a molecular structure data generation method, apparatus, device, computer storage medium, and computer program product, which can efficiently generate molecular structure data in batches, thereby supporting application scenarios such as machine learning, high-throughput computing, and high-throughput screening.

[0006] In a first aspect, an embodiment of the present application provides a method for generating molecular structure data, comprising:

[0007] Obtain an initial molecule, which is represented by an undirected graph without self-loops. The undirected graph includes nodes for representing atoms and edges for representing chemical bonds.

[0008] According to preset molecular generation rules, preset atoms and / or preset chemical bonds are added to the initial molecules to obtain multiple target molecules, wherein the molecular generation rules include atom addition rules, chemical bond addition rules, valence rules, and molecular stability rules, and the target molecules are represented by an undirected graph without self-loops;

[0009] The multiple target molecules are converted into the simplified molecular linear input specification SMILES format to obtain the structural data of the multiple target molecules.

[0010] In an optional embodiment, the preset molecule generation rule includes performing the following steps for each initial molecule:

[0011] According to the atomic addition rule, the valence rule, and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of first molecules;

[0012] For each first molecule in the plurality of first molecules, adding a chemical bond of a preset type to the first molecule according to a chemical bond addition rule, a valence rule, and a molecular stability rule to obtain a plurality of second molecules;

[0013] storing the plurality of first molecules and the plurality of second molecules as a plurality of primary target molecules into a preset target molecule set;

[0014] For each of the plurality of primary target molecules, a preset number of preset atoms are added to the primary target molecule according to the atom addition rule, the valence rule, and the molecular stability rule, to obtain a plurality of third molecules; for each of the plurality of third molecules, a preset type of chemical bond is added to the third molecule according to the chemical bond addition rule, the valence rule, and the molecular stability rule, to obtain a plurality of fourth molecules; the plurality of third molecules and the plurality of fourth molecules are used as a plurality of secondary target molecules and stored in a preset target molecule set;

[0015] Taking multiple secondary target molecules as new multiple primary target molecules, returning to each of the multiple primary target molecules, respectively, according to the atom addition rule, valence rule and molecular stability rule, adding a preset number of preset atoms to the primary target molecule to obtain multiple third molecules until the preset iteration stop condition is met.

[0016] In an alternative embodiment, the molecular stability rule includes a molecular structure diagram that does not comply with the molecular stability rule;

[0017] According to the atomic addition rule, the valence rule, and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of first molecules, including:

[0018] Adding a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain a plurality of initial first molecules;

[0019] searching for a molecular structure diagram among a plurality of initial first molecules;

[0020] The initial first molecule including the molecular structure graph is deleted to obtain a plurality of first molecules.

[0021] In an alternative embodiment, the initial molecule includes at least one non-hydrogen atom;

[0022] According to the atomic addition rule and the valence rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of initial first molecules, including performing the following steps on the initial molecules for each preset atom:

[0023] Get the number of preset atoms in the initial molecule;

[0024] When the number of preset atoms is less than a preset threshold, the following steps are performed for each non-hydrogen atom in the initial molecule: the atomic degree of the non-hydrogen atom is calculated, where the atomic degree includes the sum of the weighted values ​​of the edges connected to the corresponding nodes in the undirected graph corresponding to the initial molecule and the corresponding edge attributes, where the edge attributes are used to represent the properties of the chemical bonds corresponding to the edges; based on the correspondence between the preset atomic degree and the chemical bond type, the first chemical bond type corresponding to the atomic degree of the non-hydrogen atom is determined; and the non-hydrogen atom and the preset atom are connected according to the first chemical bond type to obtain the initial first molecule.

[0025] In an optional embodiment, after connecting the non-hydrogen atom and the predetermined atom according to the first chemical bond type to obtain the initial first molecule, the method further includes:

[0026] Conduct graph isomorphism judgment on the molecular graph level between the initial first molecule and other initial first molecules;

[0027] If there are other initial first molecules that are isomorphic to the initial first molecule graph, the initial first molecule is deleted.

[0028] In an optional embodiment, for each first molecule in the plurality of first molecules, adding a predetermined type of chemical bond to the first molecule according to a chemical bond addition rule, a valence rule, and a molecular stability rule to obtain a plurality of second molecules includes:

[0029] For each of the plurality of first molecules, adding a preset type of chemical bond in the first molecule according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules;

[0030] searching for a molecular structure diagram among a plurality of initial second molecules;

[0031] The initial second molecule including the molecular structure graph is deleted to obtain a plurality of second molecules.

[0032] In an alternative embodiment, the first molecule comprises at least two non-hydrogen atoms;

[0033] For each first molecule in the plurality of first molecules, adding a preset type of chemical bond in the first molecule according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules, including performing the following steps for each first molecule in the plurality of first molecules:

[0034] For any two non-hydrogen atoms in the first molecule, obtain the target atomic degrees and target edge attributes of the two non-hydrogen atoms respectively;

[0035] Determining the second chemical bond type corresponding to the target atomic degree of each of the two non-hydrogen atoms and the target edge attribute according to the preset correspondence between the atomic degree of each of the two non-hydrogen atoms and the edge attribute and the chemical bond type;

[0036] The two non-hydrogen atoms are connected according to the second chemical bond type to obtain an initial second molecule.

[0037] In an optional embodiment, after connecting the two non-hydrogen atoms according to the second chemical bond type to obtain the initial second molecule, the method further includes:

[0038] Performing graph isomorphism judgment on the initial second molecule, the multiple initial first molecules, and the other initial second molecules at the molecular graph level;

[0039] When there is an initial first molecule or other initial second molecule that is graph-isomorphic to the initial second molecule, the initial second molecule is deleted.

[0040] In an optional embodiment, the preset molecule generation rule includes performing the following steps for each initial molecule:

[0041] According to the atomic addition rule, the valence rule, and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of fifth molecules;

[0042] storing the plurality of fifth molecules into a preset target molecule set;

[0043] determining a target fifth molecule from a plurality of fifth molecules, where the target fifth molecule is a fifth molecule having the least number of edges among the plurality of fifth molecules;

[0044] According to the atomic addition rule, the valence rule, and the molecular stability rule, a preset number of preset atoms are added to the target fifth molecule to obtain a plurality of sixth molecules;

[0045] storing the plurality of sixth molecules into a preset target molecule set;

[0046] Taking the plurality of sixth molecules as new plurality of fifth molecules, and returning to determine a target fifth molecule from the plurality of fifth molecules until a preset iteration stop condition is satisfied;

[0047] According to the chemical bond addition rule, the valence rule, and the molecular stability rule, respectively adding chemical bonds of a preset type to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of seventh molecules;

[0048] storing the plurality of seventh molecules into a preset target molecule set;

[0049] According to the preset graph isomorphism merging rule, molecules with graph isomorphism in the preset target molecule set are merged.

[0050] In an optional implementation, the graph isomorphism merging rule includes a graph isomorphism merging rule based on a step-by-step merging algorithm.

[0051] In an optional implementation, the graph isomorphism merging rule includes a graph isomorphism merging rule based on a multi-grouping algorithm.

[0052] In an alternative embodiment, the molecular stability rule includes a molecular structure diagram that does not comply with the molecular stability rule;

[0053] According to the atomic addition rule, the valence rule, and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain multiple fifth molecules, including:

[0054] Adding a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain a plurality of initial fifth molecules;

[0055] Finding molecular structure diagrams among a plurality of initial fifth molecules;

[0056] The initial fifth molecule including the molecular structure diagram is deleted to obtain multiple fifth molecules.

[0057] In an alternative embodiment, the initial molecule includes at least one non-hydrogen atom;

[0058] According to the atomic addition rule and the valence rule, a preset number of preset atoms are added to the initial molecule to obtain multiple initial fifth molecules, including:

[0059] Get the number of preset atoms in the initial molecule;

[0060] When the number of preset atoms is less than a preset threshold, the following steps are performed for each non-hydrogen atom in the initial molecule: the atomic degree of the non-hydrogen atom is calculated, where the atomic degree includes the sum of the weighted values ​​of the edges connected to the corresponding nodes in the undirected graph corresponding to the initial molecule and the corresponding edge attributes, where the edge attributes are used to represent the properties of the chemical bonds corresponding to the edges; based on the correspondence between the preset atomic degree and the chemical bond type, the third chemical bond type corresponding to the atomic degree of the non-hydrogen atom is determined; and the non-hydrogen atom and the preset atom are connected according to the third chemical bond type to obtain the initial fifth molecule.

[0061] In an optional embodiment, according to the chemical bond addition rule, the valence rule, and the molecular stability rule, chemical bonds of a preset type are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set, respectively, to obtain the plurality of seventh molecules, including:

[0062] According to the chemical bond addition rule and the valence rule, chemical bonds of preset types are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of initial seventh molecules, and a molecular structure diagram is searched in the plurality of initial seventh molecules;

[0063] The initial seventh molecule including the molecular structure diagram is deleted to obtain multiple seventh molecules.

[0064] In an alternative embodiment, the fifth molecule and the sixth molecule each include at least two non-hydrogen atoms;

[0065] According to the chemical bond addition rule and the valence rule, chemical bonds of a preset type are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of initial seventh molecules, including performing the following steps for each of the plurality of fifth molecules and the plurality of sixth molecules:

[0066] For any two non-hydrogen atoms in a molecule, obtain the target atomic degree and target edge attributes of each of the two non-hydrogen atoms;

[0067] Determining the fourth chemical bond type corresponding to the target atomic degrees of the two non-hydrogen atoms and the target edge attributes according to the preset correspondence between the atomic degrees of the two non-hydrogen atoms and the edge attributes and the chemical bond type;

[0068] Two non-hydrogen atoms are connected according to the fourth chemical bond type to obtain multiple initial seventh molecules.

[0069] In an alternative embodiment, the molecular stability rule includes a molecular structure diagram that does not comply with the molecular stability rule;

[0070] Before adding the preset atoms and / or preset chemical bonds to the initial molecule according to the preset molecule generation rule, the method further includes:

[0071] Application environment for obtaining structural data of multiple target molecules;

[0072] According to the chemical stability rules and the application environment, determine the molecular structure diagram that does not meet the molecular stability requirements under the application environment.

[0073] In an optional embodiment, the preset atoms include at least one of carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl).

[0074] In an optional embodiment, multiple target molecules are converted into SMILES format to obtain multiple target molecule structure data, including:

[0075] Convert multiple target molecules into adjacency matrix or adjacency table format to obtain intermediate data of target molecule structure;

[0076] The target molecular structure intermediate data is converted into SMILES format to obtain multiple target molecular structure data.

[0077] In an optional embodiment, after converting the plurality of target molecules into SMILES format to obtain the plurality of target molecule structure data, the method further comprises:

[0078] A target molecular structure database is created based on multiple target molecular structure data.

[0079] In a second aspect, an embodiment of the present application provides a molecular structure data generating device, comprising:

[0080] An acquisition module is used to acquire an initial molecule, where the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds;

[0081] An adding module, configured to add preset atoms and / or preset chemical bonds to the initial molecule according to preset molecular generation rules to obtain multiple target molecules, wherein the molecular generation rules include atom addition rules, chemical bond addition rules, valence rules, and molecular stability rules, and the target molecules are represented by an undirected graph without self-loops;

[0082] The conversion module is used to convert multiple target molecules into SMILES format to obtain multiple target molecular structure data.

[0083] In a third aspect, an embodiment of the present application provides an electronic device, the device comprising: a processor and a memory storing computer program instructions;

[0084] When the processor executes the computer program instructions, it implements the molecular structure data generation method of any optional embodiment of the first aspect of the present application.

[0085] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, a method for generating molecular structure data according to any optional embodiment of the first aspect of the present application is implemented.

[0086] In a fifth aspect, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes a molecular structure data generation method as described in any optional embodiment of the first aspect of the present application.

[0087] The molecular structure data generation method, device, equipment, computer storage medium and computer program product of the embodiment of the present application can obtain an initial molecule, which is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds. In this way, by using graph theory to represent molecules, it is possible to ensure the structuring and standardization of molecular data during the molecule generation process, thereby improving the efficiency of data processing. Then, according to the preset molecule generation rules, preset atoms and / or preset chemical bonds can be added to the initial molecule to obtain multiple target molecules, wherein the molecule generation rules include atom addition rules, chemical bond addition rules, valence rules and molecular stability rules, and the target molecule is represented by an undirected graph without self-loops. On the one hand, the molecule generation rules can make the structure of the generated target molecule chemically reasonable and stable, reducing the risk of generating unreasonable molecules; on the other hand, by specifying atom addition rules and chemical bond addition rules, target molecules that meet user needs can be generated in batches, thereby improving the efficiency and flexibility of molecule generation. In this way, the generated multiple target molecules can adapt to the needs of different research scenarios. Subsequently, the multiple target molecules can also be converted into SMILES format to obtain multiple target molecular structure data. The SMILES format is a more standardized molecular representation, which is convenient for retrieval and storage, and does not have ambiguity. By converting multiple target molecules into the SMILES format, the consistency and accuracy of the target molecule structure data can be guaranteed. In this way, a solid data foundation can be provided for high-throughput computing, machine learning and high-throughput screening. The target molecule structure data generated in batches in the application embodiment are used to construct an electrolyte molecular structure database, which can effectively support the rapid development of new electrolyte systems, significantly shorten the development cycle, reduce development costs, and improve the efficiency of electrolyte design and optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0089] Figure 1 This is a flow chart of a method for generating molecular structure data provided by one embodiment of the present application;

[0090] Figure 2 is a schematic structural diagram of a molecular structure data generating device provided in yet another embodiment of the present application;

[0091] Figure 3 A schematic structural diagram of an electronic device is provided in yet another embodiment of the present application. DETAILED DESCRIPTION

[0092] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.

[0093] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.

[0094] As described in the background technology, there is currently a lack of a method that can quickly generate and update the electrolyte molecular structure database, which limits the application of high-throughput computing, high-throughput screening and machine learning technologies in electrolyte research and development.

[0095] In chemistry, a molecule is a structure composed of its constituent atoms arranged according to specific valence rules and spatial arrangements. Within a molecule, each atom interacts with the others, and when the attractive and repulsive forces of these interactions are balanced, the molecule is stable.

[0096] The bonding structure and arrangement of atoms in a molecule is called molecular structure. Methods for describing molecular structure typically include structural formula, bond-line formula, ball-and-stick formula, and ratio formula. These methods accurately describe molecular structure in chemistry. However, in computer languages, relying solely on these representations often fails to accurately describe molecular structure and can lead to numerous ambiguities (for example, when using structural formulas to represent a database, one-to-many relationships may exist).

[0097] In view of this, the inventors, after in-depth thinking, cleverly proposed a molecular structure data generating method, information receiving method, device, equipment, computer storage medium and computer program product.

[0098] In the embodiment of the present application, a molecule can be regarded as an undirected graph without self-loops in graph theory, denoted by G=<V,E> , where each atom is considered as a set of points V = {v1,v2,...,v n}, each key can be regarded as a set of edges E = {e1,e2,...,e m}, the information of atoms is stored in the attribute set A as the attribute of the point set V, denoted as A={a1,a2,...,a n}, the key information is stored as an attribute in the edge set E in the attribute set B, denoted as B = {b1, b2, ..., b m}.

[0099] In the undirected graph corresponding to the molecule, the atomic degree of a node can be expressed as the sum of the weighted values ​​of the edges connected to the node and the attributes on the edges. As an example, the atomic degree of node V1 can be recorded as d1 = b1 + b2 + ... + b p , where the number of edges connected by V1 is p, b1 to b p They can respectively represent the weighted values ​​of the attributes of the first to pth edges and their edges. In some embodiments, when the edge is a single bond, the weighted value of the edge and the attribute on the edge can be 1; when the edge is a double bond, the weighted value of the edge and the attribute on the edge can be 2; when the edge is a triple bond, the weighted value of the edge and the attribute on the edge can be 3. The node degree can be expressed as excluding the set of attributes on the edges, which can be equal to the number of edges connected to the node. Generally, the atomic degree of a node is usually less than or equal to 4.

[0100] The following, in conjunction with the accompanying drawings, describes the molecular structure data generation method provided in the embodiments of the present application through specific embodiments and their application scenarios. The molecular structure data generation method provided in the embodiments of the present application, wherein the device executing the sending can be a molecular structure data generation device, or a partial module of the molecular structure data generation device for executing the molecular structure data generation method. In the embodiments of the present application, the molecular structure data generation method provided in the embodiments of the present application is described in detail using the molecular structure data generation device executing the molecular structure data generation method as an example.

[0101] The following is combined with Figure 1 The molecular structure data generation method provided in the examples of the present application is described in detail.

[0102] Figure 1 FIG. 1 is a flow chart showing a method for generating molecular structure data according to an embodiment of the present application. Figure 1 As shown, the molecular structure data generating method may specifically include the following steps S110 to S130.

[0103] S110, obtaining an initial molecule, where the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds.

[0104] In step S110, the initial molecule may be a base molecule for molecular amplification. The initial molecule may be selected based on the user's needs and is not limited herein. The number of initial molecules may be one or more. It is understood that when there are multiple initial molecules, atomic and chemical bond amplification may be performed separately for each of the multiple initial molecules.

[0105] S120, according to the preset molecular generation rules, adding preset atoms and / or preset chemical bonds to the initial molecules to obtain multiple target molecules, wherein the molecular generation rules include atom addition rules, chemical bond addition rules, valence rules and molecular stability rules, and the target molecules are represented by an undirected graph without self-loops.

[0106] In step S120, the preset atoms can be selected according to the needs of the user and are not limited here. As an example, the preset atoms can include non-hydrogen atoms. For example, when the user needs to construct molecular structure data containing only carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl) elements, the initial molecule can be a molecule composed of at least one of carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl), and the preset atoms can include carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl). The atom addition rule can be used to specify the position, conditions, and method of adding new atoms to the molecule and adding chemical bonds accordingly. The chemical bond addition rule can be used to specify the position, conditions, and method of adding chemical bonds to the molecule. The valence rule can be used to specify the basic chemical rules that an atom needs to meet to form a bond with other atoms, so that the number of other atoms that the atom can connect to is saturated. The molecular stability rule can be used to ensure that the generated target molecule meets the chemical stability requirements. It is understandable that when adding atoms and / or chemical bonds to a molecule, the atoms connected to the newly added atoms or chemical bonds may lose a corresponding number of hydrogen atoms according to the valence rules.

[0107] In step S120, according to the molecular generation rules, different types or quantities of preset atoms and / or preset chemical bonds can be added to the initial molecule in a variety of ways, and the initial molecule can also be iteratively amplified with atoms and / or chemical bonds, thereby efficiently and automatically generating a large number of target molecules in batches.

[0108] S130, converting the multiple target molecules into SMILES format to obtain multiple target molecule structure data.

[0109] The molecular structure data generation method of the embodiment of the present application can obtain an initial molecule, which is represented by an undirected graph without self-loops. The undirected graph includes nodes for representing atoms and edges for representing chemical bonds. In this way, by using graph theory to represent molecules, it is possible to ensure the structuring and standardization of molecular data during the molecule generation process, thereby improving the efficiency of data processing. Then, according to the preset molecule generation rules, preset atoms and / or preset chemical bonds can be added to the initial molecule to obtain multiple target molecules, wherein the molecule generation rules include atom increase rules, chemical bond increase rules, valence rules and molecular stability rules, and the target molecule is represented by an undirected graph without self-loops. On the one hand, the molecule generation rules can make the structure of the generated target molecule chemically reasonable and stable, reducing the risk of generating unreasonable molecules; on the other hand, by specifying atom increase rules and chemical bond increase rules, target molecules that meet user needs can be generated in batches, thereby improving the efficiency and flexibility of molecule generation. In this way, the generated multiple target molecules can adapt to the needs of different research scenarios. Subsequently, the multiple target molecules can also be converted into SMILES format to obtain multiple target molecule structure data. The SMILES format is a more standardized molecular representation, which is convenient for retrieval and storage, and does not have ambiguity. By converting multiple target molecules into the SMILES format, the consistency and accuracy of the target molecule structure data can be guaranteed. In this way, a solid data foundation can be provided for high-throughput computing, machine learning and high-throughput screening. The target molecule structure data generated in batches in the application embodiment are used to construct an electrolyte molecular structure database, which can effectively support the rapid development of new electrolyte systems, significantly shorten the development cycle, reduce development costs, and improve the efficiency of electrolyte design and optimization.

[0110] In one embodiment, molecular structure data can be generated based on a serial molecular generation algorithm. This solution can be implemented based on a serial generation algorithm system. For ease of understanding, the serial generation algorithm system is briefly introduced below.

[0111] The serial generation algorithm system can include an initial molecule generation module, a graph constraint module, an atom augmentation module, a bond augmentation module, a format conversion module, and an auxiliary operation module. In the serial generation algorithm system, the initial molecule can be used as the raw input. New nodes are added based on possible structural combinations of atoms and bonds, and isomorphism pruning is performed simultaneously. Single, double, and triple bonds are added based on a breadth-first search combination. With the help of an auxiliary queue, the graph with the added new node is used as the root node and queued. After dequeuing, edges are added. Existing graphs are searched and isomorphisms are determined. If a graph with the current number of nodes and edges is not generated or is not isomorphic with an existing graph, it is queued and the cycle continues.

[0112] The following is a detailed description of each module described in the serial generation algorithm.

[0113] In the initial molecule generation module, the initial molecules can be represented according to graph theory rules, encapsulated by defining classes or functional modules, or stored appropriately using data structures such as adjacency matrices or adjacency lists. During the construction process, to ensure the consistency of the generated database, all elements in the database must be molecules and not contain isolated atoms. In other words, the initial molecule and each output generated iteratively must be a complete, independent molecule. To simplify the database construction algorithm and reduce both space and time complexity, hydrogen atoms are omitted from the representation of the molecules. That is, if an atom in a molecule does not meet the valence rules, the missing valence is automatically filled with hydrogen atoms. This representation method allows for automatic molecule completion using functional functions during database queries, avoiding potential ambiguity in database searches. This ensures that this molecular representation achieves both simplicity and rationality during database construction and generation.

[0114] For example, if you want to build a database that only contains carbon, hydrogen, nitrogen, oxygen, and fluorine elements, then at the beginning of the construction you only need to define the strings "C", "N", "O", and "F" and store them in a reasonable form. It should be noted that although the initial molecules look like single atoms, they actually represent "CH4", "NH3", "H2O", and "HF" molecules respectively.

[0115] The form of the initial molecule can be determined based on the system being studied and the actual situation, and the embodiments of this application do not limit this. For example, if you want to explore the relationship between the structure and properties of lithium battery electrolyte solvent molecules, and want to build an initial database for applications such as machine learning, you need to consider the atoms that may be present in the electrolyte molecules. Specifically, most electrolyte molecules usually contain functional groups such as ethers and esters, and the introduction of fluorine atoms helps to improve the stability of the electrolyte. Therefore, when constructing the initial data set, water molecules and hydrogen fluoride molecules should be introduced to provide the required oxygen and fluorine elements.

[0116] The graph constraint module can be used to delete molecules that do not meet the molecular stability rules. During the generation of the electrolyte molecule database, chemically unstable structures are inevitably generated. These unstable structures need to be found and efficiently deleted.

[0117] In mathematics, if the nodes and edges between nodes in a graph S can be mapped to another graph G, then it is considered that the graph S and the graph G satisfy the subgraph isomorphism relationship. In the embodiment of the present application, subgraph isomorphism plays an important role in searching for functional groups in molecules and finding related topological structures. Using the principle of subgraph isomorphism, molecular graphs containing unstable substructures can be efficiently deleted. Taking the structure containing carbon atoms and oxygen atoms as an example, it is recommended to delete the following molecular graphs:

[0118] (1) Molecular diagrams of small rings (three-membered rings or four-membered rings) containing double bonds or triple bonds;

[0119] (2) Molecular diagrams containing bridgehead atoms in small rings;

[0120] (3) Molecular diagram containing a spirocycle with small rings on both sides.

[0121] The main purpose of eliminating these molecular graphs is to remove stress-concentrated structures. These structures with large ring tension or topological distortion are difficult to synthesize in practice. Even if they can be constructed in a computer, they have no practical significance for research. In some embodiments, molecular graphs that are unstable in the application environment can also be deleted. The constraints of such molecular graphs can be determined based on prior knowledge of the molecular structure. For example, in a metal lithium battery system, it is usually required that the molecular structure does not contain active hydrogen (i.e., hydrogen atoms on groups such as hydroxyl and carboxyl groups). In this case, the molecular graphs containing active hydrogen can be deleted.

[0122] The atom augmentation module is used to add a preset number of new atoms to each molecule and add corresponding bonds. The organic small molecule database generation algorithm needs to call the atom augmentation module when iterating and adding points. The preset number can be set according to actual needs. For example, 1 new atom, 2 new atoms, 3 new atoms, etc. can be added to each molecule in each round, which is not limited here. Exemplarily, the preset number can be 1, and the atom augmentation module can be called only once during one round. For example, in each round of iteration, one atom can be added to the molecule of the current round through the atom augmentation module.

[0123] Taking the preset number of 1 as an example, the atomic augmentation module can take the number of atoms in the molecule in the current round as input, and output the set of molecular graphs after adding one atom and the corresponding chemical bond. When the atomic augmentation module starts running, it first determines whether the number of atoms in the current round is equal to the set upper limit of the number of atoms. If they are equal, it returns directly, indicating that the iteration is completed. If they are not equal, continue to perform subsequent operations. After the judgment, enter the loop operation, extract all the molecular graphs generated in the previous round, if the previous round is the initial molecule, extract all the initial molecular graphs, and then consider whether each atom in each molecule meets the following conditions (only common atoms are analyzed as examples here).

[0124] 1. Carbon Atom

[0125] If the currently traversed atom is a carbon atom, the first step is to check whether the number of atoms connected to the carbon atom in the current molecule has reached the upper limit. If so, the atom is skipped without skipping the current molecule graph, and the traversal continues for other atoms. If not, the subsequent operations are performed.

[0126] (1) Check the degree of the current carbon atom to determine whether it is less than 4. The degree here refers to the weighted number considering the chemical bond. If the degree is greater than or equal to 4, skip this judgment; if the degree is less than 4, continue to perform subsequent operations.

[0127] Next, consider the candidate elements (excluding hydrogen) contained in the current round's molecule. If all of these elements can form single bonds with carbon atoms, then the candidate atoms corresponding to each candidate element are sequentially connected to the carbon atom with single bonds. Then, the molecular graph with one node and one edge added is stored. For example, if the current round's molecule is the initial molecule and contains carbon, oxygen, nitrogen, and fluorine, the candidate elements can be carbon, oxygen, nitrogen, and fluorine. At this point, the atoms that can form single bonds with carbon atoms are carbon, oxygen, nitrogen, and fluorine. Therefore, the atoms and bonds that can be added are "C–C," "C–O," "C–N," and "C–F," respectively. These substructures are stored as part of the newly generated molecular graph in the graph storage set.

[0128] It is recommended to copy the original molecular graph before performing atomic augmentation, rather than directly operating on the graph during the loop. This is because during the loop, the base graph must be sequentially added with the selected atoms to generate the new molecular graph. Due to the limitations of different programming languages, not copying may lead to unexpected errors, which may cause atoms to be overwritten.

[0129] (2) Check the degree of the carbon atom. Here, the carbon atom refers to the atom without adding new atoms and new bonds, that is, the carbon atom in the cyclic base diagram, and determine whether its degree is less than 3, and check whether there is a double bond. If the degree of the carbon atom is greater than or equal to 3 or there is a double bond, skip this judgment; if the degree is less than 3 and the carbon atom is not connected to a double bond, continue to perform subsequent operations. It should be noted that the purpose of judging whether the carbon atom has a double bond here is to avoid having two double bonds on one carbon atom, that is, forming a cumulative diene. Such a structure is extremely unstable in the case of small molecular weight and is easily reacted into other substances or cannot be synthesized.

[0130] Next, it is necessary to select an atom that can form a double bond with a carbon atom from the set of candidate atoms and form a new molecular graph with the carbon atom. This means adding an atom and a double bond to a new molecule, and saving each of these molecules in the graph set in sequence. For example, if the initial molecule contains carbon, oxygen, nitrogen, and fluorine, then the atoms that can form a double bond with the carbon atom are carbon, oxygen, and nitrogen. Therefore, the atoms and bonds that can be added are "C=C," "C=O," and "C=N," respectively. The molecular graphs for this added substructure are then stored separately.

[0131] (3) Then check the degree of the carbon atom again to determine whether it is less than 2. If the degree of the carbon atom is greater than or equal to 2, skip this judgment; if the degree is less than 2, continue to perform subsequent operations. The reason why it is not necessary to consider whether it is connected to a double bond here is that if it already has a double bond, the maximum degree can only be equal to 2, and it cannot be less than 2. Therefore, only the degree restriction can eliminate the instability problem of two double bonds on a carbon atom. Next, select the atom that can form a triple bond with the carbon atom from the set of atoms to be selected, and add it to the graph set after forming a bond with it. For example, the atoms that can form a triple bond with the carbon atom at this time are carbon atoms and nitrogen atoms. Therefore, the atoms and bonds that can be added are "C≡C" and "C≡N" respectively, and then the molecular graphs with triple bonds and a new atom are stored in sequence.

[0132] 2. Nitrogen atom

[0133] If the currently traversed atom is a nitrogen atom, similar to the way of processing carbon atoms, first check whether the number of atoms connected to the nitrogen atom exceeds the given range, which will not be repeated here. Then, check the atomic degree of the nitrogen atom. If its degree is less than 3, find an atom that can form a single bond with it from the set of atoms to be selected, add atoms and bonds in sequence and store them; if its degree is less than 2, find an atom that can form a double bond with it from the set of atoms to be selected, add atoms and bonds in sequence and store them. It should be noted that at this stage, the nitrogen-nitrogen triple bond cannot be added. This is because the traversed nitrogen atom has at least one single bond connected to it in the original molecular graph to ensure the connectivity of the original molecular graph. If another triple bond is added, the nitrogen atom will be bonded to four atoms, which does not conform to the valence rules (more complex situations such as coordination bonds are not considered here).

[0134] (3) Oxygen atom

[0135] If the currently traversed atom is an oxygen atom, the first step is to check whether the number of atoms connected to the oxygen atom exceeds the given range. Next, the degree of the oxygen atom is checked to see if it is less than 2. Since oxygen has two lone pairs of electrons, it can usually only form bonds with two atoms. Therefore, only if this is the case, consider selecting an atom from the candidate set to form a single bond with the oxygen atom, and save the newly generated molecular graph to the corresponding graph set. Similar to nitrogen atoms, during the atom augmentation phase, oxygen atoms, when included in the base graph being traversed, cannot form double bonds with other atoms.

[0136] (4) Fluorine atom

[0137] If the fluorine atom is currently being traversed, you only need to consider that if this is the initial molecule, you can add an atom and a single bond, such as the fluorine gas molecule (F2). It is recommended to consider this situation when constructing the initial molecule so that these possible cases can be skipped directly.

[0138] After adding atoms and bonds, it is necessary to determine the graph isomorphism. This process should be performed before storing the molecular graph to reduce space complexity. In graph theory, graph isomorphism refers to a group of graphs with the same number of nodes and edges, and the nodes and edges have a bijective relationship. In topology, we believe that this group of graphs is the same. Specifically, you can call ready-made toolkits to determine graph isomorphism. Common graph isomorphism toolkits include NAUTY based on C / C++ language, Networkx based on Python language, etc. Graph isomorphism determination is a very time-consuming process, so we need to find a way to optimize this process. In addition to choosing a suitable and efficient algorithm, we can group and store the graphs, because only molecules with the same nodes, the same edges, and the same degree list may have graph isomorphism problems, so these three conditions can be used to make judgments in advance to reduce the large amount of time consumed by the isomorphism process. It should also be noted in this process that since non-planar graphs may appear in our generation process, we should also remove K5, K at this stage. 3,3 Graphs with subgraph isomorphism.

[0139] The four representative cases above respectively indicate that the atoms in the traversed base graph have 4, 3, 2, and 1 lone pairs of electrons, that is, they can form bonds with up to 4, 3, 2, and 1 atoms. This shows a certain degree of universality. If silicon atoms need to be considered during the construction process, only the steps for carbon atoms (C) need to be migrated to silicon (Si). If sulfur atoms (S) need to be considered, only the steps for oxygen atoms (O) need to be migrated. In other words, considering the similarity of atoms of the same main group elements allows for good algorithm migration, greatly increasing algorithm reuse and making the program simple and efficient. Therefore, the atom augmentation algorithm is universal for the generation process of various organic small molecules. It can be added to different sets of atoms of selected elements as needed, reflecting its good adaptability.

[0140] The bond augmentation module, after adding an atom and bond to a molecule, continues iteratively adding bonds until saturation is reached or further additions are prohibited due to given constraints. In the serial algorithm, the bond augmentation module is closely related to the atom augmentation module, meaning that its input is derived from its output. In each round, the bond augmentation module begins running as soon as the atom augmentation module provides an output. This ensures the efficiency of the serial algorithm and reduces waiting time between modules.

[0141] In general, the bond augmentation module in the serial algorithm adopts a breadth-first generation strategy and uses an auxiliary queue as its auxiliary data structure. The input of the module includes the molecular graph from the atom augmentation module, the current number of nodes, and the number of edges. For each input, an auxiliary queue is first constructed. Initially, the queue is empty. The input of the module, that is, a single molecular graph, is added to the queue, and a loop judgment is performed. When the queue is not empty, the following operations are performed: take a molecular graph from the queue (only one graph at a time), and then traverse all atoms on the graph. If there are two atoms that satisfy the valence of each atom is not saturated, then the two can add a bond.

[0142] The following example illustrates the execution process of the key augmentation module in detail.

[0143] 1. Carbon Atom – Carbon Atom

[0144] When the two atoms traversed are two carbon atoms, the premise is that there is no chemical bond between the two atoms, and both atoms can be connected to new atoms, that is, their valence bonds are not saturated. If the degrees of both carbon atoms are less than 4, the two carbon atoms can be connected with a single bond, and the result can be saved. It should be noted that, as described in the atomic augmentation section, it is recommended to copy the base graph from the loop to avoid superposition during the bond addition process. If the degrees of both carbon atoms are less than 3, and each carbon atom itself has no other double bonds connected, a double bond can be added between the two carbon atoms. If the degrees of both carbon atoms are less than 2, a triple bond can be added between the two carbon atoms. It is particularly emphasized that in the process of atomic augmentation or bond augmentation, all operations are serial. If the degree meets the requirements, the next step needs to continue to be judged instead of directly jumping out of the current round. Specifically, when the degree of both carbon atoms is 2 and there is no double bond connected to them, during the bond augmentation process, the degree is first determined to be less than 4. At this time, the base graph is copied and a new single bond is added between the two carbon atoms. After isomorphism judgment, it is saved in the graph collection. However, this does not mean that this round of the loop is over. It is necessary to then determine whether the degree of both carbon atoms is less than 3 and there is no double bond connected to them. At this time, the base graph is copied again and a double bond is added between the two carbon atoms. After isomorphism judgment, it is also saved. Finally, it is determined whether the degree of both carbon atoms is less than 2. If it is found that the condition is not met, the loop is jumped out and the judgment of other atom pairs is carried out.

[0145] 2. Carbon atom–nitrogen atom

[0146] When the two atoms traversed are carbon and nitrogen atoms, if the degree of the carbon atom is less than 4 and the degree of the nitrogen atom is less than 3, a single bond can be added between them. If the degree of the carbon atom is less than 3 and the degree of the nitrogen atom is less than 2, a double bond can be added between them. As mentioned above, since the two atoms are connected by at least one chemical bond to ensure the connectivity of the molecular graph during the bond expansion process, considering the valence bond rules, it is impossible to add an additional triple bond to the nitrogen atom (not considering complex situations such as coordination bonds). When generating each new molecular graph, a graph isomorphism judgment should be performed and then stored in the graph collection.

[0147] The bonding methods between other atoms are similar to the above cases and will not be discussed one by one here.

[0148] The format conversion module converts molecules represented in graph theory to the universal SMILES representation. This conversion makes molecules more accessible, facilitates better understanding for chemists, and facilitates database addition, deletion, modification, and querying. Furthermore, many chemistry toolkits, such as RDKit and Openbabel, facilitate molecular manipulation based on the SMILES format.

[0149] In the execution process of the format conversion module, first, it is necessary to define all possible atoms in the database and store them in a sequence. Secondly, define all possible bond types, such as single bonds, double bonds, triple bonds, aromatic bonds, etc. Then, read the molecular graph, extract all atoms in the molecular graph in sequence, and convert the adjacent relationship between the atoms in the molecules into an adjacency matrix. According to the numbering in the form of the graph, it can be directly corresponded to the previously predefined atoms and bonds, and the atomic information and bond information are stored in the form of an atom list and an adjacency matrix. Finally, the SMILES format is standardized using the toolkit to ensure that the molecular formula obtained in the SMILES format is unique and there is no ambiguity. This is a very important step in the database construction process, which ensures the consistency and accuracy of the data.

[0150] The auxiliary operation module may include a degree judgment function, a file reading and writing function, a counter and a progress bar, etc.

[0151] The degree determination function can be used to determine the degree of a given atomic node, which differs slightly from degree determination in traditional graph theory. In molecular graphs, edges represent chemical bonds, including single, double, and triple bonds. Therefore, graph edges contain bond information, or attributes. Due to valence rules, the maximum number of molecules an atom can connect to is determined by the number of lone pairs of electrons it carries. Therefore, when calculating the degree of an atomic node, the weighted values ​​of the edge attributes must be considered. When constructing the electrolyte database, hydrogen atoms are added at the final stage, so their influence can be ignored. For example, a carbon atom already forms a double bond and a single bond with another atom, giving it a degree of 3, meaning it can also connect to another atom via a single bond. This simplifies the calculation process by considering only the bonding of other atoms, without factoring in the influence of hydrogen atoms.

[0152] File read / write functions can be used to handle situations where the initial set of molecules is large or the generated molecules contain a large number of atoms. In such cases, relying solely on in-memory algorithms may lead to memory overflow, necessitating the placement of the generated molecules in external storage for auxiliary storage. It is important to note that because input / output (I / O) operations are time-consuming when a program accesses data from external storage, in-memory algorithms should be prioritized over file read / write functions if the number of molecules generated by the program can be stored in memory. This approach improves program efficiency while ensuring data integrity.

[0153] A counter can be used to record the number of molecules generated. When the program finishes executing, the number of molecules generated in this round needs to be reported. The program can hierarchically count the number of molecules containing a specified number of atoms, as well as the total number of molecules generated. This statistics not only helps to understand the scale of the generated problem but also allows prediction of the runtime of the next round of the program based on the generated molecule data. This approach allows for more efficient planning of computing resources and optimizes program efficiency.

[0154] A progress bar can be used to provide visual feedback when a program is generating molecules with a large number of atoms. If a program has no on-screen output for an extended period, it can be difficult to determine whether it is running normally or has stalled due to an exception. Using a progress bar can help visualize the program's progress, provide real-time insights into its execution status, and provide a rough estimate of its runtime. This not only improves the user experience but also allows for the timely identification and resolution of potential issues, ensuring smooth program operation.

[0155] Next, a method for generating molecular structure data based on a serial molecule generation algorithm is introduced. In one embodiment, the preset molecule generation rule includes executing the following steps for each initial molecule:

[0156] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of first molecules.

[0157] For each of the plurality of first molecules, a preset type of chemical bond is added to the first molecule according to a chemical bond addition rule, a valence rule, and a molecular stability rule to obtain a plurality of second molecules.

[0158] The plurality of first molecules and the plurality of second molecules are taken as a plurality of primary target molecules and stored in a preset target molecule set.

[0159] For each of the plurality of primary target molecules, a preset number of preset atoms is added to the primary target molecule according to the atom addition rule, the valence rule, and the molecular stability rule, thereby obtaining a plurality of third molecules. For each of the plurality of third molecules, a preset type of chemical bond is added to the third molecule according to the chemical bond addition rule, the valence rule, and the molecular stability rule, thereby obtaining a plurality of fourth molecules. The plurality of third molecules and the plurality of fourth molecules are stored as a plurality of secondary target molecules in a preset target molecule set.

[0160] Taking multiple secondary target molecules as new multiple primary target molecules, returning to each of the multiple primary target molecules, respectively, according to the atom addition rule, valence rule and molecular stability rule, adding a preset number of preset atoms to the primary target molecule to obtain multiple third molecules until the preset iteration stop condition is met.

[0161] In the above embodiment, the iteration stopping condition may include the number of atoms in the generated secondary target molecules reaching a preset number, the number of iterations reaching a preset number, the number of target molecules in the target molecule set reaching a preset value, etc. Those skilled in the art may select an appropriate iteration stopping condition based on actual needs, and the condition is not limited here.

[0162] According to the above implementation, a molecule can be amplified atomically, and then chemically amplified based on the results of the atomic amplification. The target molecule for the current round can then be obtained based on the results of the atomic and chemical amplification. In this way, through multiple rounds of iteration, a large amount of target molecular structure data can be efficiently generated.

[0163] In one embodiment, the molecular stability rules include molecular structure patterns that do not comply with the molecular stability rules.

[0164] According to the atomic addition rule, the valence rule, and the molecular stability rule, adding a preset number of preset atoms to the initial molecule to obtain a plurality of first molecules may specifically include:

[0165] A preset number of preset atoms are added to the initial molecule according to the atom addition rule and the valence rule to obtain a plurality of initial first molecules.

[0166] A molecular structure diagram is found among a plurality of initial first molecules.

[0167] The initial first molecule including the molecular structure graph is deleted to obtain a plurality of first molecules.

[0168] According to the above embodiment, the initial first molecule containing an unstable molecular structure can be determined through subgraph isomorphism based on a graph search method, and then deleted. In this way, unstable molecules generated during the molecular amplification process can be deleted immediately, thereby improving the quality of the target molecular structure data.

[0169] In one embodiment, the initial molecule includes at least one non-hydrogen atom.

[0170] Adding a preset number of preset atoms to the initial molecule according to the atom addition rule and the valence rule to obtain a plurality of initial first molecules may specifically include performing the following steps on the initial molecule for each preset atom:

[0171] Get the number of preset atoms in the initial molecule.

[0172] When the number of preset atoms is less than a preset threshold, the following steps are performed for each non-hydrogen atom in the initial molecule: the atomic degree of the non-hydrogen atom is calculated. The atomic degree includes the sum of the weighted values ​​of each edge connected to the corresponding node of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attribute. The edge attribute is used to represent the properties of the chemical bond corresponding to the edge. Based on the correspondence between the preset atomic degree and the chemical bond type, the first chemical bond type corresponding to the atomic degree of the non-hydrogen atom is determined. The non-hydrogen atom and the preset atom are connected according to the first chemical bond type to obtain the initial first molecule.

[0173] In the above embodiment, the correspondence between the preset atomic degree and the chemical bond type can be set according to actual needs, and can be specifically set in combination with actual needs and valence rules. For example, when the non-hydrogen atom is a C atom and the preset atoms are all C atoms, if the atomic degree of the C atom satisfies less than 4, then the corresponding first chemical bond type can be a single bond; if the atomic degree of the C atom satisfies less than 3, then the corresponding first chemical bond type can be a double bond; if the atomic degree of the C atom satisfies less than 2, then the corresponding first chemical bond type can be a triple bond. For example, when the atomic degree of the C atom is 2, the corresponding first chemical bond type can be a single bond or a double bond, then a new C atom can be connected to the C atom through a single bond and a double bond, respectively, to obtain two initial first molecules.

[0174] According to the above embodiment, a suitable bonding method can be selected based on the number of non-hydrogen atoms in the molecule and the correspondence between the preset atomic number and the chemical bond type, which is conducive to improving the chemical rationality of the initial first molecule.

[0175] In one embodiment, after the non-hydrogen atom and the predetermined atom are connected according to the first chemical bond type to obtain an initial first molecule, the method may further include:

[0176] The graph isomorphism between the initial first molecule and other initial first molecules is judged at the molecular graph level.

[0177] If there are other initial first molecules that are isomorphic to the initial first molecule graph, the initial first molecule is deleted.

[0178] According to the above embodiment, after the initial first molecule is generated, graph isomorphism can be promptly determined, thereby deleting duplicate molecules. In this way, the uniqueness of the generated initial first molecule can be guaranteed, thereby facilitating the conciseness and uniqueness of the molecular structure data.

[0179] In one embodiment, for each of the plurality of first molecules, adding a predetermined type of chemical bond to the first molecule according to a chemical bond addition rule, a valence rule, and a molecular stability rule to obtain a plurality of second molecules may specifically include:

[0180] For each of the plurality of first molecules, a preset type of chemical bond is added to the first molecule according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules.

[0181] Finding a molecular structure diagram among a plurality of initial second molecules.

[0182] The initial second molecule including the molecular structure graph is deleted to obtain a plurality of second molecules.

[0183] According to the above embodiment, the initial second molecule containing an unstable molecular structure can be determined through subgraph isomorphism based on a graph search method, and then deleted. In this way, unstable molecules generated during the molecular amplification process can be deleted immediately, thereby improving the quality of the target molecular structure data.

[0184] In one embodiment, the first molecule includes at least two non-hydrogen atoms.

[0185] For each of the plurality of first molecules, adding a preset type of chemical bond to the first molecule according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules may specifically include performing the following steps for each of the plurality of first molecules:

[0186] For any two non-hydrogen atoms in the first molecule, the target atomic degrees and target edge attributes of the two non-hydrogen atoms are obtained respectively.

[0187] According to the preset correspondence between the atomic degrees of the two non-hydrogen atoms and the edge attributes and the chemical bond type, the second chemical bond type corresponding to the target atomic degrees of the two non-hydrogen atoms and the target edge attributes is determined.

[0188] The two non-hydrogen atoms are connected according to the second chemical bond type to obtain an initial second molecule.

[0189] In the above embodiment, the atomic degrees of the two non-hydrogen atoms and the correspondence between the edge attributes and the chemical bond type can be set according to actual needs, and can be set in combination with actual needs and valence rules. For example, when the two non-hydrogen atoms in the first molecule are C atom-C atom, if there is no chemical bond between the two C atoms, and both C atoms can be connected to new atoms, that is, their valence bonds are not saturated, then a new chemical bond can be added between the two C atoms. If the atomic degrees of the two C atoms are less than 4, then the corresponding second chemical bond type can be a single bond; if the atomic degrees of the two C atoms are both less than 3, and each C atom itself has no other double bonds connected, then the corresponding second chemical bond type can be a double bond; if the atomic degrees of the two C atoms are both less than 2, then the corresponding second chemical bond type can be a triple bond. For example, when the atomic degrees of the two C atoms are both 1, the corresponding second chemical bond type can be a single bond, a double bond or a triple bond, then the two C atoms can be connected by a single bond, a double bond and a triple bond, respectively, to obtain three initial first molecules. The process of determining the type of second chemical bond that can be added between other atoms is similar to that of C atom-C atom and will not be described in detail here.

[0190] According to the above embodiment, a suitable bonding method can be selected based on the number of non-hydrogen atoms in the two molecules and the correspondence between the preset atomic number and the chemical bond type, which is conducive to improving the chemical rationality of the initial second molecule.

[0191] In one embodiment, after the two non-hydrogen atoms are connected according to the second chemical bond type to obtain an initial second molecule, the method may further include:

[0192] A graph isomorphism judgment is performed on the initial second molecule, the multiple initial first molecules, and other initial second molecules at the molecular graph level.

[0193] When there is an initial first molecule or other initial second molecule that is graph-isomorphic to the initial second molecule, the initial second molecule is deleted.

[0194] According to the above embodiment, after the initial second molecule is generated, graph isomorphism can be promptly determined, thereby deleting duplicate molecules. In this way, the uniqueness of the generated initial second molecule can be guaranteed, thereby facilitating the maintenance of the simplicity and uniqueness of the molecular structure data.

[0195] In one embodiment, molecular structure data can be generated based on a molecular incremental generation algorithm. This solution can be implemented based on a molecular incremental generation algorithm system. For ease of understanding, the molecular incremental generation algorithm system is briefly introduced below.

[0196] The incremental molecule generation algorithm system can include an initial molecule generation module, a graph constraint module, an atomic augmentation module, a bond augmentation module, a graph isomorphism determination module, a format conversion module, and an auxiliary operation module. In the incremental molecule generation algorithm system, the functions of the initial molecule generation module, the graph constraint module, the atomic augmentation module, the bond augmentation module, the format conversion module, and the auxiliary operation module are similar to those of similar modules in the serial generation algorithm, with the only difference being that the atomic augmentation module and the bond augmentation module no longer perform the graph isomorphism determination during molecule amplification; this is performed by the graph isomorphism determination module, and the input of the bond augmentation module is changed.

[0197] Specifically, the bond augmentation module generates the molecular graph of the current round based on the graph with the minimum number of edges generated in the previous round. The number of edges here is different from the weighted number of bonds. The bond is an attribute attached to the edge, that is, an edge can be a single bond, a double bond, or a triple bond, but its number of edges is only 1. In other words, the number of edges here is the node degree. The rationality and completeness of this generation method lies in that according to the atomic augmentation and molecular augmentation strategies proposed in this algorithm, any molecular graph can be generated from its subgraph. Therefore, when constructing an organic small molecule database, it is only necessary to find the minimum subgraph set of the pre-generated molecular graph, that is, a set of graphs with the minimum number of edges in the molecular graph generated in the previous round. In this way, the probability of generating graph isomorphic molecules during bond augmentation can be reduced, thereby greatly saving the time for graph isomorphism judgment.

[0198] The molecule incremental generation algorithm may include a step-by-step merging-based incremental generation algorithm and a multi-grouping-based incremental generation algorithm.

[0199] The key difference between the incremental generation algorithm based on step-by-step merging and the serial molecular graph generation algorithm lies in the abstraction of the graph isomorphism determination module. A parallel, step-by-step merging algorithm replaces the isomorphism determination process previously coupled between the atomic augmentation and bond augmentation modules. This abstraction enhances code reusability and the compartmentalization of module functions. Furthermore, it decouples the atomic augmentation and bond augmentation modules, adhering to the principle of "high cohesion and low coupling" to minimize coupling between modules and enhance connections within modules, thereby improving algorithm efficiency.

[0200] In this incremental molecular generation algorithm based on step-by-step merging, the parallelized atomic augmentation module performs multi-core computations. The module input is a set of molecular graphs with the fewest edges from the previous round. For example, if the number of nodes in the pre-generated molecular graph for this round is n, the molecular graph input to the atomic augmentation module must have n-1 nodes and n-2 edges. The output of the atomic augmentation module is a molecular graph with one additional node (atom) and one additional edge (bond, which could be a single, double, or triple bond).

[0201] The parallelized bond augmentation module sequentially adds edges to the molecular graph after atomic augmentation. Unlike the serialized algorithm, the parallelized bond augmentation module decouples atomic augmentation and bond augmentation. In the parallelized bond augmentation module, the input of the bond augmentation module comes from the entire structure of atomic augmentation, rather than being coupled to the atomic augmentation process. The output of the bond augmentation module is a set of molecular graphs with increased number of edges.

[0202] In the incremental molecular generation algorithm system based on step-by-step merging, the graph isomorphism determination module adopts a "step-by-step merging" approach compared to the serial algorithm, effectively reducing the time consumption associated with isomorphism determination. The graph isomorphism module follows the atomic augmentation module and the bond augmentation module. Since each process in the parallel algorithm is independent of each other, although the graphs generated by each process are not isomorphic, the graphs between processes may be isomorphic. Therefore, the graph isomorphism determination module takes as input the molecular graphs generated by the atomic augmentation module and the bond augmentation module, and outputs a molecular graph that does not contain any isomorphic graphs.

[0203] When determining isomorphism, if the input module has too few graphs (less than twice the number of cores on the computer), there's no need to start the gradual merging algorithm. This is because starting the process also incurs a certain amount of time overhead. When the number of graphs is small, directly determining isomorphism is inherently faster, eliminating the need for complex algorithms and parallel processing. In this case, simply aggregate the graphs from the atomic or bond augmentation modules and determine isomorphism together. Due to the small number of graphs, this process can be completed relatively quickly.

[0204] In an incremental generation algorithm based on step-by-step merging, graph isomorphism determination consumes the vast majority of the algorithm's total time. Therefore, an effective way to accelerate the algorithm is to reduce the number of graph isomorphism determinations or to reduce the number of unnecessary graphs generated. Generating a graph based on the minimum number of edges effectively avoids the generation of redundant graphs during the parallelization process due to the large number of edges in the previous round, which requires graph isomorphism determination to delete them, thus saving the time cost of program execution.

[0205] A multi-grouping-based incremental molecular generation algorithm system is based on the discovery that multi-grouping can effectively reduce the time in the graph isomorphism process. It aims to use the idea of ​​multi-grouping to reconstruct the graph isomorphism judgment module in the gradual merging algorithm, and replace the graph storage with serial numbers to reduce space complexity. The main ideas of other modules are basically the same.

[0206] The graph isomorphism determination module of the multi-grouping molecular incremental generation algorithm system first needs to be grouped pairwise. This process can be considered as a merging rate of 2. However, unlike the stepwise merging algorithm, the multi-grouping algorithm requires pairwise matching of all graph sets. For example, if 24 graph sets are generated after atomic or bond augmentation, they are labeled. Assuming a supercomputer with 24 cores and 24 child processes are created, the number of groups required in the multi-grouping algorithm is (24 × 23) / 2 = 276. Each group is labeled by the sequence number of the 24 graph sets it contains, with smaller sequence numbers first and larger sequence numbers last. Afterwards, each graph set needs to be tested for isomorphism. Assuming the system can create a maximum of 24 child processes simultaneously, a maximum of 24 graph sets can be tested for isomorphism in each round, resulting in a total of 12 rounds. During each round of judgment, the order in which isomorphic graphs appear in the sets with the earlier sequence numbers is recorded, rather than directly copying the generated graphs, to achieve in-situ elimination of isomorphic graphs. After 12 rounds of judgment, the intersection of all isomorphic sequence groups containing the same previous sequence number is taken. The resulting set is a set of all graphs in each of the 24 graph sets that may be isomorphic with other sets. The graphs in the sets with smaller sequence numbers are retained, and the graphs in the sets with larger sequence numbers are deleted. Finally, the graphs with the corresponding graph isomorphism numbers in the 24 graph sets are deleted in turn, and what remains are all non-isomorphic molecular graphs. This method can effectively reduce the number of graph isomorphism judgments and the unnecessary generation of graphs, thereby improving the overall efficiency of the algorithm.

[0207] For the database of fluorine-containing small molecules in electrolytes, the rule is that the initial atom is an oxygen atom, that is, each molecule must contain an oxygen atom. During the construction process, the atoms that can be added in the atom augmentation stage include carbon, oxygen, and fluorine atoms, and the bond augmentation stage can add carbon-carbon single bonds, carbon-oxygen single bonds, carbon-carbon double bonds, carbon-oxygen double bonds, and carbon-carbon triple bonds. At the same time, the following conditions need to be met: oxygen does not form bonds with oxygen or fluorine (no peroxide bonds are generated), two adjacent bonds cannot be double bonds at the same time, the total number of carbon atoms in a single molecule does not exceed 8, and the total number of oxygen atoms does not exceed 4.

[0208] In the actual molecular generation process, it can be observed that the number of graphs increases exponentially, and the time consumed also shows an exponential upward trend. For example, for generating a graph with 7 nodes, the incremental generation algorithm based on step-by-step merging reduces the runtime by 2 times compared to the serial generation algorithm, while the incremental generation algorithm based on multiple groups reduces the runtime by 7.6 times. For generating a graph with 8 nodes, the incremental generation algorithm based on step-by-step merging takes longer due to the high number of I / O operations involved in the design process. Comparing the serial generation algorithm with the incremental generation algorithm based on multiple groups, the latter algorithm reduces the runtime by 8.8 times. This shows that the incremental generation algorithm based on multiple groups can significantly accelerate the database generation process and save a considerable amount of time. Using this algorithm, a total of 402,830 organic small molecules for electrolytes with up to 9 atoms were generated, meeting the demand for small molecules.

[0209] Next, the molecular structure data generation method based on the molecular incremental generation algorithm is introduced.

[0210] In one embodiment, the preset molecule generation rule includes executing the following steps for each initial molecule:

[0211] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of fifth molecules.

[0212] The plurality of fifth molecules are stored in a preset target molecule set.

[0213] A target fifth molecule is determined from the plurality of fifth molecules, where the target fifth molecule is the fifth molecule having the least number of edges among the plurality of fifth molecules.

[0214] According to the atomic addition rule, the valence rule and the molecular stability rule, a preset number of preset atoms are added to the target fifth molecule to obtain a plurality of sixth molecules.

[0215] The plurality of sixth molecules are stored in a preset target molecule set.

[0216] The plurality of sixth molecules are used as new plurality of fifth molecules, and the target fifth molecule is determined from the plurality of fifth molecules until a preset iteration stop condition is satisfied.

[0217] According to the chemical bond addition rule, the valence rule and the molecular stability rule, chemical bonds of preset types are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of seventh molecules.

[0218] The plurality of seventh molecules are stored in a preset target molecule set.

[0219] According to the preset graph isomorphism merging rule, the molecules with graph isomorphism in the preset target molecule set are merged.

[0220] In the above embodiment, the target fifth molecule is the fifth molecule with the least number of edges among the multiple fifth molecules. It is understandable that there can be multiple target fifth molecules. When there are multiple target fifth molecules, the corresponding steps in the above embodiment can be performed on each target fifth molecule separately.

[0221] In the above embodiment, according to the preset graph isomorphism merging rule, the preset target molecule set contains molecules with graph isomorphism, and the same molecules can be merged. Thus, there are no duplicate molecules in the target molecule set, and the molecules in the target molecule set are the target molecules.

[0222] In the above embodiment, the iteration stopping condition may include that the number of atoms in the generated sixth molecule reaches a preset number, or that the number of iterations reaches a preset number, or that the number of molecules in the target molecule set reaches a preset value, etc. Those skilled in the art may select an appropriate iteration stopping condition based on actual needs, and the condition is not limited here.

[0223] According to the above implementation, atomic expansion is first performed on the molecule with the smallest number of edges generated in the previous round, followed by bond expansion. Generating a molecular graph based on the minimum number of edges effectively reduces the risk of redundant graphs generated from the previous round's larger number of edges, which would require graph isomorphism determination and deletion, during the parallelization process. This saves time and effort during program execution, improving the efficiency of molecular structure data generation.

[0224] In one embodiment, the graph isomorphism merging rule includes a graph isomorphism merging rule based on a step-by-step merging algorithm. The graph isomorphism merging rule based on the step-by-step merging algorithm has been described above in conjunction with the system and will not be repeated here.

[0225] According to the above implementation, the efficiency of graph isomorphism determination and merging can be improved, thereby improving the efficiency of molecular structure data generation.

[0226] In one embodiment, the graph isomorphism merging rule includes a graph isomorphism merging rule based on a multi-grouping algorithm. The graph isomorphism merging rule based on a multi-grouping algorithm has been described above in conjunction with the system and will not be repeated here.

[0227] According to the above implementation, the efficiency of graph isomorphism determination and merging can be improved, thereby improving the efficiency of molecular structure data generation.

[0228] In one embodiment, the molecular stability rules may include molecular structure patterns that do not conform to the molecular stability rules.

[0229] According to the atomic addition rule, the valence rule, and the molecular stability rule, a preset number of preset atoms are added to the initial molecule to obtain a plurality of fifth molecules, which may specifically include:

[0230] A preset number of preset atoms are added to the initial molecule according to the atom addition rule and the valence rule to obtain a plurality of initial fifth molecules.

[0231] Find molecular structure diagrams among multiple initial fifth molecules.

[0232] The initial fifth molecule including the molecular structure diagram is deleted to obtain multiple fifth molecules.

[0233] According to the above embodiment, the initial fifth molecule containing an unstable molecular structure can be determined through subgraph isomorphism based on a graph search method, and then deleted. In this way, unstable molecules generated during the molecular amplification process can be deleted immediately, thereby improving the quality of the target molecular structure data.

[0234] In one embodiment, the initial molecule includes at least one non-hydrogen atom.

[0235] Adding a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain a plurality of initial fifth molecules may specifically include:

[0236] Get the number of preset atoms in the initial molecule.

[0237] If the number of preset atoms is less than a preset threshold, the following steps are performed for each non-hydrogen atom in the initial molecule: the atomic degree of the non-hydrogen atom is calculated. The atomic degree includes the sum of the weighted values ​​of each edge connected to the corresponding node of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attribute. The edge attribute is used to represent the properties of the chemical bond corresponding to the edge. Based on the preset correspondence between atomic degree and chemical bond type, the third chemical bond type corresponding to the atomic degree of the non-hydrogen atom is determined. The non-hydrogen atom and the preset atom are connected according to the third chemical bond type to obtain an initial fifth molecule.

[0238] In the above embodiment, the method for determining the third chemical bond type is similar to the method for determining the first and third chemical bond types, and is not described in detail here.

[0239] According to the above embodiment, a suitable bonding method can be selected based on the number of non-hydrogen atoms in the molecule and the correspondence between the predetermined number of atoms and the type of chemical bond, which is beneficial to improving the chemical rationality of the initial fifth molecule.

[0240] In one embodiment, according to the chemical bond addition rule, the valence rule, and the molecular stability rule, adding a preset type of chemical bond to each of the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain the plurality of seventh molecules may specifically include:

[0241] According to the chemical bond addition rule and the valence rule, chemical bonds of preset types are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of initial seventh molecules.

[0242] Find molecular structure diagrams among multiple initial seventh molecules.

[0243] The initial seventh molecule including the molecular structure diagram is deleted to obtain multiple seventh molecules.

[0244] According to the above embodiment, the initial seventh molecule containing an unstable molecular structure can be determined through subgraph isomorphism based on a graph search method, and then deleted. In this way, unstable molecules generated during the molecular amplification process can be deleted immediately, thereby improving the quality of the target molecular structure data.

[0245] In one embodiment, the fifth molecule and the sixth molecule each include at least two non-hydrogen atoms.

[0246] According to the chemical bond addition rule and the valence rule, chemical bonds of a preset type are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain the plurality of initial seventh molecules. Specifically, the steps may include performing the following steps on each of the plurality of fifth molecules and the plurality of sixth molecules:

[0247] For any two non-hydrogen atoms in a molecule, obtain the target atomic degree and target edge attributes of each of the two non-hydrogen atoms.

[0248] According to the preset correspondence between the atomic degrees of the two non-hydrogen atoms and the edge attributes and the chemical bond types, the fourth chemical bond type corresponding to the target atomic degrees of the two non-hydrogen atoms and the target edge attributes is determined.

[0249] Two non-hydrogen atoms are connected according to the fourth chemical bond type to obtain multiple initial seventh molecules.

[0250] In the above embodiment, the method for determining the fourth chemical bond type is similar to the method for determining the second chemical bond type, and is not described in detail here.

[0251] According to the above embodiment, a suitable bonding method can be selected based on the number of non-hydrogen atoms in the two molecules and the correspondence between the predetermined atomic number and the chemical bond type, which is conducive to improving the chemical rationality of the initial seventh molecule.

[0252] In one embodiment, the molecular stability rules include molecular structure patterns that do not comply with the molecular stability rules.

[0253] Before adding the preset atoms and / or preset chemical bonds to the initial molecule according to the preset molecule generation rule, the method may further include:

[0254] An application environment for obtaining structural data of multiple target molecules.

[0255] According to the chemical stability rules and the application environment, determine the molecular structure diagram that does not meet the molecular stability requirements under the application environment.

[0256] In the above embodiments, the application environment may include the environment corresponding to the research or application system. Taking electrolyte research as an example, the application environment may include the type of battery system and the type of electrolyte system.

[0257] According to the above embodiment, molecular stability rules are constructed based on chemical stability rules and the application environment, which helps to screen out molecules that are unstable in the application environment, thereby retaining target molecules with higher stability. In this way, the target molecular structure data can be made to meet the needs of actual applications, thereby improving the flexibility of molecular structure data generation and the quality of molecular structure data.

[0258] In one embodiment, the predetermined atoms may include at least one of carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl), thereby meeting the needs of most electrolyte systems.

[0259] In one embodiment, converting multiple target molecules into SMILES format to obtain multiple target molecule structure data may specifically include:

[0260] Convert multiple target molecules into adjacency matrix or adjacency table format to obtain intermediate data of target molecule structure.

[0261] The target molecular structure intermediate data is converted into SMILES format to obtain multiple target molecular structure data.

[0262] According to the above embodiment, the multiple target molecules are converted into an adjacency matrix or adjacency table format, which is convenient for computer storage and format conversion. In this way, the intermediate data of the target molecule structure can be efficiently converted into the SMILES format.

[0263] In one embodiment, after converting the plurality of target molecules into SMILES format to obtain the plurality of target molecule structure data, the method may further include:

[0264] A target molecular structure database is created based on multiple target molecular structure data.

[0265] According to the above embodiment, a target molecular structure database is created based on multiple target molecular structure data. When combined with high-throughput computing, high-throughput screening and machine learning technologies, this database is conducive to the high-speed research and development of electrolytes.

[0266] Based on the same inventive concept as the molecular structure data generating method, an embodiment of the present application also provides a molecular structure data generating device.

[0267] like Figure 3 As shown, the molecular structure data generating device 200 may include a first acquiring module 201 , an adding module 202 and a converting module 203 .

[0268] The first acquisition module 201 is used to acquire an initial molecule. The initial molecule is represented by an undirected graph without self-loops. The undirected graph includes nodes for representing atoms and edges for representing chemical bonds.

[0269] The adding module 202 is used to add preset atoms and / or preset chemical bonds to the initial molecule according to the preset molecular generation rules to obtain multiple target molecules, wherein the molecular generation rules include atom addition rules, chemical bond addition rules, valence rules and molecular stability rules, and the target molecules are represented by an undirected graph without self-loops.

[0270] The conversion module 203 is used to convert the multiple target molecules into SMILES format to obtain multiple target molecule structure data.

[0271] The molecular structure data generation device of the embodiment of the present application can obtain an initial molecule, which is represented by an undirected graph without self-loops. The undirected graph includes nodes for representing atoms and edges for representing chemical bonds. In this way, by using graph theory to represent molecules, it is possible to ensure the structuring and standardization of molecular data during the molecule generation process, thereby improving the efficiency of data processing. Then, according to the preset molecule generation rules, preset atoms and / or preset chemical bonds can be added to the initial molecule to obtain multiple target molecules, wherein the molecule generation rules include atom addition rules, chemical bond addition rules, valence rules and molecular stability rules, and the target molecule is represented by an undirected graph without self-loops. On the one hand, the molecule generation rules can make the structure of the generated target molecule chemically reasonable and stable, reducing the risk of generating unreasonable molecules; on the other hand, by specifying atom addition rules and chemical bond addition rules, target molecules that meet user needs can be generated in batches, thereby improving the efficiency and flexibility of molecule generation. In this way, the generated multiple target molecules can adapt to the needs of different research scenarios. Subsequently, the multiple target molecules can also be converted into SMILES format to obtain multiple target molecular structure data. The SMILES format is a more standardized molecular representation, which is convenient for retrieval and storage, and does not have ambiguity. By converting multiple target molecules into the SMILES format, the consistency and accuracy of the target molecule structure data can be guaranteed. In this way, a solid data foundation can be provided for high-throughput computing, machine learning and high-throughput screening. The target molecule structure data generated in batches in the application embodiment are used to construct an electrolyte molecular structure database, which can effectively support the rapid development of new electrolyte systems, significantly shorten the development cycle, reduce development costs, and improve the efficiency of electrolyte design and optimization.

[0272] In one embodiment, the apparatus may further include a first molecule generation rule module, which may include:

[0273] The first adding submodule is used to add a preset number of preset atoms to each initial molecule according to the atom adding rule, the valence rule and the molecular stability rule to obtain multiple first molecules.

[0274] The first adding submodule is further used to add a preset type of chemical bond to each first molecule in accordance with the chemical bond addition rule, valence rule and molecular stability rule to obtain a plurality of second molecules.

[0275] The first storage submodule is used to store the plurality of first molecules and the plurality of second molecules as a plurality of primary target molecules into a preset target molecule set.

[0276] The processing submodule is configured to, for each of the plurality of primary target molecules, add a preset number of preset atoms to the primary target molecule according to the atom addition rule, the valence rule, and the molecular stability rule, to obtain a plurality of third molecules. For each of the plurality of third molecules, add a preset type of chemical bond to the third molecule according to the chemical bond addition rule, the valence rule, and the molecular stability rule, to obtain a plurality of fourth molecules. The plurality of third molecules and the plurality of fourth molecules are stored as a plurality of secondary target molecules in a preset target molecule set.

[0277] The first return submodule is used to take the multiple secondary target molecules as new multiple primary target molecules, return each of the multiple primary target molecules, and add a preset number of preset atoms to the primary target molecule according to the atom addition rule, valence rule and molecular stability rule to obtain multiple third molecules until the preset iteration stop condition is met.

[0278] In one embodiment, the molecular stability rules may include molecular structure patterns that do not conform to the molecular stability rules.

[0279] The first adding submodule is used to add a preset number of preset atoms to the initial molecule according to the atomic addition rule, the valence rule, and the molecular stability rule to obtain a plurality of first molecules, which may specifically include:

[0280] The first adding unit is used to add a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain a plurality of initial first molecules.

[0281] The first searching unit is used to search for a molecular structure diagram in a plurality of initial first molecules.

[0282] The first deleting unit is used to delete the initial first molecule including the molecular structure diagram to obtain multiple first molecules.

[0283] In one embodiment, the initial molecule may include at least one non-hydrogen atom.

[0284] The first adding unit is used to add a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain multiple initial first molecules. Specifically, it may include performing the following steps on the initial molecules for each preset atom through each subunit in the first adding unit:

[0285] The first acquisition subunit is used to acquire the number of preset atoms in the initial molecule.

[0286] The first execution subunit is configured to, when the number of preset atoms is less than a preset threshold, perform the following steps for each non-hydrogen atom in the initial molecule: calculate the atomic degree of the non-hydrogen atom, where the atomic degree comprises the sum of the weighted values ​​of each edge connected to the corresponding node of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attribute, where the edge attribute represents the properties of the chemical bond corresponding to the edge. Based on the correspondence between the preset atomic degree and the chemical bond type, determine the first chemical bond type corresponding to the atomic degree of the non-hydrogen atom. Connect the non-hydrogen atom and the preset atom according to the first chemical bond type to obtain the initial first molecule.

[0287] In one embodiment, the apparatus may further include:

[0288] The first judgment module is used to connect non-hydrogen atoms and preset atoms according to the first chemical bond type to obtain an initial first molecule, and then perform graph isomorphism judgment on the initial first molecule and other initial first molecules at the molecular graph level.

[0289] The first deleting module is configured to delete the initial first molecule if there are other initial first molecules that are isomorphic to the initial first molecule graph.

[0290] In one embodiment, the first adding submodule is configured to add a preset type of chemical bond to each of the plurality of first molecules according to a chemical bond addition rule, a valence rule, and a molecular stability rule to obtain a plurality of second molecules, and may specifically include:

[0291] The second adding unit is used to add a preset type of chemical bond to each first molecule in accordance with a chemical bond adding rule and a valence rule to obtain a plurality of initial second molecules.

[0292] The second searching unit is used to search for the molecular structure diagram in the plurality of initial second molecules.

[0293] The second deleting unit is used to delete the initial second molecule including the molecular structure diagram to obtain multiple second molecules.

[0294] In one embodiment, the first molecule may include at least two non-hydrogen atoms.

[0295] The second adding unit is configured to add a preset type of chemical bond to each of the plurality of first molecules according to a chemical bond addition rule and a valence rule to obtain a plurality of initial second molecules. Specifically, the second adding unit may include performing the following steps on each of the plurality of first molecules through each subunit of the second adding unit:

[0296] The second acquisition subunit is used to respectively acquire target atomic degrees and target edge attributes of any two non-hydrogen atoms in the first molecule.

[0297] The first determination subunit is used to determine the second chemical bond type corresponding to the target atomic degree of each of the two non-hydrogen atoms and the target edge attribute according to the preset correspondence between the atomic degree of each of the two non-hydrogen atoms and the edge attribute and the chemical bond type.

[0298] The first linker unit is used to connect two non-hydrogen atoms according to a second chemical bond type to obtain an initial second molecule.

[0299] In one embodiment, the apparatus may further include:

[0300] The second judgment module is used to perform graph isomorphism judgment on the initial second molecule and multiple initial first molecules and other initial second molecules at the molecular graph level after connecting two non-hydrogen atoms according to the second chemical bond type to obtain the initial second molecule.

[0301] The second deleting module is configured to delete the initial second molecule if there is an initial first molecule or other initial second molecule that is isomorphic to the initial second molecule graph.

[0302] In one embodiment, the apparatus may further include a second molecule generation rule module, which may include:

[0303] The second adding submodule is used to add a preset number of preset atoms to the initial molecule according to the atomic addition rule, the valence rule and the molecular stability rule to obtain a plurality of fifth molecules.

[0304] The second storage submodule is used to store the plurality of fifth molecules into a preset target molecule set.

[0305] The determination submodule is used to determine a target fifth molecule from a plurality of fifth molecules, where the target fifth molecule is the fifth molecule having the least number of edges among the plurality of fifth molecules.

[0306] The second adding submodule is further used to add a preset number of preset atoms to the target fifth molecule according to the atomic addition rule, the valence rule and the molecular stability rule to obtain multiple sixth molecules.

[0307] The second storage submodule is further used to store the plurality of sixth molecules into a preset target molecule set.

[0308] The second returning submodule is configured to use the plurality of sixth molecules as new plurality of fifth molecules and return to determine a target fifth molecule from the plurality of fifth molecules until a preset iteration stopping condition is satisfied.

[0309] The second adding submodule is also used to add preset types of chemical bonds to multiple fifth molecules and multiple sixth molecules in the preset target molecule set according to chemical bond addition rules, valence rules and molecular stability rules to obtain multiple seventh molecules.

[0310] The second storage submodule is further configured to store the plurality of seventh molecules into a preset target molecule set.

[0311] The merging submodule is used to merge molecules with graph isomorphism in a preset target molecule set according to a preset graph isomorphism merging rule.

[0312] In one embodiment, the graph isomorphism merging rule may include a graph isomorphism merging rule based on a step-by-step merging algorithm.

[0313] In one embodiment, the graph isomorphism merging rule may include a graph isomorphism merging rule based on a multi-grouping algorithm.

[0314] In one embodiment, the molecular stability rules may include molecular structure patterns that do not conform to the molecular stability rules.

[0315] The second adding submodule is used to add a preset number of preset atoms to the initial molecule according to the atomic addition rule, the valence rule, and the molecular stability rule to obtain a plurality of fifth molecules, which may specifically include:

[0316] The third adding unit is used to add a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain a plurality of initial fifth molecules.

[0317] The third searching unit is configured to search for a molecular structure diagram in a plurality of initial fifth molecules.

[0318] The third deleting unit is used to delete the initial fifth molecule including the molecular structure diagram to obtain multiple fifth molecules.

[0319] In one embodiment, the initial molecule may include at least one non-hydrogen atom.

[0320] The third adding unit is used to add a preset number of preset atoms to the initial molecule according to the atom addition rule and the valence rule to obtain a plurality of initial fifth molecules, which may specifically include:

[0321] The third acquisition subunit is used to obtain the number of preset atoms in the initial molecule.

[0322] The second execution subunit is configured to, when the number of preset atoms is less than a preset threshold, perform the following steps for each non-hydrogen atom in the initial molecule: calculate the atomic degree of the non-hydrogen atom, where the atomic degree comprises the sum of the weighted values ​​of each edge connected to the corresponding node of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attribute, where the edge attribute represents the properties of the chemical bond corresponding to the edge. Based on the preset correspondence between atomic degree and chemical bond type, determine the third chemical bond type corresponding to the atomic degree of the non-hydrogen atom. Connect the non-hydrogen atom and the preset atom according to the third chemical bond type to obtain an initial fifth molecule.

[0323] In one embodiment, the second adding submodule is configured to add chemical bonds of a preset type to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set according to the chemical bond addition rule, the valence rule, and the molecular stability rule, respectively, to obtain the plurality of seventh molecules, and may specifically include:

[0324] The fourth adding unit is used to add preset types of chemical bonds to multiple fifth molecules and multiple sixth molecules in the preset target molecule set according to the chemical bond addition rule and the valence rule, so as to obtain multiple initial seventh molecules.

[0325] The fourth searching unit is used to search for the molecular structure diagram in the plurality of initial seventh molecules.

[0326] The fourth deleting unit is used to delete the initial seventh molecule including the molecular structure diagram to obtain multiple sixth molecules.

[0327] In one embodiment, the fifth molecule and the sixth molecule may each include at least two non-hydrogen atoms.

[0328] The fourth adding unit is configured to add a preset type of chemical bond to each of the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set according to the chemical bond addition rule and the valence rule, to obtain a plurality of initial seventh molecules. Specifically, the fourth adding unit may include calling the following subunits for processing for each of the plurality of fifth molecules and the plurality of sixth molecules:

[0329] The fourth acquisition subunit is used to obtain the target atomic degree and target edge attribute of each of any two non-hydrogen atoms in the molecule.

[0330] The second determination subunit is used to determine the fourth chemical bond type corresponding to the target atomic degree of each of the two non-hydrogen atoms and the target edge attribute according to the preset correspondence between the atomic degree of each of the two non-hydrogen atoms and the edge attribute and the chemical bond type.

[0331] The second linker unit is used to connect two non-hydrogen atoms according to the fourth chemical bond type to obtain multiple initial seventh molecules.

[0332] In one embodiment, the molecular stability rules may include molecular structure patterns that do not conform to the molecular stability rules.

[0333] The apparatus may further comprise:

[0334] The second acquisition module is used to acquire an application environment of a plurality of target molecular structure data before adding preset atoms and / or preset chemical bonds to the initial molecule according to the preset molecule generation rule.

[0335] The determination module is used to determine the molecular structure diagram that does not meet the molecular stability requirements under the application environment according to the chemical stability rule and the application environment.

[0336] In one embodiment, the predetermined atoms may include at least one of carbon (C), nitrogen (N), oxygen (O), fluorine (F), silicon (Si), sulfur (S), and chlorine (Cl).

[0337] In one embodiment, the conversion module is used to convert multiple target molecules into SMILES format to obtain multiple target molecular structure data, which may specifically include:

[0338] The conversion submodule is used to convert multiple target molecules into an adjacency matrix or adjacency table format to obtain intermediate data of the target molecule structure.

[0339] The conversion submodule is also used to convert the intermediate data of the target molecular structure into the SMILES format to obtain multiple target molecular structure data.

[0340] In one embodiment, the apparatus may further include:

[0341] The creation module is used to create a target molecule structure database based on the multiple target molecule structure data after converting the multiple target molecules into SMILES format to obtain the multiple target molecule structure data.

[0342] The molecular structure data generating device provided in the embodiment of the present application can achieve Figure 1 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0343] Figure 3 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown.

[0344] The electronic device may include a processor 301 and a memory 302 storing computer program instructions.

[0345] Specifically, the processor 301 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0346] The memory 302 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 302 may include removable or non-removable (or fixed) media. Where appropriate, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid-state memory.

[0347] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0348] The processor 301 reads and executes computer program instructions stored in the memory 302 to implement any one of the molecular structure data generation methods in the above embodiments.

[0349] As an example, the electronic device may further include a communication interface 303 and a bus 310. Figure 3 As shown, the processor 301 , the memory 302 , and the communication interface 303 are connected via a bus 310 and communicate with each other.

[0350] The communication interface 303 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0351] Bus 310 comprises hardware, software or both, and the parts of molecular structure data generation equipment are coupled to each other.For example, and not limitation, bus can comprise accelerated graphics port (AGP) or other graphic buses, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations.In suitable cases, bus 310 can comprise one or more buses.Although the present application embodiment describes and shows specific bus, the application considers any suitable bus or interconnection.

[0352] The electronic device can execute the molecular structure data generation method in the embodiment of the present application, thereby realizing the combination Figure 1 and Figure 2 The present invention relates to a method and apparatus for generating molecular structure data.

[0353] In addition, in conjunction with the molecular structure data generation method in the above embodiments, embodiments of the present application may provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the molecular structure data generation methods in the above embodiments is implemented.

[0354] An embodiment of the present application also provides a computer program product, including a computer program, which, when processed and executed, implements any one of the methods for generating molecular structure data in the above embodiments.

[0355] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.

[0356] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0357] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0358] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or flowchart and the combination of the boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0359] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.

Claims

1. A method for generating molecular structure data, characterized in that: include: Obtaining an initial molecule, wherein the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds; According to preset molecular generation rules, adding preset atoms and / or preset chemical bonds to the initial molecules to obtain multiple target molecules, wherein the molecular generation rules include atom addition rules, chemical bond addition rules, valence rules, and molecular stability rules, and the target molecules are represented by an undirected graph without self-loops; Converting the plurality of target molecules into a simplified molecular linear input specification (SMILES) format to obtain structure data of the plurality of target molecules; The preset molecule generation rule includes executing the following steps for each initial molecule: According to the atomic addition rule, the valence rule, and the molecular stability rule, adding a preset number of preset atoms to the initial molecule to obtain a plurality of fifth molecules; storing the plurality of fifth molecules into a preset target molecule set; determining a target fifth molecule from the plurality of fifth molecules, the target fifth molecule being the fifth molecule having the least number of edges among the plurality of fifth molecules; According to the atomic addition rule, the valence rule, and the molecular stability rule, adding a preset number of preset atoms to the target fifth molecule to obtain a plurality of sixth molecules; storing the plurality of sixth molecules into a preset target molecule set; taking the plurality of sixth molecules as new plurality of fifth molecules, and returning to the step of determining a target fifth molecule from the plurality of fifth molecules until a preset iterative stopping condition is satisfied; According to the chemical bond addition rule, the valence rule, and the molecular stability rule, respectively adding chemical bonds of a preset type to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of seventh molecules; storing the plurality of seventh molecules into a preset target molecule set; According to a preset graph isomorphism merging rule, molecules with graph isomorphism in the preset target molecule set are merged.

2. The method according to claim 1, characterized in that The graph isomorphism merging rule includes a graph isomorphism merging rule based on a step-by-step merging algorithm.

3. The method according to claim 1, characterized in that The graph isomorphism merging rule includes a graph isomorphism merging rule based on a multi-grouping algorithm.

4. The method according to claim 1, wherein The molecular stability rules include molecular structure diagrams that do not comply with molecular stability; The step of adding a preset number of preset atoms to the initial molecule according to the atomic addition rule, the valence rule, and the molecular stability rule to obtain a plurality of fifth molecules includes: adding a preset number of preset atoms to the initial molecule according to an atom addition rule and a valence rule to obtain a plurality of initial fifth molecules; searching for the molecular structure diagram in the plurality of initial fifth molecules; The initial fifth molecule including the molecular structure diagram is deleted to obtain the plurality of fifth molecules.

5. The method according to claim 4, characterized in that The initial molecule includes at least one non-hydrogen atom; The step of adding a preset number of preset atoms to the initial molecule according to the atomic addition rule and the valence rule to obtain a plurality of initial fifth molecules includes: Obtaining the number of the preset atoms in the initial molecule; When the number of the preset atoms is less than a preset threshold, the following steps are performed for each non-hydrogen atom in the initial molecule: the atomic degree of the non-hydrogen atom is calculated, the atomic degree includes the sum of the weighted values ​​of the edges connected to the corresponding nodes of the non-hydrogen atom in the undirected graph corresponding to the initial molecule and the corresponding edge attributes, and the edge attributes are used to represent the attributes of the chemical bond corresponding to the edge; according to the correspondence between the preset atomic degree and the chemical bond type, the third chemical bond type corresponding to the atomic degree of the non-hydrogen atom is determined; and the non-hydrogen atom and the preset atom are connected according to the third chemical bond type to obtain the initial fifth molecule.

6. The method according to claim 4, characterized in that According to the chemical bond addition rule, the valence rule, and the molecular stability rule, chemical bonds of preset types are added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of seventh molecules, including: According to the chemical bond addition rule and the valence rule, adding a preset type of chemical bond to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set, respectively, to obtain a plurality of initial seventh molecules; searching for the molecular structure diagram in the plurality of initial seventh molecules; The initial seventh molecule including the molecular structure diagram is deleted to obtain the plurality of seventh molecules.

7. The method according to claim 6, characterized in that The fifth molecule and the sixth molecule each include at least two non-hydrogen atoms; According to the chemical bond addition rule and the valence rule, a preset type of chemical bond is added to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set to obtain a plurality of initial seventh molecules, including performing the following steps for each of the plurality of fifth molecules and the plurality of sixth molecules: For any two non-hydrogen atoms in the molecule, respectively obtain the target atomic degree and target edge attribute of the two non-hydrogen atoms; Determining a fourth chemical bond type corresponding to the target atomic degree and target edge attribute of each of the two non-hydrogen atoms according to a preset correspondence between the atomic degree and edge attribute of each of the two non-hydrogen atoms and the chemical bond type; The two non-hydrogen atoms are connected according to the fourth chemical bond type to obtain a plurality of initial seventh molecules.

8. The method according to any one of claims 1 to 7, characterized in that The molecular stability rules include molecular structure diagrams that do not comply with molecular stability; Before adding preset atoms and / or preset chemical bonds to the initial molecule according to the preset molecule generation rule, the method further includes: an application environment for obtaining the plurality of target molecular structure data; According to the chemical stability rule and the application environment, a molecular structure diagram that does not meet the molecular stability requirements under the application environment is determined.

9. The method according to any one of claims 1 to 7, characterized in that The predetermined atoms include at least one of carbon, nitrogen, oxygen, fluorine, silicon, sulfur, and chlorine.

10. The method according to any one of claims 1 to 7, characterized in that The multiple target molecules are converted into a simplified molecular linear input specification SMILES format to obtain multiple target molecular structure data, including: Converting the plurality of target molecules into an adjacency matrix or adjacency table format to obtain intermediate target molecule structure data; The target molecular structure intermediate data is converted into SMILES format to obtain the multiple target molecular structure data.

11. The method according to any one of claims 1 to 7, characterized in that After converting the plurality of target molecules into a simplified molecular linear input specification (SMILES) format to obtain a plurality of target molecular structure data, the method further includes: A target molecular structure database is created based on multiple target molecular structure data.

12. A molecular structure data generating device, characterized in that: include: An acquisition module, configured to acquire an initial molecule, wherein the initial molecule is represented by an undirected graph without self-loops, and the undirected graph includes nodes for representing atoms and edges for representing chemical bonds; an adding module, configured to add preset atoms and / or preset chemical bonds to the initial molecule according to preset molecular generation rules to obtain a plurality of target molecules, wherein the molecular generation rules include atom addition rules, chemical bond addition rules, valence rules, and molecular stability rules, and the target molecules are represented by an undirected graph without self-loops; A conversion module, configured to convert the plurality of target molecules into a SMILES format to obtain a plurality of target molecule structure data; The device further includes a second molecule generation rule module, wherein the second molecule generation rule module includes: a second adding submodule, configured to add a preset number of preset atoms to the initial molecule according to an atom adding rule, a valence rule, and a molecular stability rule, to obtain a plurality of fifth molecules; A second storage submodule, configured to store the plurality of fifth molecules into a preset target molecule set; a determination submodule, configured to determine a target fifth molecule from the plurality of fifth molecules, wherein the target fifth molecule is the fifth molecule having the least number of edges among the plurality of fifth molecules; The second adding submodule is further configured to add a preset number of preset atoms to the target fifth molecule according to an atom adding rule, a valence rule, and a molecular stability rule to obtain a plurality of sixth molecules; The second storage submodule is further used to store the plurality of sixth molecules into a preset target molecule set; a second returning submodule, configured to use the plurality of sixth molecules as new plurality of fifth molecules and return the target fifth molecule determined from the plurality of fifth molecules until a preset iteration stopping condition is satisfied; The second adding submodule is further configured to add chemical bonds of a preset type to the plurality of fifth molecules and the plurality of sixth molecules in the preset target molecule set according to the chemical bond adding rule, the valence rule, and the molecular stability rule, to obtain a plurality of seventh molecules; The second storage submodule is further used to store the plurality of seventh molecules into a preset target molecule set; The merging submodule is used to merge the molecules with graph isomorphism in the preset target molecule set according to the preset graph isomorphism merging rule.

13. An electronic device, characterized in that: The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the method for generating molecular structure data according to any one of claims 1 to 11 is implemented.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the method for generating molecular structure data according to any one of claims 1 to 11 is implemented.

15. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the molecular structure data generating method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Directional molecule generation method based on graph neural network

    CN113140267A

  • Molecular map generation method and device

    CN114822721A