Molecular property prediction method and device, electrolyte design method, secondary battery

Through a prediction model based on knowledge embedding, the melting point, boiling point and flash point of secondary battery electrolyte solvents can be predicted quickly and accurately, solving the problem of poor versatility in electrolyte property prediction in existing technologies, achieving automation and efficiency in electrolyte design, and developing safer secondary batteries.

CN119252363BActive Publication Date: 2025-09-23TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411313929.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2025-09-23
Estimated Expiration
2044-09-19

AI Technical Summary

Technical Problem

In the existing technology, the prediction methods of the molecular properties of electrolyte solvents for secondary batteries have poor versatility and require a lot of manual intervention. It is difficult to quickly and accurately obtain the melting point, boiling point and flash point. In addition, the electrolyte design method is time-consuming and costly, and it is difficult to operate safely over a wide temperature range.

Method used

A knowledge embedding-based prediction model is adopted to predict the properties of electrolyte solvent molecules through the knowledge embedding of atomic structure vectors, bond structure vectors and molecular structure vectors, using the deep learning model Uni-Mol to achieve fast and accurate prediction of melting point, boiling point and flash point, and then screen out suitable solvents.

Benefits of technology

The prediction of the properties of electrolyte solvent molecules has been made universal and automated, which reduces the complexity of prediction, improves the efficiency of electrolyte design, and develops secondary batteries that can operate safely over a wide temperature range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119252363B_ABST
    Figure CN119252363B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for predicting molecular properties, an electrolyte design method, and a secondary battery. The property prediction method utilizes a prediction model based on knowledge embedding to process prediction tasks. The processing includes obtaining a prediction task for a target molecule, determining a target property to be predicted, wherein the target property includes at least one of a melting point, a boiling point, and a flash point, and the target molecule is one of the optional solvents for the electrolyte to be prepared; determining an atomic structure vector, a bond structure vector, and a molecular structure vector based on the molecular structure of the target molecule; performing knowledge embedding on the atomic structure vector and the bond structure vector based on the knowledge vector to obtain a first embedding result; performing knowledge embedding on a vector determined based on the first embedding result and the molecular structure vector based on the knowledge vector to obtain a second embedding result; and determining the target property of the target molecule based on the second embedding result. The present disclosure can quickly and accurately obtain the melting point, boiling point, and flash point of a solvent molecule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of secondary batteries, and in particular to a method and device for predicting molecular properties, a method for designing an electrolyte, and a secondary battery. Background Art

[0002] In recent years, the requirements for the battery performance of secondary batteries have been continuously improved. However, in some specific application scenarios, the performance of secondary batteries is severely limited. Therefore, it is particularly important to develop secondary batteries that can operate safely over a wide temperature range. The operating temperature range and safety performance of secondary batteries depend to a large extent on the physicochemical properties of the electrolyte, and the development of new chemical compositions of electrolytes is crucial. However, in the related art, the molecular property prediction methods of solvent molecules in electrolytes have poor versatility, are often limited to specific systems or single properties, and require a lot of manual intervention. In order to quickly and accurately obtain the melting point, boiling point and flash point of solvent molecules in electrolytes, there is an urgent need for a new method to realize a universal and automated molecular property prediction method. Summary of the Invention

[0003] In view of this, the present disclosure proposes a molecular property prediction method and device, an electrolyte design method, and a secondary battery, which can quickly and accurately obtain the melting point, boiling point, and flash point of the solvent molecules in the electrolyte to be prepared, thereby realizing universal and automated molecular property prediction.

[0004] According to one aspect of the present disclosure, a method for predicting properties of a molecule is provided, wherein the method utilizes a prediction model based on knowledge embedding to process a prediction task for a target molecule, and the processing comprises the following steps: obtaining a prediction task for the target molecule, determining a target property to be predicted according to the prediction task, wherein the target property comprises at least one of a melting point, a boiling point, and a flash point, and the target molecule is one of the optional solvents for the electrolyte to be prepared; determining an atomic structure vector, a bond structure vector, and a molecular structure vector based on the molecular structure of the target molecule, wherein the atomic structure vector represents the structural information of all atoms in the target molecule, and the bond structure vector represents the structural information of all chemical bonds in the target molecule. The molecular structure vector represents the overall structural information of the target molecule; based on the pre-obtained knowledge vector, the atomic structure vector and the bond structure vector are knowledge-embedded to obtain a first embedding result; based on the knowledge vector, the vector determined according to the first embedding result and the molecular structure vector is knowledge-embedded to obtain a second embedding result, wherein the knowledge vector is determined based on the molecular characteristics of each sample molecule, each of the molecular characteristics includes the number of atoms, chemical bond properties, functional group category, and electronic properties, and each of the sample molecules is a solvent in an existing electrolyte; based on the second embedding result, the target properties of the target molecule are determined, so as to screen the solvent of the electrolyte according to the target properties of the target molecule.

[0005] In one possible implementation, knowledge embedding is performed on the atomic structure vector and the bond structure vector based on a pre-obtained knowledge vector to obtain a first embedding result; knowledge embedding is performed on the vector determined according to the first embedding result and the molecular structure vector based on the knowledge vector to obtain a second embedding result, including: determining first knowledge used for knowledge embedding of the atomic structure vector, second knowledge used for knowledge embedding of the bond structure vector, and third knowledge used for knowledge embedding of the molecular structure vector based on the knowledge vector; embedding the first knowledge and the second knowledge into the atomic structure vector and the bond structure vector respectively to obtain the first embedding result; and embedding the third knowledge into the vector determined according to the first embedding result and the molecular structure vector to obtain the second embedding result.

[0006] In one possible implementation, the knowledge vector has multiple dimensions, and each dimension has at least one knowledge point; wherein, the method further includes: performing Shapley value analysis based on the knowledge vector to obtain the contribution value corresponding to each dimension in the knowledge vector, and the contribution value corresponding to each dimension represents the degree of influence of the knowledge point of the dimension on the molecular properties; determining at least one target dimension whose contribution value is greater than a preset threshold from all dimensions of the knowledge vector; and determining the first knowledge, the second knowledge, and the third knowledge based on part or all of the knowledge points in each target dimension.

[0007] In one possible implementation, the first knowledge and the second knowledge are correspondingly embedded into the atomic structure vector and the bond structure vector to obtain the first embedding result, including: embedding the first knowledge into the atomic structure vector to obtain a first embedding vector; and embedding the second knowledge into the bond structure vector to obtain a second embedding vector; and using the first embedding vector and the second embedding vector as the first embedding result.

[0008] In one possible implementation, the first knowledge and the second knowledge are correspondingly embedded into the atomic structure vector and the bond structure vector to obtain the first embedding result, including: embedding the first knowledge into the atomic structure vector to obtain a first embedding vector; embedding the second knowledge into a vector determined according to the first embedding vector and the bond structure vector to obtain a third embedding vector; and using the first embedding vector and the third embedding vector as the first embedding result.

[0009] In one possible implementation, the method also includes a training process for the prediction model, and the training process includes the following steps: obtaining molecular data of each of the sample molecules, each of the molecular data is used to represent the molecular structure and molecular properties of the corresponding sample molecule, and the molecular properties include melting point, boiling point, and flash point; performing feature extraction on each of the molecular data to obtain molecular features of the corresponding sample molecules, and determining the knowledge vector based on all molecular features, wherein the knowledge vector has multiple dimensions, and each dimension has at least one knowledge point; extracting some knowledge points in each of the dimensions to obtain a first knowledge point set, and using the first knowledge point set and the sample set to train the original prediction model to obtain a first model; using the second knowledge point set and the sample set to train the first model to obtain a trained prediction model, wherein the second knowledge point set includes knowledge points under the target dimension, and the target dimension is obtained based on Shapley value analysis of the knowledge vector.

[0010] In a possible implementation, the method further includes: determining an error range corresponding to the target property of the target molecule based on the second embedding result.

[0011] According to another aspect of the present disclosure, a device for predicting properties of a molecule is provided, comprising: an acquisition module for acquiring a prediction task for a target molecule, and determining a target property to be predicted based on the prediction task, wherein the target property includes at least one of a melting point, a boiling point, and a flash point, and the target molecule is one of the optional solvents for the electrolyte to be prepared; a determination module for determining an atomic structure vector, a bond structure vector, and a molecular structure vector based on the molecular structure of the target molecule, wherein the atomic structure vector represents the structural information of all atoms in the target molecule, the bond structure vector represents the structural information of all chemical bonds in the target molecule, and the molecular structure vector represents the overall structural information of the target molecule. ; A knowledge embedding module, used to perform knowledge embedding on the atomic structure vector and the bond structure vector based on the pre-obtained knowledge vector to obtain a first embedding result; based on the knowledge vector, perform knowledge embedding on the vector determined according to the first embedding result and the molecular structure vector to obtain a second embedding result, wherein the knowledge vector is determined based on the molecular characteristics of each sample molecule, each of the molecular characteristics includes the number of atoms, chemical bond properties, functional group category, and electronic properties, and each of the sample molecules is a solvent in an existing electrolyte; A prediction module, used to determine the target properties of the target molecule based on the second embedding result, so as to screen the solvent of the electrolyte according to the target properties of the target molecule.

[0012] In one possible implementation, knowledge embedding is performed on the atomic structure vector and the bond structure vector based on a pre-obtained knowledge vector to obtain a first embedding result; knowledge embedding is performed on the vector determined according to the first embedding result and the molecular structure vector based on the knowledge vector to obtain a second embedding result, including: determining first knowledge used for knowledge embedding of the atomic structure vector, second knowledge used for knowledge embedding of the bond structure vector, and third knowledge used for knowledge embedding of the molecular structure vector based on the knowledge vector; embedding the first knowledge and the second knowledge into the atomic structure vector and the bond structure vector respectively to obtain the first embedding result; and embedding the third knowledge into the vector determined according to the first embedding result and the molecular structure vector to obtain the second embedding result.

[0013] In one possible implementation, the knowledge vector has multiple dimensions, and each dimension has at least one knowledge point; wherein, the device also includes an analysis module, which is used to: perform Shapley value analysis based on the knowledge vector to obtain the contribution value corresponding to each dimension in the knowledge vector, and the contribution value corresponding to each dimension represents the degree of influence of the knowledge point of the dimension on the molecular properties; determine at least one target dimension whose contribution value is greater than a preset threshold from all dimensions of the knowledge vector; and determine the first knowledge, the second knowledge, and the third knowledge based on part or all of the knowledge points in each target dimension.

[0014] In one possible implementation, the first knowledge and the second knowledge are correspondingly embedded into the atomic structure vector and the bond structure vector to obtain the first embedding result, including: embedding the first knowledge into the atomic structure vector to obtain a first embedding vector; and embedding the second knowledge into the bond structure vector to obtain a second embedding vector; and using the first embedding vector and the second embedding vector as the first embedding result.

[0015] In one possible implementation, the first knowledge and the second knowledge are correspondingly embedded into the atomic structure vector and the bond structure vector to obtain the first embedding result, including: embedding the first knowledge into the atomic structure vector to obtain a first embedding vector; embedding the second knowledge into a vector determined according to the first embedding vector and the bond structure vector to obtain a third embedding vector; and using the first embedding vector and the third embedding vector as the first embedding result.

[0016] In one possible implementation, the device also includes a training module for performing a training process for the prediction model, and the training process includes: obtaining molecular data of each of the sample molecules, each of the molecular data is used to represent the molecular structure and molecular properties of the corresponding sample molecule, and the molecular properties include melting point, boiling point, and flash point; performing feature extraction on each of the molecular data to obtain molecular features of the corresponding sample molecules, and determining the knowledge vector based on all molecular features, wherein the knowledge vector has multiple dimensions, and each dimension has at least one knowledge point; extracting some knowledge points in each of the dimensions to obtain a first knowledge point set, and using the first knowledge point set and the sample set to train the original prediction model to obtain a first model; using the second knowledge point set and the sample set to train the first model to obtain a trained prediction model, wherein the second knowledge point set includes knowledge points under the target dimension, and the target dimension is obtained based on Shapley value analysis of the knowledge vector.

[0017] In a possible implementation, the device further includes a range determination module, configured to determine an error range corresponding to the target property of the target molecule based on the second embedding result.

[0018] According to another aspect of the present disclosure, a method for designing an electrolyte is provided, comprising: obtaining a target molecule, which is an optional solvent for the electrolyte to be prepared; determining target properties of the target molecule using the above-mentioned molecular property prediction method, wherein the target properties include at least one of a melting point, a boiling point, and a flash point; and if the target properties meet preset electrolyte design conditions, preparing the electrolyte based on the target molecule.

[0019] According to another aspect of the present disclosure, a secondary battery is provided, wherein the electrolyte used in the secondary battery is determined by the above-mentioned electrolyte design method.

[0020] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above-mentioned molecular property prediction method when executing the instructions stored in the memory.

[0021] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above-mentioned molecular property prediction method when executed by a processor.

[0022] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0023] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.

[0025] Figure 1 A flowchart illustrating a method for predicting molecular properties according to an embodiment of the present disclosure is shown.

[0026] Figures 2 to 3 A block diagram illustrating a device for predicting molecular properties according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0027] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0028] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0029] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0030] In order to facilitate those skilled in the art to understand the technical solution provided by the embodiments of the present disclosure, the technical environment in which the technical solution is implemented is described below.

[0031] In recent years, the widespread use of secondary batteries such as lithium-ion batteries in electric vehicles and portable electronic devices has led to increasing demands for battery performance. However, in certain specific application scenarios, such as harsh cold environments, desert areas, and special requirements such as spacecraft, underground exploration, and medical equipment sterilization, the performance of secondary batteries is severely limited. Under high temperature conditions, side reactions are intensified, leading to rapid consumption of electrolyte and active materials, and even triggering thermal runaway of the battery. Under low temperature conditions, the reaction kinetics are significantly slowed, and lithium dendrites may form on the graphite negative electrode, posing a safety hazard. Therefore, it is particularly important to develop secondary batteries that can operate safely over a wide temperature range.

[0032] The operating temperature range and safety performance of batteries depend heavily on the physicochemical properties of the electrolyte, making the development of novel electrolyte chemical compositions and design strategies crucial. However, traditional research methods based on trial and error are time-consuming and difficult to find electrolytes that simultaneously meet the requirements of low-temperature adaptability, high-temperature stability, and non-flammability, which seriously hinders electrolyte development.

[0033] With the rapid development of high-performance computing and data-driven technologies, simulation prediction and high-throughput virtual screening have become effective means to accelerate electrolyte development. Currently, several methods exist for predicting the melting, boiling, and flash points of electrolytes. For example, early molecular property estimation methods employed specific mathematical functions for fitting, but these methods often required the introduction of additional physical parameters and simplified assumptions for calculation, resulting in limited applicability. To improve the versatility of property prediction, the group contribution method assumes that the contributions of individual functional groups are consistent across molecules, rapidly deriving molecular properties through linear summation. However, this method lacks consideration of interactions between different groups, resulting in limited accuracy. With the advancement of computational chemistry, molecular dynamics simulations can now predict the melting and boiling points of electrolytes with reasonable accuracy. However, predicting phase transition temperatures often requires specialized force fields and accurate crystal structures, increasing the complexity of the simulations. With the rapid advancement of machine learning techniques, quantitative structure-activity relationship methods have achieved success in drug design, but they still rely heavily on the precise construction of molecular descriptors. In addition, developing an electrolyte with a wide temperature range and high safety often requires a large amount of experimental design. For example, researchers need to design the material composition and ratio of a variety of different electrolytes based on experience, and then verify the performance of the electrolyte, or optimize the formula based on existing electrolytes.

[0034] In summary, traditional molecular property prediction methods have limited versatility, are often limited to specific systems or single properties, and require extensive manual intervention. Furthermore, electrolyte design methods in related technologies not only have long R&D cycles, high costs, and poor mobility, but also struggle to capture the inherent principles of electrolyte design and development. In order to quickly and accurately determine the melting, boiling, and flash points of solvent molecules in electrolytes, and thereby develop secondary batteries that can operate safely over a wide temperature range, a new method for universal and automated property prediction is urgently needed.

[0035] In order to solve the above technical problems, an embodiment of the present disclosure provides a method for predicting the properties of a molecule, which uses a prediction model based on knowledge embedding to process a prediction task for a target molecule, and the processing includes the following steps: obtaining a prediction task for the target molecule, determining the target property to be predicted according to the prediction task, the target property including at least one of melting point, boiling point, and flash point, and the target molecule is one of the optional solvents for the electrolyte to be prepared; based on the molecular structure of the target molecule, determining the atomic structure vector, the bond structure vector, and the molecular structure vector, the atomic structure vector represents the structural information of all atoms in the target molecule, the bond structure vector represents the structural information of all chemical bonds in the target molecule, and the molecular structure vector represents the overall structural information of the target molecule; based on the pre-obtained knowledge vector, the atomic structure vector and the bond structure vector are knowledge embedded to obtain a first embedding structure. result; based on the knowledge vector, the vector determined according to the first embedding result and the molecular structure vector is embedded with knowledge to obtain a second embedding result, wherein the knowledge vector is determined based on the molecular characteristics of each sample molecule, and each molecular characteristic includes the number of atoms, chemical bond properties, functional group category, and electronic properties, and each sample molecule is a solvent in an existing electrolyte; based on the second embedding result, the target properties of the target molecule are determined, and the solvent of the electrolyte is screened according to the target properties of the target molecule. In this way, there is no need to introduce additional physical parameters and simplifying assumptions for calculation. The knowledge vector that describes the inherent laws of electrolyte design and development in advance is used to perform knowledge embedding at the atomic, bond, and entire molecular levels corresponding to the target molecule, so that the melting point, boiling point, and flash point of the solvent molecule in the electrolyte can be obtained quickly and accurately, thereby realizing universal and automated molecular property prediction.

[0036] Figure 1 Flowchart showing the method for predicting the properties of molecules according to an embodiment of the present disclosure. Figure 1 As shown, the molecular property prediction method utilizes a prediction model based on knowledge embedding to process a prediction task for a target molecule, and the process may include the following steps S101 to S104.

[0037] Step S101: Obtain a prediction task for a target molecule, and determine the target property to be predicted according to the prediction task.

[0038] The target property includes at least one of a melting point, a boiling point, and a flash point. The target molecule is one of the optional solvents of the electrolyte to be prepared.

[0039] Step S102: Based on the molecular structure of the target molecule, determine the atomic structure vector, the bond structure vector, and the molecular structure vector.

[0040] In some embodiments, the molecular structure information of the target molecule can be converted into a format that can be processed or understood by a computer, that is, the target molecule is molecularly encoded, and vectors corresponding to the three levels of atoms, chemical bonds (bonds for short), and the entire molecule can be obtained based on the encoding results, that is, atomic structure vectors, bond structure vectors, and molecular structure vectors. The atomic structure vector can represent the structural information of all atoms in the target molecule, the bond structure vector can represent the structural information of all chemical bonds in the target molecule, and the molecular structure vector can represent the overall structural information of the target molecule. Taking the atomic structure vector as an example, the structural information of the atom may include but is not limited to the type of atom and the mass of the atom. In fact, the structural information at different levels represented by the atomic structure vector, the bond structure vector, and the molecular structure vector can be completely determined according to the actual molecular encoding method selected, and the embodiments of the present disclosure do not limit this.

[0041] Step S103: Based on the pre-obtained knowledge vector, the atomic structure vector and the bond structure vector are knowledge-embedded to obtain a first embedding result; based on the knowledge vector, the vector determined according to the first embedding result and the molecular structure vector is knowledge-embedded to obtain a second embedding result.

[0042] The molecular property prediction method of the embodiment of the present disclosure performs knowledge embedding at three levels: atoms, bonds, and entire molecules. The knowledge vector used for knowledge embedding is determined based on the molecular characteristics of each sample molecule. The knowledge vector has multiple dimensions, and each dimension may have at least one knowledge point. Each sample molecule is a solvent in an existing electrolyte. Each molecular feature may include the number of atoms (such as the specified number of atoms, the ratio of the number of atoms of different categories, etc.), chemical bond properties (such as whether it contains double bonds, triple bonds, aromatic bonds, etc.), functional group categories (or molecular group categories, such as hydroxyl, carboxyl, ester, aldehyde, ketone, etc.), and electronic properties (such as electronegativity, affinity, first ionization energy, etc.). In this way, based on the rules and principles of chemical field knowledge, these four types of features are selected to construct the knowledge vector to enhance the chemical rationality of the molecular representation.

[0043] In some embodiments, step S103 may include: based on the knowledge vector, determining the first knowledge used for knowledge embedding into the atomic structure vector, the second knowledge used for knowledge embedding into the bond structure vector, and the third knowledge used for knowledge embedding into the molecular structure vector, wherein the first knowledge, the second knowledge, and the third knowledge may be completely identical, not completely identical, or completely different, and may be flexibly set according to actual needs, and the embodiments of the present disclosure do not limit this; embedding the first knowledge and the second knowledge into the atomic structure vector and the bond structure vector respectively to obtain a first embedding result; embedding the third knowledge into the vector determined according to the first embedding result and the molecular structure vector to obtain a second embedding result.

[0044] The order of embedding knowledge at the atomic and bond levels can be flexibly set, and can be atoms first and then bonds, bonds first and then atoms, or can be performed simultaneously. In the above step S103, the first knowledge and the second knowledge are embedded into the atomic structure vector and the bond structure vector respectively to obtain a first embedding result, which may include: embedding the first knowledge into the atomic structure vector to obtain a first embedding vector; and embedding the second knowledge into the bond structure vector to obtain a second embedding vector; using the first embedding vector and the second embedding vector as the first embedding result. In this way, the second embedding vector corresponding to the bond level does not contain information at the atomic level.

[0045] Alternatively, the above-mentioned step S103 of embedding the first and second knowledge into the atomic structure vector and the bond structure vector to obtain the first embedding result may include: embedding the first knowledge into the atomic structure vector to obtain the first embedding vector; embedding the second knowledge into a vector determined based on the first embedding vector and the bond structure vector to obtain a third embedding vector; and using the first embedding vector and the third embedding vector as the first embedding result. In this way, the first embedding result corresponding to the atomic level is first combined with the bond structure vector, and then the bond-level knowledge is embedded based on the combined vector, which helps to obtain more accurate prediction results.

[0046] The present molecular property prediction method may also include the following steps for determining the first knowledge, the second knowledge, and the third knowledge: performing Shapley value analysis based on the knowledge vector to obtain the contribution value corresponding to each dimension in the knowledge vector, the contribution value corresponding to each dimension indicates the degree of influence of the knowledge point of the dimension on the molecular properties, the larger the contribution value of a dimension, the greater the influence of the dimension on the molecular properties, and at the same time, using Shapley value analysis, it can also be known how each dimension affects the molecular properties, that is, whether there is a positive or negative correlation between each dimension and the molecular properties; determining at least one target dimension whose contribution value is greater than a preset threshold from all dimensions of the knowledge vector, and the preset threshold can be set according to actual needs; determining the first knowledge, the second knowledge, and the third knowledge based on part or all of the knowledge points in each target dimension, the first knowledge, the second knowledge, and the third knowledge can be completely the same, completely different, or not completely the same. For example, if the first knowledge, the second knowledge, and the third knowledge are completely the same, then a target knowledge can be directly determined based on part or all of the knowledge points in each target dimension, and the target knowledge is used as the first knowledge, the second knowledge, and the third knowledge.

[0047] Step S104 : determining the target properties of the target molecule based on the second embedding result, so as to screen the solvent of the electrolyte according to the target properties of the target molecule.

[0048] In some embodiments, the prediction model can be a deep learning Uni-Mol model. The second embedding result is processed by the Uni-Mol model's prediction head to obtain the target property of the target molecule. The processing of the Uni-Mol model is determined based on the actual structure used. If necessary, reference can be made to related art, and this article will not elaborate further.

[0049] Through steps S101 to S104, the melting point, boiling point and flash point of the solvent molecules in the electrolyte can be obtained quickly and accurately without introducing additional physical parameters and simplified assumptions for calculation. This method has high applicability, reduces the complexity of the prediction, and significantly improves the efficiency and performance of the electrolyte design, thereby developing secondary batteries that can operate safely within a wide temperature range.

[0050] In some embodiments, the property prediction method may also include a prediction model training process, which may include the following steps: obtaining molecular data of each sample molecule, each molecular data is used to represent the molecular structure and molecular properties of the corresponding sample molecule, and the molecular properties include melting point, boiling point, and flash point. For example, a data collection, analysis, and statistics module can be used to extract existing electrolyte molecular structures and their corresponding molecular property data (melting point, boiling point, and flash point) from public databases and reported papers, and the collected data can be preprocessed, filtered, and organized to generate a structured data set for model training; feature extraction is performed on each molecular data to obtain the molecular features of the corresponding sample molecule, and a knowledge vector is determined based on all molecular features. The knowledge vector has multiple dimensions, and each dimension has at least one knowledge point. For example, a The feature engineering and explainable knowledge discovery module extracts molecular features from structured data sets, and uses explainable machine learning algorithms to complete the knowledge discovery process based on molecular features; extracts some knowledge points in each dimension to obtain a first knowledge point set, and uses the first knowledge point set and the sample set to train the original prediction model to obtain a first model; uses the second knowledge point set and the sample set to train the first model to obtain a trained prediction model, wherein the second knowledge point set includes knowledge points under the target dimension, and the target dimension is obtained based on Shapley value analysis of the knowledge vector. For example, a knowledge embedding and property prediction module can be used to train the prediction model using a machine learning algorithm, and the knowledge discovered in the feature engineering and explainable knowledge discovery modules can be embedded into the prediction model to improve the prediction performance of the prediction model.

[0051] Below above-mentioned data collection, analysis and statistical module are described in detail.In the present embodiment, first utilize open database interface (including but not limited to PubChem interface and ChemSpider interface) to automatically collect the molecular structure of sample molecules in database and corresponding property data thereof in batches.At the same time, the data in the paper of public report can be collected by integrated software.In order to ensure the accuracy and consistency of data, chemical structural formula is carried out standardization using chemical information toolkit (such as RDKit tool).According to identifiers such as International Union of Pure and Applied Chemistry (IUPAC) name, Chemical Abstracts Service (CAS) number and PubChem ID, SMILES expression of molecular structure is obtained by open database interface retrieval.Subsequently, utilize chemical information toolkit to delete the molecule that cannot be converted or incomplete, and complete conversion based on SMILES standardization algorithm.

[0052] After data preprocessing is completed, the properties of melting point, boiling point, and flash point are statistically analyzed. To facilitate intuitive observation of distribution and properties, in some embodiments, the properties to be observed (such as melting point, boiling point, flash point, number of heavy atoms, and molecular weight) are normalized to a range of 0 to 1 using the minimum-maximum normalization method, and a statistical tool (such as the Python program Matplotlib) is used to create a histogram for data binning and counting. The number of equal-width bins is set to 50 to ensure that the statistical analysis results of the data are representative.

[0053] In order to further analyze the relationship between molecular structure and molecular properties, in some embodiments, the solvent molecules in the electrolyte are classified. For example, they can be divided into three categories: hydrocarbons, oxygen-containing functional groups, and molecules containing other functional groups. Hydrocarbons include molecules containing only carbon and hydrogen elements, which are further subdivided into linear and cyclic hydrocarbons; oxygen-containing functional groups include molecules containing carbon, oxygen, and hydrogen elements, which are subdivided into alcohols, phenols, aldehydes, ketones, carboxylic acids, ethers, and esters; molecules containing other functional groups are counted according to the heteroatoms they contain, and all molecules are statistically counted in detail by category.

[0054] In order to explore the relationship between molecular properties and macroscopic categories, statistical descriptive methods (such as half-violin plots, etc.) are used for illustration. Half-violin plots show the scatter distribution and kernel density distribution of the data to enhance the understanding and analysis of the data set.

[0055] In order to visualize the data set distribution, in certain embodiments, use molecular fingerprint to represent molecular features, then utilize cluster analysis algorithm that molecule is reduced to two-dimensional plane and utilize correlation property to be colored.For example, can use the MACCS molecular fingerprint in the DeepChem toolkit that all sample molecules are characterized, the molecular structure of sample molecules is encoded as a 167-dimensional binary string, is made up of 166 features plus a placeholder, and each feature is corresponding to a predefined structural feature.If there is predefined molecular structure feature, then this feature is set to 1, otherwise is set to 0.Subsequently, use the t distribution random neighbor embedding algorithm in the Scikit-learn toolkit to carry out cluster analysis on these binary strings.T distribution random neighbor embedding algorithm is a kind of nonlinear dimensionality reduction technique, is used for high dimensional data embedding low dimensional space, retains the local structure of data point. These 167-dimensional binary strings are embedded in a two-dimensional space. The embedding is initialized using principal component analysis, with the perplexity set to 35 and the random state set to 0. The generated two-dimensional vectors are colored according to the melting point, boiling point, and flash point of the molecule, thereby visualizing the data distribution of all sample molecules, in order to gain a preliminary understanding of the distribution of the melting point, boiling point, and flash point of the electrolyte solvent molecules.

[0056] In some embodiments, after obtaining the molecular data of each sample molecule, data that does not meet the requirements of a given scenario can be removed according to pre-set rules (for example, chemical stability is not met, non-planar molecules, and ring tension is large), and outliers in the data can be removed at the same time to ensure data quality.

[0057] In this way, the function of automatically collecting data through the application programming interface improves the efficiency and comprehensiveness of data collection. Moreover, by normalizing the data to the same scale according to the standardization steps, the stability and accuracy of the prediction model training can be improved. In addition, by visually displaying the obtained raw data, the understanding and analysis of the data set can be enhanced, and statistical description can be performed, including statistical analysis and drawing according to given requirements (such as specified functional groups, specified substructures, etc.).

[0058] The above-mentioned feature engineering and explainable knowledge discovery modules are described in detail below. In this embodiment, molecular feature descriptors are extracted by chemical informatics tools, which are mainly divided into four parts: the number of atoms, the nature of the bond, the functional group, and the electronic properties. For example, the number of carbon and oxygen atoms is constructed by a string search method, and the number of carbon branches is defined as the number of branches on the longest carbon chain. The molecule is converted into an adjacency matrix using a graph theory algorithm, and all possible simple paths between each pair of atoms are found using a graph algorithm library to determine the atomic index list of the longest carbon chain. The algorithm also identifies the number of atoms in the largest ring in the molecule. Among them, the electronic property descriptor is defined as the average ionization energy, average electron affinity, and average electronegativity of each element multiplied by the number of atoms of the element in the molecule, and then the total number of atoms in the molecule is averaged. The calculation formula for the electronic properties of a sample molecule can be found in Formula 1 below:

[0059]

[0060] Where A, B, and C represent the different elements in the sample molecule, E(A) represents the first ionization energy of the atom corresponding to element A, E(B) represents the electron affinity of the atom corresponding to element B, E(C) represents the electronegativity of the atom corresponding to element C, M represents the electronic properties of the molecule, x represents the number of atoms corresponding to element A, y represents the number of atoms corresponding to element B, and z represents the number of atoms corresponding to element C.

[0061] The above-mentioned knowledge embedding and property prediction modules are described in detail below. In this embodiment, first, 10 molecular conformations are randomly generated using the RDKit toolkit, and these conformations and property data are stored in the Lightning Memory Mapped Database Manager (LMDB). The Uni-Mol model uses these packaged data as input, and the pre-training part uses the weights trained by the original author and is fine-tuned according to the given molecular property data. The optimizer uses Adam, the loss function is the smoothed mean absolute error, and the validation set evaluation uses the mean absolute error. The batch size is set to 32, the maximum number of iterations is 500, and an early stopping mechanism is set. If the loss on the validation set does not improve after several training iterations, the training is stopped early. The learning rate uses polynomial decay, and the initial value is set to 10 -4 .

[0062] Initially, molecules are input in SMILES format and encoded as representations of atoms and bonds. Subsequently, 64-dimensional molecular feature vectors extracted by the feature engineering and interpretable knowledge discovery modules are embedded into these representations. The proportion of knowledge is adjusted by a knowledge purity controller and a knowledge flow controller. The knowledge purity controller is used to adjust the accuracy of the embedded knowledge. It performs Shapley value analysis on the knowledge vector to obtain the contribution value corresponding to each dimension, thereby selectively controlling the embedding of certain dimensions in the knowledge vector into these representations. The specific number of knowledge points embedded is controlled by the knowledge flow controller. Initially, the original knowledge vector is encoded to 512 dimensions through a linear layer by the knowledge purity controller. It is then transformed through a Gaussian Error Linear Unit (GELU) nonlinear activation function, followed by random pooling and layer normalization, and then passed through another linear layer to obtain the transformed knowledge. The knowledge flow controller automatically controls the representation vector (x) of atoms, bonds, or the entire molecule to the transformed knowledge vector through learnable parameters in the neural network. Parameters are set to control the amount of knowledge vector embedded. The knowledge embedding process of the knowledge flow controller can be expressed as follows:

[0063] output = x + α × knowledge (Equation 2)

[0064] In Equation 2, x represents the structure vector of an atom, bond, or entire molecule (i.e., atomic structure vector, bond structure vector, molecular structure vector), knowledge corresponding to x represents the converted knowledge vector (the dimension is first controlled by the knowledge purity controller, and then the number of knowledge points is controlled by the knowledge flow controller), α is a learnable parameter, and output represents the embedding result of the atom, bond, or entire molecule. The correspondence between x and knowledge means that when x is an atomic structure vector, the knowledge added to x corresponds to the atomic level, and the same applies to bonds and entire molecules. After being processed by the prediction head of the prediction model, the output obtains at least one of the predicted melting point, boiling point, and flash point, depending on the prediction task.

[0065] During the training of the prediction model, the knowledge purity controller was initially set to maximum, fully embedding all 64-dimensional knowledge vectors into the molecular representation to enhance knowledge retention. Simultaneously, the knowledge flow controller was set to automatically adjust the proportion of knowledge currently embedded in the molecular representation in a learnable manner. To prevent overfitting, cross-validation was used for model evaluation to minimize the impact of data splitting on model performance. Final performance was evaluated using mean absolute error (MAE) on an unseen test set. For the test cases melting point, boiling point, and flash point, the MAEs were 11.3 Kelvin (K), 5.2 K, and 5.4 K, respectively. Compared to the non-knowledge-embedded version, the model with 64-dimensional knowledge embedding improved performance by 5%–11%, demonstrating the effectiveness of knowledge embedding. Next, to improve the predictive performance of the prediction model, the knowledge purity controller was fine-tuned to adjust the effective purity of the embedded knowledge, specifically selecting knowledge points in the target dimension that contribute most to the accuracy of molecular property predictions. As more features are embedded, the introduction of less relevant knowledge increases the learning burden, leading to a decrease in prediction accuracy. On the baseline with the best performance of all knowledge embeddings, the prediction accuracy of the prediction model was measured by increasing the number of dimensions by 10 each time. For the prediction of melting point, boiling point, and flash point, the model achieved the best performance when the number of embedded knowledge was 20, 20, and 10, respectively, with mean absolute error values ​​of 10.5K, 4.6K, and 4.8K, respectively. Compared with the model without embedded knowledge, the mean absolute error in the prediction of melting point, boiling point, and flash point improved by 6.7%, 14.7%, and 17.8%, respectively. Even compared with the random forest model that relies only on knowledge, the improvements in the prediction of melting point, boiling point, and flash point were 51.9%, 68.2%, and 55.5%, respectively.

[0066] In order to further quantify the effectiveness of the embedded knowledge, an ablation study was conducted. Based on the results, it was found that the performance of the prediction model was best when the knowledge was embedded at the atomic, bond and entire molecule levels. The model performance can be evaluated by comparing the mean absolute error, mean square error, root mean square error and coefficient of determination of the predicted and actual measured values. In order to evaluate the generalization of the method, the data set in the reported method was used as a baseline, involving methods such as group contribution method, multivariate linear regression based on quantitative structure-activity relationship method and machine learning. Optimal results were achieved in 18 of the 20 data sets collected. The supplementary materials provide detailed information on the molecular types, scale and methods. In this embodiment, the generalization ability of the model can also be verified by an independent data set.

[0067] During the prediction model training process, various supervised learning methods, including linear regression, support vector machine regression, random forest regression, and gradient boosting regression, can be used to correlate molecular descriptors with molecular properties (such as melting point, boiling point, and flash point). The prediction model inputs are the molecular structure and the corresponding molecular features (i.e., the extracted 64-dimensional molecular feature vector), as well as the melting point, boiling point, and flash point. Model parameters are continuously optimized throughout the training process, and the best-performing model on the validation set is saved as the prediction model for actual predictions.

[0068] In some embodiments, the property prediction method may further include: determining an error range corresponding to the target property of the target molecule based on the second embedding result, so as to screen the solvent of the electrolyte based on the target property and the error range.

[0069] The property prediction method provided by the embodiment of the present disclosure is a knowledge-data dual-driven method for predicting the melting point, boiling point and flash point of the solvent molecules in the electrolyte to be prepared. The machine learning method based on the dual drive of chemical field knowledge and data is applied to the prediction of the melting point, boiling point and flash point of the electrolyte molecules, which can effectively promote the design of electrolyte molecules, and can also reveal the intrinsic connection between the structure and performance of electrolyte molecules, and realize efficient and accurate molecular property prediction, thereby improving the efficiency and performance of electrolyte design. Powerful prediction and generalization capabilities can accurately predict melting point, boiling point and flash point. This property prediction method can be used to quickly predict a large number of candidate molecules, screen out target molecules that meet specific property requirements, and search for the molecules most similar to the target molecules in the specific property space through molecular neighbor search to provide potential optimization directions and design ideas. This property prediction method, from a data-driven perspective, enables automated data collection and structured processing. Its high interpretability allows quantitative identification of four molecular features influencing melting, boiling, and flash points, making it a powerful tool for constructing structure-activity relationships and discovering knowledge. Furthermore, it incorporates a knowledge-driven perspective, combining discovered chemical knowledge with deep learning models to achieve highly accurate predictions of molecular properties. Mean absolute errors of 10.4 K, 4.6 K, and 4.8 K were achieved in the predictions of melting, boiling, and flash points, respectively. This method can be applied to the development of rechargeable batteries with a wide temperature range and high safety. In summary, this method not only accurately predicts molecular properties and deepens our understanding of structure-activity relationships, but also serves as a versatile and practical tool. By integrating domain knowledge with data-driven algorithms, it applies artificial intelligence to specific scenarios, providing new approaches and insights for the design and optimization of electrolyte molecules.

[0070] The disclosed embodiments also provide a device for predicting molecular properties. Figure 2 FIG. 1 is a block diagram of a device for predicting molecular properties according to an embodiment of the present disclosure. Figure 2As shown, the molecular property prediction device 100 may include the following acquisition module 101, determination module 102, knowledge embedding module 103, and prediction module 104.

[0071] The acquisition module 101 is used to obtain a prediction task for a target molecule, and determine the target property to be predicted based on the prediction task. The target property includes at least one of a melting point, a boiling point, and a flash point. The target molecule is one of the optional solvents of the electrolyte to be prepared.

[0072] Determination module 102 is used to determine the atomic structure vector, bond structure vector, and molecular structure vector based on the molecular structure of the target molecule, wherein the atomic structure vector represents the structural information of all atoms in the target molecule, the bond structure vector represents the structural information of all chemical bonds in the target molecule, and the molecular structure vector represents the overall structural information of the target molecule.

[0073] The knowledge embedding module 103 is configured to perform knowledge embedding on the atomic structure vector and the bond structure vector based on a pre-obtained knowledge vector to obtain a first embedding result; and perform knowledge embedding on a vector determined based on the first embedding result and the molecular structure vector based on the knowledge vector to obtain a second embedding result, wherein the knowledge vector is determined based on the molecular characteristics of each sample molecule, each of the molecular characteristics including the number of atoms, chemical bond properties, functional group category, and electronic properties, and each of the sample molecules is a solvent in an existing electrolyte.

[0074] The prediction module 104 is configured to determine the target property of the target molecule based on the second embedding result, so as to screen the solvent of the electrolyte according to the target property of the target molecule.

[0075] In one possible implementation, knowledge embedding is performed on the atomic structure vector and the bond structure vector based on a pre-obtained knowledge vector to obtain a first embedding result; knowledge embedding is performed on the vector determined according to the first embedding result and the molecular structure vector based on the knowledge vector to obtain a second embedding result, including: determining first knowledge used for knowledge embedding of the atomic structure vector, second knowledge used for knowledge embedding of the bond structure vector, and third knowledge used for knowledge embedding of the molecular structure vector based on the knowledge vector; embedding the first knowledge and the second knowledge into the atomic structure vector and the bond structure vector respectively to obtain the first embedding result; and embedding the third knowledge into the vector determined according to the first embedding result and the molecular structure vector to obtain the second embedding result.

[0076] In one possible implementation, the knowledge vector has multiple dimensions, and each dimension has at least one knowledge point; wherein, the device also includes an analysis module, which is used to: perform Shapley value analysis based on the knowledge vector to obtain the contribution value corresponding to each dimension in the knowledge vector, and the contribution value corresponding to each dimension represents the degree of influence of the knowledge point of the dimension on the molecular properties; determine at least one target dimension whose contribution value is greater than a preset threshold from all dimensions of the knowledge vector; and determine the first knowledge, the second knowledge, and the third knowledge based on part or all of the knowledge points in each target dimension.

[0077] In one possible implementation, the first knowledge and the second knowledge are correspondingly embedded into the atomic structure vector and the bond structure vector to obtain the first embedding result, including: embedding the first knowledge into the atomic structure vector to obtain a first embedding vector; and embedding the second knowledge into the bond structure vector to obtain a second embedding vector; and using the first embedding vector and the second embedding vector as the first embedding result.

[0078] In one possible implementation, the first knowledge and the second knowledge are correspondingly embedded into the atomic structure vector and the bond structure vector to obtain the first embedding result, including: embedding the first knowledge into the atomic structure vector to obtain a first embedding vector; embedding the second knowledge into a vector determined according to the first embedding vector and the bond structure vector to obtain a third embedding vector; and using the first embedding vector and the third embedding vector as the first embedding result.

[0079] In one possible implementation, the device also includes a training module for performing a training process for the prediction model, and the training process includes: obtaining molecular data of each of the sample molecules, each of the molecular data is used to represent the molecular structure and molecular properties of the corresponding sample molecule, and the molecular properties include melting point, boiling point, and flash point; performing feature extraction on each of the molecular data to obtain molecular features of the corresponding sample molecules, and determining the knowledge vector based on all molecular features, wherein the knowledge vector has multiple dimensions, and each dimension has at least one knowledge point; extracting some knowledge points in each of the dimensions to obtain a first knowledge point set, and using the first knowledge point set and the sample set to train the original prediction model to obtain a first model; using the second knowledge point set and the sample set to train the first model to obtain a trained prediction model, wherein the second knowledge point set includes knowledge points under the target dimension, and the target dimension is obtained based on Shapley value analysis of the knowledge vector.

[0080] In a possible implementation, the device further includes a range determination module, configured to determine an error range corresponding to the target property of the target molecule based on the second embedding result.

[0081] In some embodiments, the functions or modules included in the molecular property prediction device provided in the embodiments of the present disclosure can be used to execute the method described in the above method embodiment. Its specific implementation can refer to the description of the above molecular property prediction method embodiment. For the sake of brevity, it will not be repeated here.

[0082] The present disclosure also provides a method for designing an electrolyte, comprising: obtaining a target molecule, which is an optional solvent for the electrolyte to be prepared; determining the target properties of the target molecule using the above-mentioned molecular property prediction method, wherein the target properties include at least one of a melting point, a boiling point, and a flash point; if the target properties meet the preset electrolyte design conditions, the electrolyte is prepared based on the target molecule. In some embodiments, the functions or modules included in the electrolyte design method provided by the present disclosure can be used to execute the method described in the above method embodiment. Its specific implementation can refer to the description of the above molecular property prediction method embodiment. For the sake of brevity, it will not be repeated here.

[0083] The present disclosure also provides a secondary battery, wherein the electrolyte used in the secondary battery is determined by the electrolyte design method described above. In some embodiments, the functions or modules included in the secondary battery provided by the present disclosure can be used to perform the method described in the method embodiment above. The specific implementation thereof can be referred to the description of the molecular property prediction method embodiment above, and for the sake of brevity, it is not further described here.

[0084] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.

[0085] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0086] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0087] Figure 31900 is a block diagram of a molecular property prediction device according to an embodiment of the present disclosure. For example, the device 1900 can be provided as a server or a terminal device. Figure 3 The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0088] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2003. TM , MacOS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.

[0089] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the apparatus 1900 to perform the above-described method.

[0090] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0091] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0092] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0093] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0094] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0095] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0096] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0097] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0098] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for predicting molecular properties, characterized in that: The method utilizes a prediction model based on knowledge embedding to process a prediction task for a target molecule, and the process comprises the following steps: Obtaining a prediction task for a target molecule, and determining a target property to be predicted based on the prediction task, wherein the target property includes at least one of a melting point, a boiling point, and a flash point, and the target molecule is one of optional solvents for an electrolyte to be prepared; Based on the molecular structure of the target molecule, determining an atomic structure vector, a bond structure vector, and a molecular structure vector, wherein the atomic structure vector represents structural information of all atoms in the target molecule, the bond structure vector represents structural information of all chemical bonds in the target molecule, and the molecular structure vector represents overall structural information of the target molecule; performing knowledge embedding on the atomic structure vector and the bond structure vector based on a pre-obtained knowledge vector to obtain a first embedding result; performing knowledge embedding on a vector determined based on the first embedding result and the molecular structure vector based on the knowledge vector to obtain a second embedding result, wherein the knowledge vector is determined based on molecular characteristics of each sample molecule, each of the molecular characteristics including the number of atoms, chemical bond properties, functional group type, and electronic properties, and each of the sample molecules is a solvent in an existing electrolyte; The target property of the target molecule is determined based on the second embedding result, so as to screen the solvent of the electrolyte according to the target property of the target molecule.

2. The method according to claim 1, characterized in that Performing knowledge embedding on the atomic structure vector and the bond structure vector based on the pre-obtained knowledge vector to obtain a first embedding result; Performing knowledge embedding on a vector determined according to the first embedding result and the molecular structure vector based on the knowledge vector to obtain a second embedding result includes: Determining, based on the knowledge vectors, first knowledge for embedding knowledge into the atomic structure vector, second knowledge for embedding knowledge into the bond structure vector, and third knowledge for embedding knowledge into the molecular structure vector; Embedding the first knowledge and the second knowledge into the atomic structure vector and the bond structure vector respectively to obtain the first embedding result; The third knowledge is embedded into a vector determined according to the first embedding result and the molecular structure vector to obtain the second embedding result.

3. The method according to claim 2, characterized in that The knowledge vector has multiple dimensions, each dimension having at least one knowledge point; wherein the method further comprises: Performing Shapley value analysis based on the knowledge vector to obtain a contribution value corresponding to each dimension in the knowledge vector, wherein the contribution value corresponding to each dimension represents the degree of influence of the knowledge point in the dimension on the molecular properties; Determining at least one target dimension whose contribution value is greater than a preset threshold from all dimensions of the knowledge vector; Based on part or all of the knowledge points in each of the target dimensions, the first knowledge, the second knowledge, and the third knowledge are determined.

4. The method according to claim 2 or 3, characterized in that Embedding the first knowledge and the second knowledge into the atomic structure vector and the bond structure vector respectively to obtain the first embedding result includes: Embedding the first knowledge into the atomic structure vector to obtain a first embedding vector; and embedding the second knowledge into the bond structure vector to obtain a second embedding vector; The first embedding vector and the second embedding vector are used as the first embedding result.

5. The method according to claim 2 or 3, characterized in that Embedding the first knowledge and the second knowledge into the atomic structure vector and the bond structure vector respectively to obtain the first embedding result includes: Embedding the first knowledge into the atomic structure vector to obtain a first embedding vector; Embedding the second knowledge into a vector determined according to the first embedding vector and the key structure vector to obtain a third embedding vector; The first embedding vector and the third embedding vector are used as the first embedding result.

6. The method according to claim 1 or 2, characterized in that The method further includes a training process of the prediction model, wherein the training process includes the following steps: Acquiring molecular data of each sample molecule, wherein each molecular data is used to represent the molecular structure and molecular properties of the corresponding sample molecule, wherein the molecular properties include melting point, boiling point, and flash point; Performing feature extraction on each of the molecular data to obtain molecular features of the corresponding sample molecules, and determining the knowledge vector based on all the molecular features, wherein the knowledge vector has multiple dimensions, and each dimension has at least one knowledge point; Extracting some knowledge points from each dimension to obtain a first knowledge point set, and training an original prediction model using the first knowledge point set and the sample set to obtain a first model; The first model is trained using a second knowledge point set and the sample set to obtain a trained prediction model, wherein the second knowledge point set includes knowledge points under a target dimension, and the target dimension is obtained based on a Shapley value analysis of the knowledge vector.

7. The method according to claim 1 or 2, characterized in that The method further comprises: An error range corresponding to the target property of the target molecule is determined based on the second embedding result.

8. A method for designing an electrolyte, characterized in that: include: obtaining a target molecule, wherein the target molecule is an optional solvent for an electrolyte to be prepared; Determining a target property of the target molecule using the method according to any one of claims 1 to 7, wherein the target property comprises at least one of a melting point, a boiling point, and a flash point; If the target properties meet the preset electrolyte design conditions, the electrolyte is prepared based on the target molecule.

9. A secondary battery, characterized in that: The electrolyte used in the secondary battery is determined by the electrolyte design method according to claim 8.

10. A molecular property prediction device, characterized in that: include: an acquisition module, configured to acquire a prediction task for a target molecule, and determine a target property to be predicted based on the prediction task, wherein the target property includes at least one of a melting point, a boiling point, and a flash point, and the target molecule is one of the optional solvents for the electrolyte to be prepared; a determination module for determining an atomic structure vector, a bond structure vector, and a molecular structure vector based on the molecular structure of the target molecule, wherein the atomic structure vector represents the structural information of all atoms in the target molecule, the bond structure vector represents the structural information of all chemical bonds in the target molecule, and the molecular structure vector represents the overall structural information of the target molecule; a knowledge embedding module, configured to perform knowledge embedding on the atomic structure vector and the bond structure vector based on a pre-obtained knowledge vector to obtain a first embedding result; performing knowledge embedding on a vector determined according to the first embedding result and the molecular structure vector based on the knowledge vector to obtain a second embedding result, wherein the knowledge vector is determined based on molecular features of each sample molecule, each of the molecular features including the number of atoms, chemical bond properties, functional group type, and electronic properties, and each of the sample molecules is a solvent in an existing electrolyte; A prediction module is used to determine the target property of the target molecule based on the second embedding result, so as to screen the solvent of the electrolyte according to the target property of the target molecule.

Citation Information

Patent Citations

  • Knowledge tracking model training method and device, knowledge tracking method and device, equipment and medium

    CN114707775A

  • Molecular property prediction method based on chemical element knowledge graph and functional group prompt

    CN115762657A