A process and apparatus for predicting physical properties of electrolyte formulation components
By constructing a molecular property prediction model based on the Uni-Mol model and nonlinear regression model, the problems of high complexity and long cycle in the physical property analysis of electrolyte formula components were solved, and efficient physical property prediction was achieved.
Patent Information
- Application Number
- CN202411715296.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-11-27
AI Technical Summary
The existing technology for analyzing the physical properties of electrolyte formula components is complex, has a long cycle, and is inefficient.
A molecular property prediction model based on the Uni-Mol model and seven nonlinear regression models is constructed. Through model training, seven types of physical properties of one-, two-, and three-dimensional molecular structures are predicted, and a property prediction sequence is output.
It reduces the analysis complexity, shortens the analysis cycle, and improves the analysis efficiency.
Smart Images

Figure CN119560051B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a processing method and device for predicting physical properties of electrolyte formula components. BACKGROUND
[0002] An electrolyte formula is composed of three components, namely solvent, electrolyte and additive, each of which corresponds to a molecular structure. When designing and developing an electrolyte formula, the physical properties (such as melting point, boiling point, vapor pressure, dielectric constant, refractive index, density, and syntheticity) of each component of the formula need to be analyzed. Currently, there are two general analysis methods: simulation-based analysis and experiment-based analysis. The simulation-based analysis method generally includes the following steps: first, using a chemical informatics tool (such as OpenBabel, RDKit, etc.) to perform three-dimensional modeling based on the molecular information (such as one-dimensional molecular sequence and two-dimensional molecular topology) of each formula component, then using a molecular dynamics (MD) simulation tool (such as LAMMPS, GROMACS, NAMD, etc.) to simulate the molecular motion based on the modeling structure to obtain the corresponding simulation structure, calculating the properties based on the trajectory composed of multiple simulation structures and the energy, force, dipole moment, etc. generated in the simulation, or using a quantum chemistry calculation tool (such as Gaussian, GAMESS, etc.) to calculate the properties based on the simulation structure. The experiment-based analysis method generally includes the following steps: first, preparing the components based on the molecular information (such as one-dimensional molecular sequence) of the formula components, and analyzing the syntheticity based on the preparation effect, and when the syntheticity analysis result meets the standard, measuring other physical properties (such as melting point, boiling point, vapor pressure, dielectric constant, refractive index, density, etc.) of the formula components based on a series of experimental measurement methods. In practice, we found that both of the above analysis methods have the defects of complex analysis operation, long analysis period, and low analysis efficiency. SUMMARY
[0003] The present application aims at the defects of the prior art, and provides a processing method and device for predicting physical properties of electrolyte formula components, electronic equipment and computer readable storage medium. The present application first constructs a molecular property prediction model based on a Uni-Mol model and seven nonlinear regression models, and then trains the model to enable the prediction of seven types of physical properties (melting point, boiling point, vapor pressure, dielectric constant, refractive index, density and synthesizability) of one / two / three-dimensional molecular structure and output corresponding property prediction sequence. After the model training is completed, the molecular property prediction model is used to analyze the seven types of physical properties of all formula components of any electrolyte formula input by the user, and the obtained property analysis report is fed back to the user. Through the component property analysis of the electrolyte formula by the present application, the analysis complexity can be reduced, the analysis period can be shortened, and the analysis efficiency can be improved.
[0004] To achieve the above-mentioned object, the first aspect of the embodiment of the present application provides a processing method for predicting physical properties of electrolyte formula components, which comprises:
[0005] A molecular property prediction model is constructed. The molecular property prediction model is used to predict seven types of physical properties according to input formula component information and output corresponding property prediction sequence. The formula component information includes molecular structure type and molecular structure information. The molecular structure type includes one-dimensional, two-dimensional and three-dimensional types. When the molecular structure type is one-dimensional type, the corresponding molecular structure information is a one-dimensional molecular sequence. When the molecular structure type is two-dimensional type, the corresponding molecular structure information is a two-dimensional molecular topology graph. When the molecular structure type is three-dimensional type, the corresponding molecular structure information is a three-dimensional molecular conformation. The seven types of physical properties include melting point, boiling point, vapor pressure, dielectric constant, refractive index, density and synthesizability. The property prediction sequence includes melting point prediction value, boiling point prediction value, vapor pressure prediction value, dielectric constant prediction value, refractive index prediction value, density prediction value and synthesizability prediction value.
[0006] training the molecular property prediction model based on a preset first data set; wherein the first data set comprises a plurality of first data records; the first data record comprises a first training molecular structure type, first training molecular structure information and a first property label sequence; the first training molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the first training molecular structure type is one-dimensional, the corresponding first training molecular structure information is a one-dimensional molecular sequence; when the first training molecular structure type is two-dimensional, the corresponding first training molecular structure information is a two-dimensional molecular topology graph; when the first training molecular structure type is three-dimensional, the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first property label sequence comprises a first melting point label value, a first boiling point label value, a first vapor pressure label value, a first dielectric constant label value, a first refractive index label value, a first density label value and a first synthesizability label value;
[0007] After the model training is completed, an arbitrary electrolyte formula input by a user is received as a corresponding first formula; and the molecular property prediction model is used to perform electrolyte formula component physical property prediction processing on the first formula to obtain a corresponding first property analysis report to feed back to the user; wherein the first formula comprises a plurality of first formula component information; the first formula component information comprises a first molecular structure type and first molecular structure information; the first molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the first molecular structure type is one-dimensional, the corresponding first molecular structure information is a one-dimensional molecular sequence; when the first molecular structure type is two-dimensional, the corresponding first molecular structure information is a two-dimensional molecular topology graph; when the first molecular structure type is three-dimensional, the corresponding first molecular structure information is a three-dimensional molecular conformation; the first property analysis report comprises a plurality of first component analysis records; the first component analysis record comprises the first formula component information and a corresponding first property analysis sequence.
[0008] Preferably, the model input end of the molecular property prediction model is used to receive the input formula component information, and the model output end is used to output the corresponding property prediction sequence;
[0009] The molecular property prediction model comprises a molecular structure initialization module, a molecular structure optimization module, a melting point prediction module, a boiling point prediction module, a vapor pressure prediction module, a dielectric constant prediction module, a refractive index prediction module, a density prediction module, a synthesizability prediction module and a molecular property output module;
[0010] The input end of the molecular structure initialization module is connected with the model input end, and the output end is connected with the input end of the molecular structure optimization module; the output end of the molecular structure optimization module is connected with the input ends of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module and the synthesizability prediction module; the seven input ends of the molecular property output module are connected with the output ends of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module and the synthesizability prediction module, and the output end of the molecular property output module is connected with the model output end;
[0011] The molecular structure initialization module is used to extract the corresponding molecular structure type and molecular structure information from the formula ingredient information as the corresponding current type and current structure information, and identify the current type. If the current type is a one-dimensional type, the current structure information is taken as a corresponding first molecular sequence, and a one-dimensional sequence processing interface provided by a preset chemoinformatics tool is used to perform three-dimensional molecular structure modeling processing according to the first molecular sequence to obtain a corresponding initial molecular conformation. If the current type is a two-dimensional type, the current structure information is taken as a corresponding first molecular topology graph, and a two-dimensional topology graph processing interface provided by the chemoinformatics tool is used to perform three-dimensional molecular structure modeling processing according to the first molecular topology graph to obtain the corresponding initial molecular conformation. If the current type is a three-dimensional type, the current structure information is taken as the initial molecular conformation. The obtained initial molecular conformation is sent to the molecular structure optimization module. The chemoinformatics tool at least includes OpenBabel and RDKit. The one-dimensional sequence processing interface is a processing interface provided by the chemoinformatics tool, which is used to convert a one-dimensional molecular sequence input by the interface into a corresponding three-dimensional molecular conformation and output the obtained three-dimensional molecular conformation as interface output data. The two-dimensional topology graph processing interface is a processing interface provided by the chemoinformatics tool, which is used to convert a two-dimensional molecular topology graph input by the interface into a corresponding three-dimensional molecular conformation and output the obtained three-dimensional molecular conformation as interface output data.
[0012] The molecular structure optimization module is implemented based on a Uni-Mol model that has completed model pre-training; the molecular structure optimization module is used to perform three-dimensional structure optimization processing on the initialized molecular conformation to obtain a corresponding optimized molecular conformation; and the optimized molecular conformation is sent to the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module, and the synthesizability prediction module, respectively;
[0013] The melting point prediction module is implemented based on a first nonlinear regression model; the melting point prediction module is used to perform melting point prediction processing according to the optimized molecular conformation to obtain a corresponding melting point prediction value, which is sent to the molecular property output module; the first nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model;
[0014] The boiling point prediction module is implemented based on a second nonlinear regression model; the boiling point prediction module is used to perform boiling point prediction processing according to the optimized molecular conformation to obtain a corresponding boiling point prediction value, which is sent to the molecular property output module; the second nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model;
[0015] The vapor pressure prediction module is implemented based on a third nonlinear regression model; the vapor pressure prediction module is used to perform vapor pressure prediction processing according to the optimized molecular conformation to obtain a corresponding vapor pressure prediction value, which is sent to the molecular property output module; the third nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model;
[0016] The dielectric constant prediction module is implemented based on a fourth nonlinear regression model; the dielectric constant prediction module is used to perform dielectric constant prediction processing according to the optimized molecular conformation to obtain a corresponding dielectric constant prediction value, which is sent to the molecular property output module; the fourth nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model;
[0017] The refractive index prediction module is implemented based on a fifth nonlinear regression model; the refractive index prediction module is used to perform refractive index prediction processing according to the optimized molecular conformation to obtain a corresponding refractive index prediction value, which is sent to the molecular property output module; the fifth nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model;
[0018] The density prediction module is implemented based on a sixth nonlinear regression model; the density prediction module is configured to perform density prediction processing according to the optimized molecular conformation to obtain the corresponding density prediction value and send the density prediction value to the molecular property output module; the sixth nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0019] The synthetic degree prediction module is implemented based on a seventh nonlinear regression model; the synthetic degree prediction module is configured to perform synthetic degree prediction processing according to the optimized molecular conformation to obtain the corresponding synthetic degree prediction value and send the synthetic degree prediction value to the molecular property output module; the seventh nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0020] The molecular property output module is configured to obtain the corresponding property prediction sequence composed of the melting point prediction value, the boiling point prediction value, the vapor pressure prediction value, the dielectric constant prediction value, the refractive index prediction value, the density prediction value and the synthetic degree prediction value and output the property prediction sequence.
[0021] Preferably, the molecular property prediction model is trained based on a preset first data set, and specifically includes:
[0022] Step 31, the first data set is randomly divided based on a preset first test evaluation ratio to obtain a corresponding first test set and a first evaluation set;
[0023] The first test set and the first evaluation set are both composed of a plurality of first data records; the ratio of the total number of records of the first test set to the total number of records of the first evaluation set satisfies the first test evaluation ratio;
[0024] Step 32, setting a corresponding first debugging object as a nonlinear regression model; and extracting a first first data record of the first test set as a corresponding current test record;
[0025] The first debugging object includes a nonlinear regression model and a structure optimization model.
[0026] Step 33, a corresponding first training formula component information is composed of the first training molecular structure type and the first training molecular structure information of the current test record;
[0027] Step 34, the first training formula component information is input into the molecular property prediction model to perform seven types of physical property prediction and obtain a corresponding first property prediction sequence;
[0028] The first physical property prediction sequence includes a first melting point prediction value, a first boiling point prediction value, a first vapor pressure prediction value, a first dielectric constant prediction value, a first refractive index prediction value, a first density prediction value, and a first synthesizability prediction value.
[0029] Step 35, the first physical property prediction sequence and the first physical property label sequence of the current test record are brought into a preset first model loss function for calculation to obtain a corresponding first loss value;
[0030] The first model loss function at least includes an L1 loss function, an L2 loss function, and a cross-entropy loss function.
[0031] Step 36, whether the first loss value meets a preset first loss value range is identified; if the first loss value meets the first loss value range, whether the current test record is the last first data record of the first test set is identified, if yes, step 37 is turned to, if not, the next first data record of the first test set is extracted as a new current test record and step 33 is returned; if the first loss value does not meet the first loss value range, the first debugging object is identified, if the first debugging object is a nonlinear regression model, the model parameters of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module, and the synthesizability prediction module are modulated for one round based on a preset first model parameter optimizer towards the direction of minimizing the first model loss function on the premise that the model parameters of the molecular structure optimization module remain unchanged, and step 34 is returned at the end of this round of parameter modulation, if the first debugging object is a structure optimization model, the model parameters of the molecular structure optimization module are fine-tuned for one round based on a preset second model parameter optimizer towards the direction of minimizing the first model loss function on the premise that the model parameters of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module, and the synthesizability prediction module remain unchanged, and step 34 is returned at the end of this round of parameter fine-tuning.
[0032] The first model parameter optimizer at least includes an SGD optimizer, an ADAM optimizer; and the second model parameter optimizer at least includes an RMSprop optimizer, an AdamW optimizer, and an ADAM optimizer.
[0033] Step 37, identifying whether the first debugging object is a structure optimization model; if yes, going to step 38; if no, setting the first debugging object as a structure optimization model, extracting the first data record of the first test set as a new current test record, and returning to step 33;
[0034] Step 38, traversing all the first data records of the first evaluation set; and in the traversal process, taking the first data record currently traversed as a corresponding current evaluation record; and taking the first training molecular structure type and the first training molecular structure information of the current evaluation record to form a corresponding second training formula component information; and inputting the second training formula component information into the molecular property prediction model to perform seven types of physical property prediction and obtain a corresponding second property prediction sequence; and taking the second property prediction sequence and the first property label sequence of the current test record to form a corresponding prediction-label pair; and at the end of the traversal, taking all the obtained prediction-label pairs into a preset first model evaluation function to obtain a corresponding first evaluation value;
[0035] Among them, the first model evaluation function at least includes the RMSE function;
[0036] Step 39, identifying whether the first evaluation value meets a preset first evaluation value range; if not, returning to step 31 to continue training; if yes, stopping training and confirming that the model training is completed.
[0037] Preferably, the user is fed back the corresponding first property analysis report obtained by using the molecular property prediction model to perform the physical property prediction processing on the first formula, specifically including:
[0038] Inputting each first formula component information of the first formula into the molecular property prediction model to perform seven types of physical property prediction and taking the obtained property prediction sequence this time as a corresponding first property analysis sequence; and taking each first formula component information and the corresponding first property analysis sequence to form a corresponding first component analysis record; and taking all the obtained first component analysis records to form a corresponding first property analysis report to feed back to the user.
[0039] The second aspect of the embodiment of the application provides a device for implementing the processing method for predicting the physical properties of electrolyte formula components in the first aspect, and the device comprises a model construction module, a model training module and a model application module.
[0040] The model construction module is configured to construct a molecular property prediction model; the molecular property prediction model is configured to perform seven types of physical property prediction and output corresponding property prediction sequences according to inputted formula ingredient information; wherein the formula ingredient information comprises a molecular structure type and molecular structure information; the molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the molecular structure type is one-dimensional type, the corresponding molecular structure information is a one-dimensional molecular sequence; when the molecular structure type is two-dimensional type, the corresponding molecular structure information is a two-dimensional molecular topology graph; when the molecular structure type is three-dimensional type, the corresponding molecular structure information is a three-dimensional molecular conformation; the seven types of physical properties comprise melting point, boiling point, vapor pressure, dielectric constant, refractive index, density and synthesizability; the property prediction sequence comprises a melting point prediction value, a boiling point prediction value, a vapor pressure prediction value, a dielectric constant prediction value, a refractive index prediction value, a density prediction value and a synthesizability prediction value;
[0041] The model training module is configured to perform model training on the molecular property prediction model based on a preset first data set; wherein the first data set comprises a plurality of first data records; the first data record comprises a first training molecular structure type, first training molecular structure information and a first property label sequence; the first training molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the first training molecular structure type is one-dimensional type, the corresponding first training molecular structure information is a one-dimensional molecular sequence; when the first training molecular structure type is two-dimensional type, the corresponding first training molecular structure information is a two-dimensional molecular topology graph; when the first training molecular structure type is three-dimensional type, the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first property label sequence comprises a first melting point label value, a first boiling point label value, a first vapor pressure label value, a first dielectric constant label value, a first refractive index label value, a first density label value and a first synthesizability label value;
[0042] The model application module is configured to receive a user inputted arbitrary electrolyte formula as a corresponding first formula after model training is completed, and use the molecular property prediction model to perform electrolyte formula component physical property prediction processing on the first formula to obtain a corresponding first physical property analysis report for feedback to the user, wherein the first formula comprises a plurality of first formula component information, the first formula component information comprises a first molecular structure type and first molecular structure information, the first molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types, the first molecular structure information corresponding to the one-dimensional type is a one-dimensional molecular sequence, the first molecular structure information corresponding to the two-dimensional type is a two-dimensional molecular topology graph, and the first molecular structure information corresponding to the three-dimensional type is a three-dimensional molecular conformation, and the first physical property analysis report comprises a plurality of first component analysis records, and the first component analysis record comprises the first formula component information and a corresponding first physical property analysis sequence.
[0043] The electronic device provided in the third aspect of the embodiment of the present application comprises a memory, a processor and a transceiver.
[0044] The processor is coupled with the memory, reads and executes instructions in the memory to realize the method steps of the first aspect.
[0045] The transceiver is coupled with the processor, and the transceiver is controlled by the processor to perform message transceiving.
[0046] The fourth aspect of the embodiment of the present application provides a computer readable storage medium, which stores computer instructions, and when the computer instructions are executed by a computer, the computer instructions make the computer execute the instructions of the method of the first aspect.
[0047] The embodiment of the present application provides a processing method and device for predicting physical properties of electrolyte formula components, an electronic device and a computer readable storage medium. As known from the above, the embodiment of the present application first constructs a molecular property prediction model based on a Uni-Mol model and seven nonlinear regression models, then through model training, the model can predict seven types of physical properties (melting point, boiling point, vapor pressure, dielectric constant, refractive index, density and synthesizability) of one / two / three-dimensional molecular structures and output corresponding physical property prediction sequences, and after model training is completed, the molecular property prediction model is used to analyze seven types of physical properties of all formula components of an arbitrary electrolyte formula inputted by a user and feed back the obtained physical property analysis report to the user. Through the component physical property analysis of the electrolyte formula in the embodiment of the present application, the analysis complexity is reduced, the analysis period is shortened, and the analysis efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 A processing method for predicting physical properties of electrolyte formula components provided for the first embodiment of the present application is shown in the schematic diagram;
[0049] Figure 2 A module schematic diagram of the molecular property prediction model provided for the first embodiment of the present application is shown in the schematic diagram;
[0050] Figure 3 A module structure diagram of the processing device for predicting physical properties of electrolyte formula components provided for the second embodiment of the present application is shown in the schematic diagram;
[0051] Figure 4 A structure schematic diagram of an electronic device provided for the third embodiment of the present application is shown in the schematic diagram. DETAILED DESCRIPTION
[0052] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0053] The first embodiment of the present application provides a processing method for predicting physical properties of electrolyte formula components, which is shown in the schematic diagram. Figure 1 The processing method for predicting physical properties of electrolyte formula components provided for the first embodiment of the present application is shown in the schematic diagram, and the method mainly includes the following steps:
[0054] Step 1, constructing a molecular property prediction model.
[0055] Here, the molecular property prediction model of the embodiment of the present application is used to perform seven types of physical property prediction according to the inputted formula component information and output corresponding property prediction sequences; wherein the formula component information comprises molecular structure type and molecular structure information; the molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; the molecular structure information corresponding to the one-dimensional type is a one-dimensional molecular sequence (such as SMILES sequence), the molecular structure information corresponding to the two-dimensional type is a two-dimensional molecular topology graph, and the molecular structure information corresponding to the three-dimensional type is a three-dimensional molecular conformation; the seven types of physical properties comprise melting point, boiling point, vapor pressure, dielectric constant, refractive index, density and synthesizability; the property prediction sequence comprises melting point prediction value, boiling point prediction value, vapor pressure prediction value, dielectric constant prediction value, refractive index prediction value, density prediction value and synthesizability prediction value.
[0056] As Figure 2 As shown in the module schematic diagram of the molecular property prediction model provided for the first embodiment of the present application, the model input end of the molecular property prediction model of the embodiment of the present application is used to receive the inputted formula component information, and the model output end is used to output the corresponding property prediction sequence; and the molecular property prediction model comprises a molecular structure initialization module, a molecular structure optimization module, a melting point prediction module, a boiling point prediction module, a vapor pressure prediction module, a dielectric constant prediction module, a refractive index prediction module, a density prediction module, a synthesizability prediction module and a molecular property output module.
[0057] The connection relationship of the modules of the molecular property prediction model is that the input end of the molecular structure initialization module is connected with the model input end, and the output end is connected with the input end of the molecular structure optimization module; the output end of the molecular structure optimization module is connected with the input end of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module and the synthesizability prediction module respectively; the seven input ends of the molecular property output module are connected with the output end of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module and the synthesizability prediction module respectively, and the output end of the molecular property output module is connected with the model output end.
[0058] It should be noted that the molecular structure initialization module of the molecular property prediction model of the embodiment of the application is used to extract the corresponding molecular structure type and molecular structure information from the formula ingredient information as the corresponding current type and current structure information, and identify the current type. If the current type is a one-dimensional type, the current structure information is used as the corresponding first molecular sequence, and a one-dimensional sequence processing interface provided by a preset chemoinformatics tool is used to perform three-dimensional molecular structure modeling processing according to the first molecular sequence to obtain the corresponding initialization molecular conformation. If the current type is a two-dimensional type, the current structure information is used as the corresponding first molecular topology graph, and a two-dimensional topology graph processing interface provided by the chemoinformatics tool is used to perform three-dimensional molecular structure modeling processing according to the first molecular topology graph to obtain the corresponding initialization molecular conformation. If the current type is a three-dimensional type, the current structure information is used as the corresponding initialization molecular conformation. The obtained initialization molecular conformation is sent to the molecular structure optimization module. The chemoinformatics tool of the embodiment of the application at least includes OpenBabel and RDKit. The one-dimensional sequence processing interface is a processing interface provided by a chemoinformatics tool, which is used to convert the one-dimensional molecular sequence input by the interface into a corresponding three-dimensional molecular conformation and output the obtained three-dimensional molecular conformation as interface output data. The two-dimensional topology graph processing interface is a processing interface provided by a chemoinformatics tool, which is used to convert the two-dimensional molecular topology graph input by the interface into a corresponding three-dimensional molecular conformation and output the obtained three-dimensional molecular conformation as interface output data.
[0059] It should be noted that the molecular structure optimization module of the molecular property prediction model of the embodiment of the application is implemented based on a Uni-Mol model that has completed model pre-training. The molecular structure optimization module is used to perform three-dimensional structure optimization processing on the initialization molecular conformation to obtain the corresponding optimized molecular conformation, and send the optimized molecular conformation to the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module and the synthesizability prediction module, respectively. The Uni-Mol model mentioned here can be understood with reference to the paper “UNI-MOL:A UNIVERSAL 3D MOLECULAR REPRESENTATION LEARNING FRAMEWORK”. The paper describes the pre-training method of the model, and the pre-training of the Uni-Mol model can be completed by referring to the paper content. The paper also points out that the Uni-Mol model can perform structure optimization on the three-dimensional molecular conformation, so the embodiment of the application implements the module function of the molecular structure optimization module based on a Uni-Mol model that has completed model pre-training.
[0060] It should be noted that the melting point prediction module of the molecular property prediction model in the embodiment of the present application is realized based on a first nonlinear regression model; the melting point prediction module is used to perform melting point prediction processing according to the optimized molecular conformation to obtain a corresponding melting point prediction value and send the melting point prediction value to the molecular property output module; here, the optional model types of the first nonlinear regression model in the embodiment of the present application at least include an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0061] It should be noted that the boiling point prediction module of the molecular property prediction model in the embodiment of the present application is realized based on a second nonlinear regression model; the boiling point prediction module is used to perform boiling point prediction processing according to the optimized molecular conformation to obtain a corresponding boiling point prediction value and send the boiling point prediction value to the molecular property output module; here, the optional model types of the second nonlinear regression model in the embodiment of the present application at least include an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0062] It should be noted that the vapor pressure prediction module of the molecular property prediction model in the embodiment of the present application is realized based on a third nonlinear regression model; the vapor pressure prediction module is used to perform vapor pressure prediction processing according to the optimized molecular conformation to obtain a corresponding vapor pressure prediction value and send the vapor pressure prediction value to the molecular property output module; here, the optional model types of the third nonlinear regression model in the embodiment of the present application at least include an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0063] It should be noted that the dielectric constant prediction module of the molecular property prediction model in the embodiment of the present application is realized based on a fourth nonlinear regression model; the dielectric constant prediction module is used to perform dielectric constant prediction processing according to the optimized molecular conformation to obtain a corresponding dielectric constant prediction value and send the dielectric constant prediction value to the molecular property output module; here, the optional model types of the fourth nonlinear regression model in the embodiment of the present application at least include an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0064] It should be noted that the refractive index prediction module of the molecular property prediction model in the embodiment of the present application is realized based on a fifth nonlinear regression model; the refractive index prediction module is used to perform refractive index prediction processing according to the optimized molecular conformation to obtain a corresponding refractive index prediction value and send the refractive index prediction value to the molecular property output module; here, the optional model types of the fifth nonlinear regression model in the embodiment of the present application at least include an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0065] It should be noted that the density prediction module of the molecular property prediction model in the embodiment of the application is implemented based on a sixth nonlinear regression model; the density prediction module is used for performing density prediction processing according to the optimized molecular conformation to obtain a corresponding density prediction value and sending the density prediction value to the molecular property output module; here, the optional model types of the sixth nonlinear regression model in the embodiment of the application at least include an XGBoost model, a GBDT model, a Random Forest model and an MLP model.
[0066] It should be noted that the synthetic degree prediction module of the molecular property prediction model in the embodiment of the application is implemented based on a seventh nonlinear regression model; the synthetic degree prediction module is used for performing synthetic degree prediction processing according to the optimized molecular conformation to obtain a corresponding synthetic degree prediction value and sending the synthetic degree prediction value to the molecular property output module; here, the optional model types of the seventh nonlinear regression model in the embodiment of the application at least include an XGBoost model, a GBDT model, a Random Forest model and an MLP model.
[0067] It should be noted that the molecular property output module of the molecular property prediction model in the embodiment of the application is used for outputting a corresponding property prediction sequence composed of the obtained melting point prediction value, boiling point prediction value, vapor pressure prediction value, dielectric constant prediction value, refractive index prediction value, density prediction value and synthetic degree prediction value.
[0068] Step 2, model training is performed on the molecular property prediction model based on a preset first data set;
[0069] The first data set in the embodiment of the application is a model training data set constructed in advance through big data collection and is composed of a plurality of first data records; the first data record includes a first training molecular structure type, first training molecular structure information and a first property label sequence; the first training molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; when the first training molecular structure type is a one-dimensional type, the corresponding first training molecular structure information is a one-dimensional molecular sequence, when the first training molecular structure type is a two-dimensional type, the corresponding first training molecular structure information is a two-dimensional molecular topological graph, and when the first training molecular structure type is a three-dimensional type, the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first property label sequence includes a first melting point label value, a first boiling point label value, a first vapor pressure label value, a first dielectric constant label value, a first refractive index label value, a first density label value and a first synthetic degree label value.
[0070] Specifically, step 21, the first data set is randomly divided based on a preset first test evaluation ratio to obtain a corresponding first test set and a first evaluation set;
[0071] The first test evaluation ratio is a preset ratio value, for example, 8:2; the first test set and the first evaluation set each comprise a plurality of first data records; a ratio of a total number of records in the first test set to a total number of records in the first evaluation set satisfies the first test evaluation ratio;
[0072] Step 22, setting the corresponding first debugging object as a nonlinear regression model; and extracting the first data record of the first test set as the corresponding current test record;
[0073] The first debugging object comprises a nonlinear regression model and a structure optimization model.
[0074] Step 23, a corresponding first training formula component information is composed of the first training molecular structure type and the first training molecular structure information of the current test record;
[0075] Step 24, the first training formula component information is input into a molecular property prediction model to perform seven types of physical property prediction and obtain a corresponding first physical property prediction sequence;
[0076] The first physical property prediction sequence comprises a first melting point prediction value, a first boiling point prediction value, a first vapor pressure prediction value, a first dielectric constant prediction value, a first refractive index prediction value, a first density prediction value and a first synthesizability prediction value.
[0077] Step 25, the first physical property prediction sequence and the first physical property label sequence of the current test record are brought into a preset first model loss function to obtain a corresponding first loss value;
[0078] The first model loss function at least comprises an L1 loss function, an L2 loss function and a cross-entropy loss function.
[0079] Step 26, identify whether the first loss value meets the preset first loss value range; if the first loss value meets the first loss value range, identify whether the current test record is the last first data record of the first test set, if yes, go to step 27, if no, extract the next first data record of the first test set as a new current test record and return to step 23; if the first loss value does not meet the first loss value range, identify the first debugging object, if the first debugging object is a nonlinear regression model, based on the preset first model parameter optimizer, on the premise that the model parameters of the molecular structure optimization module remain unchanged, one round of model parameter modulation is performed on the model parameters of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module and the synthesizability prediction module towards the direction of making the first model loss function reach the minimum value, and return to step 24 at the end of the round of parameter modulation, if the first debugging object is a structure optimization model, based on the preset second model parameter optimizer, on the premise that the model parameters of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module and the synthesizability prediction module remain unchanged, one round of model parameter fine tuning is performed on the model parameters of the molecular structure optimization module towards the direction of making the first model loss function reach the minimum value, and return to step 24 at the end of the round of parameter fine tuning;
[0080] Wherein, the first loss value range is a pre-set loss value range; the first model parameter optimizer at least includes an SGD optimizer, an ADAM optimizer; the second model parameter optimizer at least includes an RMSprop optimizer, an AdamW optimizer and an ADAM optimizer;
[0081] Step 27, identify whether the first debugging object is a structure optimization model; if yes, go to step 28; if no, set the first debugging object as a structure optimization model, and extract the first first data record of the first test set as a new current test record, and return to step 23;
[0082] Here, as can be seen from the current step, the embodiment of the present application will adopt a two-stage training method to reduce the training difficulty when training the molecular property prediction model; 1) in the first stage, set the first debugging object = nonlinear regression model, only the parameters of the 7 nonlinear regression models (melting point prediction module, boiling point prediction module, vapor pressure prediction module, dielectric constant prediction module, refractive index prediction module, density prediction module and synthesizability prediction module) are modulated, and the model parameters of the Uni-Mol model (molecular structure optimization module) are not processed; 2) in the second stage, set the first debugging object = structure optimization model, do not modulate the parameters of the 7 nonlinear regression models (melting point prediction module, boiling point prediction module, vapor pressure prediction module, dielectric constant prediction module, refractive index prediction module, density prediction module and synthesizability prediction module), and only fine-tune the model parameters of the Uni-Mol model (molecular structure optimization module);
[0083] Step 28, traverse all first data records of the first evaluation set; and in the traversal process, take the currently traversed first data record as the corresponding current evaluation record; and form a corresponding second training formula component information from the first training molecular structure type and the first training molecular structure information of the current evaluation record; and input the second training formula component information into the molecular property prediction model to perform seven types of physical property prediction and obtain a corresponding second physical property prediction sequence; and form a corresponding prediction-label pair from the second physical property prediction sequence and the first physical property label sequence of the current test record; and at the end of the traversal, all prediction-label pairs obtained are brought into the preset first model evaluation function to obtain a corresponding first evaluation value;
[0084] Among them, the first model evaluation function at least includes the RMSE function;
[0085] Step 29, identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 21 to continue training; if yes, stop training and confirm that the model training is completed;
[0086] Among them, the first evaluation value range is a pre-set evaluation value range.
[0087] Step 3, after the model training is completed, receive the user input of any electrolyte formula as a corresponding first formula; and use the molecular property prediction model to perform electrolyte formula component physical property prediction processing on the first formula to obtain a corresponding first physical property analysis report to feedback to the user;
[0088] Specifically includes: step 31, after the model training is completed, receive the user input of any electrolyte formula as a corresponding first formula;
[0089] The first formula includes a plurality of first formula component information; the first formula component information includes a first molecular structure type and first molecular structure information; the first molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the first molecular structure type corresponds to a one-dimensional molecular sequence when being a one-dimensional type, a two-dimensional molecular topology map when being a two-dimensional type, and a three-dimensional molecular conformation when being a three-dimensional type;
[0090] In step 32, the electrolyte formula component physical property prediction model is used to perform electrolyte formula component physical property prediction processing on the first formula to obtain a corresponding first physical property analysis report, which is fed back to the user.
[0091] The first physical property analysis report includes a plurality of first component analysis records; the first component analysis record includes first formula component information and a corresponding first physical property analysis sequence.
[0092] Specifically, each first formula component information of the first formula is input into the molecular physical property prediction model to perform seven types of physical property prediction, and the obtained physical property prediction sequence is taken as the corresponding first physical property analysis sequence; each first formula component information and the corresponding first physical property analysis sequence form a corresponding first component analysis record; and all obtained first component analysis records form a corresponding first physical property analysis report, which is fed back to the user.
[0093] Figure 3 A module structure diagram of a processing device for predicting physical properties of electrolyte formula components is provided for the second embodiment of the present application. The device is a terminal device or a server for implementing the foregoing method embodiments, or a device capable of enabling the foregoing terminal device or server to implement the foregoing method embodiments, such as a device or a chip system of the foregoing terminal device or server. As shown in the figure, the device includes a model construction module 201, a model training module 202 and a model application module 203. Figure 3
[0094] The model construction module 201 is configured to construct a molecular property prediction model; the molecular property prediction model is configured to perform seven types of physical property prediction according to inputted formula component information and output corresponding property prediction sequences; wherein the formula component information comprises molecular structure types and molecular structure information; the molecular structure types comprise one-dimensional, two-dimensional and three-dimensional types; when the molecular structure type is one-dimensional type, the corresponding molecular structure information is a one-dimensional molecular sequence; when the molecular structure type is two-dimensional type, the corresponding molecular structure information is a two-dimensional molecular topology graph; when the molecular structure type is three-dimensional type, the corresponding molecular structure information is a three-dimensional molecular conformation; the seven types of physical properties comprise melting point, boiling point, vapor pressure, dielectric constant, refractive index, density and synthesizability; the property prediction sequence comprises a melting point prediction value, a boiling point prediction value, a vapor pressure prediction value, a dielectric constant prediction value, a refractive index prediction value, a density prediction value and a synthesizability prediction value.
[0095] The model training module 202 is configured to perform model training on the molecular property prediction model based on a preset first data set; wherein the first data set comprises a plurality of first data records; the first data record comprises a first training molecular structure type, first training molecular structure information and a first property label sequence; the first training molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the first training molecular structure type is one-dimensional type, the corresponding first training molecular structure information is a one-dimensional molecular sequence; when the first training molecular structure type is two-dimensional type, the corresponding first training molecular structure information is a two-dimensional molecular topology graph; when the first training molecular structure type is three-dimensional type, the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first property label sequence comprises a first melting point label value, a first boiling point label value, a first vapor pressure label value, a first dielectric constant label value, a first refractive index label value, a first density label value and a first synthesizability label value.
[0096] The model application module 203 is configured to, after the model training is completed, receive a user inputted arbitrary electrolyte formula as a corresponding first formula; and perform electrolyte formula component physical property prediction processing on the first formula using the molecular property prediction model to obtain a corresponding first property analysis report and feed back to the user; wherein the first formula comprises a plurality of first formula component information; the first formula component information comprises a first molecular structure type and first molecular structure information; the first molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the first molecular structure type is one-dimensional type, the corresponding first molecular structure information is a one-dimensional molecular sequence; when the first molecular structure type is two-dimensional type, the corresponding first molecular structure information is a two-dimensional molecular topology graph; when the first molecular structure type is three-dimensional type, the corresponding first molecular structure information is a three-dimensional molecular conformation; the first property analysis report comprises a plurality of first component analysis records; the first component analysis record comprises the first formula component information and a corresponding first property analysis sequence.
[0097] The processing device for predicting physical properties of electrolyte formula components provided by the embodiment of the present application can execute the method steps in the method embodiments, and has similar implementation principles and technical effects, which will not be described here again.
[0098] It should be noted that the division of each module of the above device is only a logical division of functions, and all or part of the actual implementation can be integrated into one physical entity, or can be physically separated. And these modules can all be implemented in the form of software called by the processing element; all can be implemented in the form of hardware; some modules can be implemented in the form of software called by the processing element, and some modules can be implemented in the form of hardware. For example, the model construction module can be a separate processing element, or it can be integrated into a chip of the above device, in addition, it can also be stored in the form of program code in the memory of the above device, and the function of the above determination module is called and executed by a processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capability. In the implementation process, each step of the above method or each module can be completed by the integrated logic circuit of the hardware in the processor element or the instruction in the form of software.
[0099] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), etc. For another example, when a certain module above is implemented in the form of program code scheduled by the processing element, the processing element can be a general purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules can be integrated together to implement in the form of system on a chip (SOC).
[0100] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the above method embodiments are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The above-mentioned computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above-mentioned computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.) means. The above-mentioned computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The above-mentioned available medium can be a magnetic medium (such as a floppy disk, hard disk, tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0101] Figure 4 This is a schematic diagram of the structure of an electronic device provided in the third embodiment of the present invention. The electronic device can be a terminal device or server that implements the method of the aforementioned embodiment, or it can be a terminal device or server that implements the method of the aforementioned embodiment connected to the aforementioned terminal device or server. Figure 4 As shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver 303's transceiver actions. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the aforementioned embodiment method. Preferably, the electronic device involved in the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The above-mentioned communication port 306 is used for connecting and communicating between the electronic device and other peripherals.
[0102] exist Figure 4The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus. The communication interface is used to realize the communication between the database access device and other devices (such as clients, read-write libraries and read-only libraries). The memory can include Random Access Memory (RAM), and can also include Non-Volatile Memory, such as at least one disk memory.
[0103] The processor mentioned above can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; can also be a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0104] It should be noted that the embodiments of the present application also provide a computer readable storage medium, which stores instructions, when the instructions are run on a computer, the computer executes the method and process provided in the above embodiments.
[0105] The embodiments of the present application provide a processing method and device for predicting physical properties of electrolyte formula components, an electronic device and a computer readable storage medium. From the above content, it can be known that the embodiments of the present application first construct a molecular physical property prediction model based on a Uni-Mol model and seven nonlinear regression models, then make the model trained to be able to predict seven types of physical properties (melting point, boiling point, vapor pressure, dielectric constant, refractive index, density and synthesizability) of one / two / three-dimensional molecular structure and output corresponding physical property prediction sequences; after the model training is completed, the molecular physical property prediction model is used to analyze seven types of physical properties of all formula components of an arbitrary electrolyte formula input by a user and feed back a physical property analysis report obtained to the user. Through the component physical property analysis of the electrolyte formula in the embodiments of the present application, the analysis complexity is reduced, the analysis period is shortened, and the analysis efficiency is improved.
[0106] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and
[0107] The above detailed description describes the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A process for predicting physical properties of electrolyte formulation components, characterized by, The method comprises: constructing a molecular property prediction model; the molecular property prediction model is used for seven types of physical property prediction and output corresponding physical property prediction sequence according to input formula ingredient information; wherein, the formula ingredient information includes molecular structure type and molecular structure information; the molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the molecular structure type is one-dimensional type, and the corresponding molecular structure information is a one-dimensional molecular sequence; the molecular structure type is two-dimensional type, and the corresponding molecular structure information is a two-dimensional molecular topology graph; the molecular structure type is three-dimensional type, and the corresponding molecular structure information is a three-dimensional molecular conformation; the seven types of physical properties include melting point, boiling point, vapor pressure, dielectric constant, refractive index, density and synthesizability; the physical property prediction sequence includes melting point prediction value, boiling point prediction value, vapor pressure prediction value, dielectric constant prediction value, refractive index prediction value, density prediction value and synthesizability prediction value; based on a preset first data set, the molecular property prediction model is trained; wherein, the first data set includes a plurality of first data records; the first data record includes first training molecular structure type, first training molecular structure information and first physical property label sequence; the first training molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the first training molecular structure type is one-dimensional type, and the corresponding first training molecular structure information is a one-dimensional molecular sequence; the first training molecular structure type is two-dimensional type, and the corresponding first training molecular structure information is a two-dimensional molecular topology graph; the first training molecular structure type is three-dimensional type, and the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first physical property label sequence includes first melting point label value, first boiling point label value, first vapor pressure label value, first dielectric constant label value, first refractive index label value, first density label value and first synthesizability label value; after model training, receiving user input of any electrolyte formula as corresponding first formula; and using the molecular property prediction model to perform electrolyte formula ingredient physical property prediction processing on the first formula to obtain corresponding first physical property analysis report to feedback to the user; wherein, the first formula includes a plurality of first formula ingredient information; the first formula ingredient information includes first molecular structure type and first molecular structure information; the first molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the first molecular structure type is one-dimensional type, and the corresponding first molecular structure information is a one-dimensional molecular sequence; the first molecular structure type is two-dimensional type, and the corresponding first molecular structure information is a two-dimensional molecular topology graph; the first molecular structure type is three-dimensional type, and the corresponding first molecular structure information is a three-dimensional molecular conformation; the first physical property analysis report includes a plurality of first ingredient analysis records; the first ingredient analysis record includes the first formula ingredient information and the corresponding first physical property analysis sequence.
2. The processing method of claim 1, wherein the molecular property prediction model is configured to receive the inputted formula component information at a model input end and output corresponding property prediction sequences at a model output end. The molecular property prediction model comprises a molecular structure initialization module, a molecular structure optimization module, a melting point prediction module, a boiling point prediction module, a vapor pressure prediction module, a dielectric constant prediction module, a refractive index prediction module, a density prediction module, a synthesizability prediction module, and a molecular property output module. The input end of the molecular structure initialization module is connected to the model input end, and the output end is connected to the input end of the molecular structure optimization module. The output end of the molecular structure optimization module is connected to the input ends of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module, and the synthesizability prediction module. The seven input ends of the molecular property output module are connected to the output ends of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module, and the synthesizability prediction module. The molecular structure initialization module is configured to extract the molecular structure type and the molecular structure information from the formula component information as current type and current structure information, and identify the current type. If the current type is one-dimensional, the current structure information is taken as a first molecular sequence, and a one-dimensional sequence processing interface provided by a preset chemoinformatics tool is used to perform three-dimensional molecular structure modeling processing based on the first molecular sequence to obtain an initial molecular conformation. If the current type is two-dimensional, the current structure information is taken as a first molecular topology graph, and a two-dimensional topology graph processing interface provided by the chemoinformatics tool is used to perform three-dimensional molecular structure modeling processing based on the first molecular topology graph to obtain the initial molecular conformation. If the current type is three-dimensional, the current structure information is taken as the initial molecular conformation. The initial molecular conformation is sent to the molecular structure optimization module, and the chemoinformatics tool at least includes OpenBabel and RDKit. The one-dimensional sequence processing interface is a processing interface provided by a chemoinformatics tool, which is configured to convert a one-dimensional molecular sequence inputted by an interface into a three-dimensional molecular conformation and output the three-dimensional molecular conformation as interface output data. The two-dimensional topology graph processing interface is a processing interface provided by a chemoinformatics tool, which is configured to convert a two-dimensional molecular topology graph inputted by an interface into a three-dimensional molecular conformation and output the three-dimensional molecular conformation as interface output data. The molecular structure optimization module is implemented based on a Uni-Mol model that has completed model pre-training; the molecular structure optimization module is configured to perform three-dimensional structure optimization processing on the initialized molecular conformation to obtain a corresponding optimized molecular conformation; and the optimized molecular conformation is sent to the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module, and the synthesizability prediction module, respectively; The melting point prediction module is implemented based on a first nonlinear regression model; the melting point prediction module is configured to perform melting point prediction processing on the optimized molecular conformation to obtain a corresponding melting point prediction value, which is sent to the molecular property output module; the first nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model; The boiling point prediction module is implemented based on a second nonlinear regression model; the boiling point prediction module is configured to perform boiling point prediction processing on the optimized molecular conformation to obtain a corresponding boiling point prediction value, which is sent to the molecular property output module; the second nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model; The vapor pressure prediction module is implemented based on a third nonlinear regression model; the vapor pressure prediction module is configured to perform vapor pressure prediction processing on the optimized molecular conformation to obtain a corresponding vapor pressure prediction value, which is sent to the molecular property output module; the third nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model; The dielectric constant prediction module is implemented based on a fourth nonlinear regression model; the dielectric constant prediction module is configured to perform dielectric constant prediction processing on the optimized molecular conformation to obtain a corresponding dielectric constant prediction value, which is sent to the molecular property output module; the fourth nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model; The refractive index prediction module is implemented based on a fifth nonlinear regression model; the refractive index prediction module is configured to perform refractive index prediction processing on the optimized molecular conformation to obtain a corresponding refractive index prediction value, which is sent to the molecular property output module; the fifth nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model; The density prediction module is implemented based on a sixth nonlinear regression model; the density prediction module is configured to perform density prediction processing on the optimized molecular conformation to obtain a corresponding density prediction value, which is sent to the molecular property output module; the sixth nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model; The synthesizability prediction module is realized based on a seventh nonlinear regression model; the synthesizability prediction module is used for performing synthesizability prediction processing according to the optimized molecular conformation to obtain the corresponding synthesizability prediction value and sending the synthesizability prediction value to the molecular property output module; the seventh nonlinear regression model at least includes an XGBoost model, a GBDT model, a RandomForest model and an MLP model; The molecular property output module is used for composing the corresponding property prediction sequence from the obtained melting point prediction value, boiling point prediction value, vapor pressure prediction value, dielectric constant prediction value, refractive index prediction value, density prediction value and synthesizability prediction value and outputting the property prediction sequence.
3. The process for predicting physical properties of electrolyte formulation components according to claim 2, wherein, The molecular property prediction model is trained based on a preset first data set, and specifically includes the following steps: Step 31, the first data set is randomly divided based on a preset first test evaluation ratio to obtain a corresponding first test set and a first evaluation set; Wherein, the first test set and the first evaluation set are both composed of a plurality of first data records; the ratio of the total number of records of the first test set to the total number of records of the first evaluation set satisfies the first test evaluation ratio; Step 32, setting a corresponding first debugging object as a nonlinear regression model; and extracting a first first data record of the first test set as a corresponding current test record; Wherein, the first debugging object includes a nonlinear regression model and a structure optimization model; Step 33, a corresponding first training formula component information is composed of the first training molecular structure type and the first training molecular structure information of the current test record; Step 34, the first training formula component information is input into the molecular property prediction model to perform seven types of physical property prediction and obtain a corresponding first property prediction sequence; Wherein, the first property prediction sequence includes a first melting point prediction value, a first boiling point prediction value, a first vapor pressure prediction value, a first dielectric constant prediction value, a first refractive index prediction value, a first density prediction value and a first synthesizability prediction value; Step 35, the first property prediction sequence and the first property label sequence of the current test record are brought into a preset first model loss function to obtain a corresponding first loss value; Wherein, the first model loss function at least includes an L1 loss function, an L2 loss function and a cross-entropy loss function; Step 36, identify whether the first loss value meets the preset first loss value range; if the first loss value meets the first loss value range, identify whether the current test record is the last first data record of the first test set, if yes, go to step 37, if no, extract the next first data record of the first test set as a new current test record and return to step 33; if the first loss value does not meet the first loss value range, identify the first debugging object, if the first debugging object is a nonlinear regression model, based on a preset first model parameter optimizer, on the premise that the model parameters of the molecular structure optimization module remain unchanged, one round of model parameter modulation is performed on the model parameters of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module and the synthesizability prediction module towards the direction of minimizing the first model loss function, and at the end of this round of parameter modulation, return to step 34, if the first debugging object is a structure optimization model, based on a preset second model parameter optimizer, on the premise that the model parameters of the melting point prediction module, the boiling point prediction module, the vapor pressure prediction module, the dielectric constant prediction module, the refractive index prediction module, the density prediction module and the synthesizability prediction module remain unchanged, one round of model parameter fine-tuning is performed on the model parameters of the molecular structure optimization module towards the direction of minimizing the first model loss function, and at the end of this round of parameter fine-tuning, return to step 34; Wherein, the first model parameter optimizer at least includes SGD optimizer, ADAM optimizer; the second model parameter optimizer at least includes RMSprop optimizer, AdamW optimizer and ADAM optimizer; Step 37, identify whether the first debugging object is a structure optimization model; if yes, go to step 38; if no, set the first debugging object as a structure optimization model, and extract the first first data record of the first test set as a new current test record, and return to step 33; Step 38, traverse all the first data records of the first evaluation set; and in the traversal process, the first data record currently traversed is taken as a corresponding current evaluation record; and a corresponding second training formula component information is composed of the first training molecular structure type and the first training molecular structure information of the current evaluation record; and the second training formula component information is input into the molecular property prediction model to predict seven types of physical properties and obtain a corresponding second physical property prediction sequence; and a corresponding prediction-label pair is composed of the second physical property prediction sequence and the first physical property label sequence of the current test record; and at the end of the traversal, all the prediction-label pairs obtained are brought into a preset first model evaluation function to obtain a corresponding first evaluation value; The first model evaluation function at least includes an RMSE function; Step 39, whether the first evaluation value meets the preset first evaluation value range is identified; if not, return to step 31 to continue training; if yes, stop training and confirm that the model training is completed.
4. The process of predicting physical properties of electrolyte formulation components of claim 1, wherein, The corresponding first physical property analysis report is fed back to the user, specifically including: The first formula ingredient information of the first formula is input into the molecular physical property prediction model to perform seven types of physical property prediction, and the obtained physical property prediction sequence is taken as the corresponding first physical property analysis sequence; each first formula ingredient information and the corresponding first physical property analysis sequence form a corresponding first ingredient analysis record; and all the obtained first ingredient analysis records form the corresponding first physical property analysis report to feed back to the user.
5. An apparatus for performing the process of predicting the physical properties of the components of an electrolyte formulation according to any one of claims 1-4, characterized in that, The device comprises a model construction module, a model training module and a model application module; The model construction module is used to construct a molecular physical property prediction model; the molecular physical property prediction model is used to perform seven types of physical property prediction according to input formula ingredient information and output a corresponding physical property prediction sequence; wherein the formula ingredient information includes molecular structure type and molecular structure information; the molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; when the molecular structure type is one-dimensional type, the corresponding molecular structure information is a one-dimensional molecular sequence, when the molecular structure type is two-dimensional type, the corresponding molecular structure information is a two-dimensional molecular topology graph, and when the molecular structure type is three-dimensional type, the corresponding molecular structure information is a three-dimensional molecular conformation; the seven types of physical properties include melting point, boiling point, vapor pressure, dielectric constant, refractive index, density and synthesizability; the physical property prediction sequence includes melting point prediction value, boiling point prediction value, vapor pressure prediction value, dielectric constant prediction value, refractive index prediction value, density prediction value and synthesizability prediction value; The model training module is used to perform model training on the molecular physical property prediction model based on a preset first data set; wherein the first data set comprises a plurality of first data records; the first data record comprises a first training molecular structure type, first training molecular structure information and a first physical property label sequence; the first training molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; when the first training molecular structure type is one-dimensional type, the corresponding first training molecular structure information is a one-dimensional molecular sequence, when the first training molecular structure type is two-dimensional type, the corresponding first training molecular structure information is a two-dimensional molecular topology graph, and when the first training molecular structure type is three-dimensional type, the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first physical property label sequence includes a first melting point label value, a first boiling point label value, a first vapor pressure label value, a first dielectric constant label value, a first refractive index label value, a first density label value and a first synthesizability label value; The model application module is used to receive a user inputted arbitrary electrolyte formula as a corresponding first formula after the end of model training; and use the molecular property prediction model to perform electrolyte formula component physical property prediction processing on the first formula to obtain a corresponding first physical property analysis report to feed back to the user; wherein the first formula includes a plurality of first formula component information; the first formula component information includes a first molecular structure type and first molecular structure information; the first molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; when the first molecular structure type is a one-dimensional type, the corresponding first molecular structure information is a one-dimensional molecular sequence, when the first molecular structure type is a two-dimensional type, the corresponding first molecular structure information is a two-dimensional molecular topology graph, and when the first molecular structure type is a three-dimensional type, the corresponding first molecular structure information is a three-dimensional molecular conformation; the first physical property analysis report includes a plurality of first component analysis records; the first component analysis record includes the first formula component information and a corresponding first physical property analysis sequence.
6. An electronic device, comprising: Comprise: a memory, a processor and a transceiver; the processor is used to couple with the memory, read and execute instructions in the memory to realize the method of any one of claims 1-4; the transceiver is coupled with the processor, and the transceiver is controlled by the processor to perform message transmission and reception.
7. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, when the computer instructions are executed by the computer, the computer executes the method of any one of claims 1-4.
Citation Information
Patent Citations
Electrolyte design method, device, equipment, medium and program product
CN114255826A
Processing system for predicting electrochemical window of lithium battery electrolyte
CN117854625A