A process and apparatus for predicting the redox properties of electrolyte formulation components
By constructing a molecular redox property prediction model using the Uni-Mol model and a nonlinear regression model, the problems of high complexity and long cycle in the analysis of redox properties of electrolyte formulation components are solved, and efficient property prediction is achieved.
Patent Information
- Application Number
- CN202411723260.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-11-28
AI Technical Summary
Existing technologies for predicting the redox properties of electrolyte formulation components are complex, time-consuming, and inefficient.
A molecular redox property prediction model based on the Uni-Mol model and four nonlinear regression models was constructed. Through model training, four sets of redox properties of one-, two-, and three-dimensional molecular structures were predicted, and the redox property prediction sequence was output.
It reduces the complexity of analysis, shortens the analysis cycle, and improves analysis efficiency.
Smart Images

Figure CN119649952B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a processing method and device for predicting oxidation-reduction properties of electrolyte formula components. BACKGROUND
[0002] An electrolyte formula is composed of three components, namely a solvent, an electrolyte and an additive, each of which corresponds to a molecular structure. When designing and developing an electrolyte formula, the oxidation-reduction properties (such as HOMO / LUMO, VIP / VEA, AIP / AEA, oxidation / reduction potential, etc.) of each component of the formula need to be analyzed. At present, the conventional analysis method generally includes the following steps: first, three-dimensional modeling is performed on the basis of the molecular information (such as one-dimensional molecular sequence and two-dimensional molecular topological structure) of each formula component by using a chemical informatics tool (such as OpenBabel, RDKit, etc.), then molecular motion simulation is performed on the basis of the modeling structure by using a molecular dynamics (MD) simulation tool (such as LAMMPS, GROMACS, NAMD, etc.) to obtain the corresponding simulation structure, and then property calculation is performed on the basis of the simulation structure by using a quantum chemistry calculation tool (such as Gaussian, GAMESS, etc.). In practice, we found that the conventional analysis method has defects such as complex analysis operation, long analysis period and low analysis efficiency. SUMMARY
[0003] The present application is directed to the defects of the prior art, and provides a processing method and device for predicting oxidation-reduction properties of electrolyte formula components, an electronic device and a computer readable storage medium. The present application first constructs a molecular oxidation-reduction property prediction model on the basis of a Uni-Mol model and four nonlinear regression models, then trains the model to enable it to predict and output the corresponding oxidation-reduction property prediction sequence for four groups of oxidation-reduction properties (HOMO / LUMO, VIP / VEA, AIP / AEA, oxidation / reduction potential) of one / two / three-dimensional molecular structure, and finally uses the molecular oxidation-reduction property prediction model to analyze the four groups of oxidation-reduction properties of all formula components of any electrolyte formula input by the user after the model training is completed and feeds back the obtained property analysis report to the user. By analyzing the oxidation-reduction properties of the formula components of the electrolyte formula by the present application, the analysis complexity can be reduced, the analysis period can be shortened, and the analysis efficiency can be improved.
[0004] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application provides a processing method for predicting oxidation-reduction properties of electrolyte formula components, which comprises:
[0005] construct a molecular redox property prediction model; the molecular redox property prediction model is used for four-group redox property prediction and output of corresponding redox property prediction sequence according to input formula ingredient information; wherein, the formula ingredient information includes molecular structure type and molecular structure information; the molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the molecular structure type is one-dimensional type, and the corresponding molecular structure information is a one-dimensional molecular sequence; the molecular structure type is two-dimensional type, and the corresponding molecular structure information is a two-dimensional molecular topology graph; the molecular structure type is three-dimensional type, and the corresponding molecular structure information is a three-dimensional molecular conformation; the redox property prediction sequence includes HOMO prediction value, LUMO prediction value, VIP prediction value, VEA prediction value, AIP prediction value, AEA prediction value, oxidation potential prediction value and reduction potential prediction value;
[0006] model training is carried out on the molecular redox property prediction model based on a preset first data set; wherein, the first data set includes a plurality of first data records; the first data record includes first training molecular structure type, first training molecular structure information and first property label sequence; the first training molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the first training molecular structure type is one-dimensional type, and the corresponding first training molecular structure information is a one-dimensional molecular sequence; the first training molecular structure type is two-dimensional type, and the corresponding first training molecular structure information is a two-dimensional molecular topology graph; the first training molecular structure type is three-dimensional type, and the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first property label sequence includes first HOMO label value, first LUMO label value, first VIP label value, first VEA label value, first AIP label value, first AEA label value, first oxidation potential label value and first reduction potential label value;
[0007] After the model training is completed, an arbitrary electrolyte formula input by a user is recorded as a corresponding first formula; and a redox property prediction model of a molecule is used to analyze the redox property of a formula component of the first formula to obtain a corresponding first property analysis report to feed back to the user; wherein the first formula comprises a plurality of first formula component information; the first formula component information comprises a first molecular structure type and first molecular structure information; the first molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the first molecular structure type is a one-dimensional type, the corresponding first molecular structure information is a one-dimensional molecular sequence, when the first molecular structure type is a two-dimensional type, the corresponding first molecular structure information is a two-dimensional molecular topology graph, and when the first molecular structure type is a three-dimensional type, the corresponding first molecular structure information is a three-dimensional molecular conformation; the first property analysis report comprises a plurality of first component analysis records; the first component analysis record comprises the first formula component information and a corresponding first redox property analysis sequence.
[0008] Preferably, the model input end of the molecular redox property prediction model is used to receive the input formula component information, and the model output end is used to output the corresponding redox property prediction sequence;
[0009] The molecular redox property prediction model comprises a molecular structure initialization module, a molecular structure optimization module, a HOMO / LUMO prediction module, a VIP / VEA prediction module, an AIP / AEA prediction module, an oxidation / reduction potential prediction module and a molecular property output module;
[0010] The input end of the molecular structure initialization module is connected with the model input end, and the output end is connected with the input end of the molecular structure optimization module; the output end of the molecular structure optimization module is connected with the input end of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module respectively; the four input ends of the molecular property output module are connected with the output ends of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module respectively, and the output end of the molecular property output module is connected with the model output end;
[0011] The molecular structure initialization module is configured to extract the corresponding molecular structure type and molecular structure information from the formula ingredient information as the corresponding current type and current structure information, identify the current type, if the current type is a one-dimensional type, take the current structure information as a corresponding first molecular sequence, and use a one-dimensional sequence processing interface provided by a preset chemoinformatics tool to perform three-dimensional molecular structure modeling processing according to the first molecular sequence to obtain a corresponding initialized molecular conformation, if the current type is a two-dimensional type, take the current structure information as a corresponding first molecular topology graph, and use a two-dimensional topology graph processing interface provided by the chemoinformatics tool to perform three-dimensional molecular structure modeling processing according to the first molecular topology graph to obtain the corresponding initialized molecular conformation, if the current type is a three-dimensional type, take the current structure information as the corresponding initialized molecular conformation, and send the obtained initialized molecular conformation to the molecular structure optimization module, wherein the chemoinformatics tool at least includes OpenBabel and RDKit, the one-dimensional sequence processing interface is a processing interface provided by the chemoinformatics tool, and the one-dimensional sequence processing interface is configured to perform corresponding three-dimensional molecular conformation conversion according to one-dimensional molecular sequence input by an interface and output the obtained three-dimensional molecular conformation as interface output data, and the two-dimensional topology graph processing interface is a processing interface provided by the chemoinformatics tool, and the two-dimensional topology graph processing interface is configured to perform corresponding three-dimensional molecular conformation conversion according to a two-dimensional molecular topology graph input by an interface and output the obtained three-dimensional molecular conformation as interface output data;
[0012] The molecular structure optimization module is implemented based on a Uni-Mol model that has completed model pre-training, and is configured to perform three-dimensional structure optimization processing on the initialized molecular conformation to obtain a corresponding optimized molecular conformation, and send the optimized molecular conformation to the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module, and the oxidation / reduction potential prediction module, respectively;
[0013] The HOMO / LUMO prediction module is implemented based on a first nonlinear regression model, and is configured to perform prediction processing on the highest occupied molecular orbital and the lowest unoccupied molecular orbital of the optimized molecular conformation to obtain a corresponding HOMO prediction value and a LUMO prediction value, and send a corresponding HOMO / LUMO prediction vector composed of the HOMO prediction value and the LUMO prediction value to the molecular property output module, wherein the first nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model;
[0014] The VIP / VEA prediction module is implemented based on a second nonlinear regression model; the VIP / VEA prediction module is configured to predict the vertical ionization energy and the vertical electron affinity of the optimized molecular conformation to obtain corresponding VIP prediction values and VEA prediction values, and send a corresponding VIP / VEA prediction vector to the molecular property output module; the second nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model;
[0015] The AIP / AEA prediction module is implemented based on a third nonlinear regression model; the AIP / AEA prediction module is configured to predict the adiabatic ionization energy and the adiabatic electron affinity of the optimized molecular conformation to obtain corresponding AIP prediction values and AEA prediction values, and send a corresponding AIP / AEA prediction vector to the molecular property output module; the third nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model;
[0016] The oxidation / reduction potential prediction module is implemented based on a fourth nonlinear regression model; the oxidation / reduction potential prediction module is configured to predict the oxidation potential and the reduction potential of the optimized molecular conformation to obtain corresponding oxidation potential prediction values and reduction potential prediction values, and send a corresponding oxidation / reduction potential prediction vector to the molecular property output module; the fourth nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model;
[0017] The molecular property output module is configured to extract the HOMO prediction value, the LUMO prediction value, the VIP prediction value, the VEA prediction value, the AIP prediction value, the AEA prediction value, the oxidation potential prediction value, and the reduction potential prediction value from the received HOMO / LUMO prediction vector, VIP / VEA prediction vector, AIP / AEA prediction vector, and oxidation / reduction potential prediction vector to obtain a corresponding redox property prediction sequence and output.
[0018] Preferably, the molecular redox property prediction model is trained based on a preset first data set, specifically including:
[0019] Step 31, the first data set is randomly divided based on a preset first test evaluation ratio to obtain a corresponding first test set and a first evaluation set;
[0020] The first test set and the first evaluation set are both composed of a plurality of first data records; a ratio of a total number of records of the first test set to a total number of records of the first evaluation set meets the first test evaluation ratio;
[0021] Step 32, setting a corresponding first debugging object as a nonlinear regression model; and extracting a first first data record of the first test set as a corresponding current test record;
[0022] The first debugging object includes a nonlinear regression model and a structure optimization model.
[0023] Step 33, a corresponding first training formula component information is composed of the first training molecular structure type and the first training molecular structure information of the current test record;
[0024] Step 34, the first training formula component information is input into the molecular redox property prediction model to perform four-group redox property prediction and obtain a corresponding first redox property prediction sequence;
[0025] The first redox property prediction sequence includes a first HOMO prediction value, a first LUMO prediction value, a first VIP prediction value, a first VEA prediction value, a first AIP prediction value, a first AEA prediction value, a first oxidation potential prediction value and a first reduction potential prediction value.
[0026] Step 35, the first redox property prediction sequence and the first property label sequence of the current test record are brought into a preset first model loss function to obtain a corresponding first loss value through calculation;
[0027] The first model loss function at least includes an L1 loss function, an L2 loss function and a cross-entropy loss function.
[0028] Step 36, identify whether the first loss value meets the preset first loss value range; if the first loss value meets the first loss value range, identify whether the current test record is the last first data record of the first test set, if yes, go to step 37, if no, extract the next first data record of the first test set as a new current test record and return to step 33; if the first loss value does not meet the first loss value range, identify the first debugging object, if the first debugging object is a nonlinear regression model, based on a preset first model parameter optimizer, on the premise that the model parameters of the molecular structure optimization module remain unchanged, one round of model parameter modulation is performed on the model parameters of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module towards the direction of making the first model loss function reach the minimum value, and at the end of this round of parameter modulation, return to step 34, if the first debugging object is a structure optimization model, based on a preset second model parameter optimizer, on the premise that the model parameters of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module remain unchanged, one round of model parameter fine tuning is performed on the model parameters of the molecular structure optimization module towards the direction of making the first model loss function reach the minimum value, and at the end of this round of parameter fine tuning, return to step 34;
[0029] Wherein, the first model parameter optimizer at least includes an SGD optimizer, an ADAM optimizer; the second model parameter optimizer at least includes an RMSprop optimizer, an AdamW optimizer and an ADAM optimizer;
[0030] Step 37, identify whether the first debugging object is a structure optimization model; if yes, go to step 38; if no, set the first debugging object as a structure optimization model, extract the first first data record of the first test set as a new current test record, and return to step 33;
[0031] Step 38, traversing all the first data records of the first evaluation set; and in the process of traversal, taking the first data record currently traversed as a corresponding current evaluation record; and taking the first training molecular structure type and the first training molecular structure information of the current evaluation record to form a corresponding second training formula component information; and inputting the second training formula component information into the molecular redox property prediction model to perform four-group redox property prediction and obtain a corresponding second redox property prediction sequence; and taking the second redox property prediction sequence and the first property label sequence of the current test record to form a corresponding prediction-label pair; and at the end of traversal, taking all the prediction-label pairs obtained into a preset first model evaluation function to obtain a corresponding first evaluation value;
[0032] The first model evaluation function at least includes an RMSE function.
[0033] Step 39, identifying whether the first evaluation value meets a preset first evaluation value range; if not, returning to step 31 to continue training; if yes, stopping training and confirming that the model training is completed.
[0034] Preferably, the first property analysis report obtained by analyzing the redox properties of the formula components of the first formula using the molecular redox property prediction model is fed back to the user, specifically including:
[0035] Inputting each first formula component information of the first formula into the molecular redox property prediction model to perform four-group redox property prediction and taking the redox property prediction sequence obtained this time as a corresponding first redox property analysis sequence; and taking each first formula component information and the corresponding first redox property analysis sequence to form a corresponding first component analysis record; and taking all the first component analysis records obtained to form a corresponding first property analysis report to feed back to the user.
[0036] The second aspect of the embodiment of the application provides a device for implementing the processing method for predicting the redox properties of formula components of electrolyte, which is used to implement the processing method for predicting the redox properties of formula components of electrolyte as described in the first aspect.
[0037] The model construction module is configured to construct a molecular redox property prediction model; the molecular redox property prediction model is configured to perform four-group redox property prediction according to inputted formula ingredient information and output a corresponding redox property prediction sequence; wherein the formula ingredient information comprises a molecular structure type and molecular structure information; the molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the molecular structure type is a one-dimensional type, the corresponding molecular structure information is a one-dimensional molecular sequence; when the molecular structure type is a two-dimensional type, the corresponding molecular structure information is a two-dimensional molecular topology graph; and when the molecular structure type is a three-dimensional type, the corresponding molecular structure information is a three-dimensional molecular conformation; the redox property prediction sequence comprises a HOMO prediction value, a LUMO prediction value, a VIP prediction value, a VEA prediction value, an AIP prediction value, an AEA prediction value, an oxidation potential prediction value and a reduction potential prediction value;
[0038] The model training module is configured to perform model training on the molecular redox property prediction model based on a preset first data set; wherein the first data set comprises a plurality of first data records; the first data record comprises a first training molecular structure type, first training molecular structure information and a first property label sequence; the first training molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the first training molecular structure type is a one-dimensional type, the corresponding first training molecular structure information is a one-dimensional molecular sequence; when the first training molecular structure type is a two-dimensional type, the corresponding first training molecular structure information is a two-dimensional molecular topology graph; and when the first training molecular structure type is a three-dimensional type, the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first property label sequence comprises a first HOMO label value, a first LUMO label value, a first VIP label value, a first VEA label value, a first AIP label value, a first AEA label value, a first oxidation potential label value and a first reduction potential label value;
[0039] The model application module is used to receive a user inputted arbitrary electrolyte formula as a corresponding first formula after the end of model training, and analyze the redox properties of the formula components of the first formula using the molecular redox property prediction model to obtain a corresponding first property analysis report to feed back to the user; wherein the first formula includes a plurality of first formula component information; the first formula component information includes a first molecular structure type and first molecular structure information; the first molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; when the first molecular structure type is a one-dimensional type, the corresponding first molecular structure information is a one-dimensional molecular sequence, when the first molecular structure type is a two-dimensional type, the corresponding first molecular structure information is a two-dimensional molecular topology graph, and when the first molecular structure type is a three-dimensional type, the corresponding first molecular structure information is a three-dimensional molecular conformation; the first property analysis report includes a plurality of first component analysis records; the first component analysis record includes the first formula component information and a corresponding first redox property analysis sequence.
[0040] The third aspect of the embodiment of the present application provides an electronic device, comprising a memory, a processor and a transceiver;
[0041] The processor is used to be coupled with the memory, read and execute instructions in the memory, so as to realize the method steps of the first aspect;
[0042] The transceiver is coupled with the processor, and the transceiver is controlled by the processor to perform message transceiving.
[0043] The fourth aspect of the embodiment of the present application provides a computer readable storage medium, the computer readable storage medium stores computer instructions, when the computer instructions are executed by the computer, the computer instructions make the computer execute the instructions of the method of the first aspect.
[0044] The embodiment of the present application provides a processing method, device, electronic equipment and computer readable storage medium for predicting oxidation-reduction properties of electrolyte formula components. According to the above content, the embodiment of the present application first constructs a molecular oxidation-reduction property prediction model based on a Uni-Mol model and four nonlinear regression models, then makes the model trained to be capable of predicting four groups of oxidation-reduction properties (HOMO / LUMO, VIP / VEA, AIP / AEA, oxidation / reduction potential) of one / two / three-dimensional molecular structure and outputting corresponding oxidation-reduction property prediction sequences, and finally analyzes four groups of oxidation-reduction properties of all formula components of an arbitrary electrolyte formula input by a user after the model training is completed and feeds back a property analysis report obtained to the user. Through the oxidation-reduction property analysis of the formula components of the electrolyte formula by the embodiment of the present application, the analysis complexity is reduced, and the analysis cycle is shortened and the analysis efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 A processing method for predicting oxidation-reduction properties of electrolyte formula components provided by the embodiment one of the present application is shown in a schematic diagram.
[0046] Figure 2 A module schematic diagram of a molecular oxidation-reduction property prediction model provided by the embodiment one of the present application is shown.
[0047] Figure 3 A module structure diagram of a processing device for predicting oxidation-reduction properties of electrolyte formula components provided by the embodiment two of the present application is shown.
[0048] Figure 4 A structure schematic diagram of an electronic equipment provided by the embodiment three of the present application is shown. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.
[0050] The embodiment one of the present application provides a processing method for predicting oxidation-reduction properties of electrolyte formula components, which comprises the following steps. Figure 1 A processing method for predicting oxidation-reduction properties of electrolyte formula components provided by the embodiment one of the present application is shown in a schematic diagram, and the method mainly comprises the following steps.
[0051] Step 1, constructing a molecular oxidation-reduction property prediction model.
[0052] The molecular redox property prediction model of the embodiment of the present application is used to perform four-group redox property prediction according to the input formula component information and output the corresponding redox property prediction sequence; wherein the formula component information includes molecular structure type and molecular structure information; the molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the molecular structure information corresponding to the one-dimensional type is a one-dimensional molecular sequence, the molecular structure information corresponding to the two-dimensional type is a two-dimensional molecular topology graph, and the molecular structure information corresponding to the three-dimensional type is a three-dimensional molecular conformation; the redox property prediction sequence includes the highest occupied molecular orbital (HOMO) prediction value, the lowest unoccupied molecular orbital (LUMO) prediction value, the vertical ionization potential (VIP) prediction value, the vertical electron affinity (VEA) prediction value, the adiabatic ionization potential (AIP) prediction value, the adiabatic electron affinity (AEA) prediction value, the oxidation potential (OP) prediction value and the reduction potential (RP) prediction value.
[0053] As Figure 2 As shown in the module schematic diagram of the molecular redox property prediction model provided by the first embodiment of the present application, the model input end of the molecular redox property prediction model of the embodiment of the present application is used to receive the input formula component information, and the model output end is used to output the corresponding redox property prediction sequence; and the molecular redox property prediction model includes a molecular structure initialization module, a molecular structure optimization module, a HOMO / LUMO prediction module, a VIP / VEA prediction module, an AIP / AEA prediction module, an oxidation / reduction potential prediction module and a molecular property output module;
[0054] The connection relationship of each module of the molecular oxidation-reduction property prediction model is as follows: the input end of the molecular structure initialization module is connected with the model input end, and the output end is connected with the input end of the molecular structure optimization module; the output end of the molecular structure optimization module is connected with the input end of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module respectively; the four input ends of the molecular property output module are connected with the output ends of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module respectively, and the output end of the molecular property output module is connected with the model output end;
[0055] The molecular structure initialization module of the molecular oxidation-reduction property prediction model is used for extracting the corresponding molecular structure type and molecular structure information from the formula component information as the corresponding current type and current structure information, and identifying the current type; if the current type is one-dimensional type, the current structure information is taken as the corresponding first molecular sequence, and a one-dimensional sequence processing interface provided by a preset chemoinformatics tool is used to perform three-dimensional molecular structure modeling processing according to the first molecular sequence to obtain the corresponding initialization molecular conformation; if the current type is two-dimensional type, the current structure information is taken as the corresponding first molecular topological graph, and a two-dimensional topological graph processing interface provided by the chemoinformatics tool is used to perform three-dimensional molecular structure modeling processing according to the first molecular topological graph to obtain the corresponding initialization molecular conformation; if the current type is three-dimensional type, the current structure information is taken as the corresponding initialization molecular conformation; and the obtained initialization molecular conformation is sent to the molecular structure optimization module; wherein the chemoinformatics tool used in the embodiment of the application at least includes OpenBabel and RDKit; the one-dimensional sequence processing interface is a processing interface provided by the above chemoinformatics tool, which is used for converting the one-dimensional molecular sequence input by the interface into the corresponding three-dimensional molecular conformation and outputting the obtained three-dimensional molecular conformation as interface output data; the two-dimensional topological graph processing interface is a processing interface provided by the above chemoinformatics tool, which is used for converting the two-dimensional molecular topological graph input by the interface into the corresponding three-dimensional molecular conformation and outputting the obtained three-dimensional molecular conformation as interface output data;
[0056] The molecular structure optimization module of the molecular oxidation-reduction property prediction model of the embodiment of the application is implemented based on a Uni-Mol model that has completed model pre-training; the molecular structure optimization module is used for performing three-dimensional structure optimization processing on the initialized molecular conformation to obtain a corresponding optimized molecular conformation; and the optimized molecular conformation is sent to the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module respectively; the Uni-Mol model can be understood with reference to the public technical document UNI-MOL:A UNIVERSAL 3D MOLECULAR REPRESENTATION LEARNING FRAMEWORK, which describes the pre-training method of the model; the pre-training of the Uni-Mol model can be completed by referring to the content of the reference paper; the technical document also points out that the Uni-Mol model can perform structure optimization on the three-dimensional molecular conformation; therefore, the module function of the molecular structure optimization module is implemented based on a Uni-Mol model that has completed model pre-training in the embodiment of the application;
[0057] The HOMO / LUMO prediction module of the molecular oxidation-reduction property prediction model of the embodiment of the application is implemented based on a first nonlinear regression model; the HOMO / LUMO prediction module is used for performing prediction processing on the highest occupied molecular orbital and the lowest unoccupied molecular orbital of the optimized molecular conformation to obtain a corresponding HOMO prediction value and a LUMO prediction value to form a corresponding HOMO / LUMO prediction vector, which is sent to the molecular property output module; the optional model types of the first nonlinear regression model of the embodiment of the application include at least an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0058] The VIP / VEA prediction module of the molecular oxidation-reduction property prediction model of the embodiment of the application is implemented based on a second nonlinear regression model; the VIP / VEA prediction module is used for performing prediction processing on the vertical ionization energy and the vertical electron affinity of the optimized molecular conformation to obtain a corresponding VIP prediction value and a VEA prediction value to form a corresponding VIP / VEA prediction vector, which is sent to the molecular property output module; the optional model types of the second nonlinear regression model of the embodiment of the application include at least an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0059] The AIP / AEA prediction module of the molecular oxidation-reduction property prediction model is realized based on a third nonlinear regression model; the AIP / AEA prediction module is used for performing prediction processing on the adiabatic ionization energy and the adiabatic electron affinity of the optimized molecular conformation to obtain corresponding AIP prediction values and AEA prediction values to form a corresponding AIP / AEA prediction vector and send the AIP / AEA prediction vector to the molecular property output module; the optional model types of the third nonlinear regression model of the embodiment of the application include at least an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0060] The oxidation / reduction potential prediction module of the molecular oxidation-reduction property prediction model is realized based on a fourth nonlinear regression model; the oxidation / reduction potential prediction module is used for performing prediction processing on the oxidation potential and the reduction potential of the optimized molecular conformation to obtain corresponding oxidation potential prediction values and reduction potential prediction values to form a corresponding oxidation / reduction potential prediction vector and send the oxidation / reduction potential prediction vector to the molecular property output module; the optional model types of the fourth nonlinear regression model of the embodiment of the application include at least an XGBoost model, a GBDT model, a Random Forest model and an MLP model;
[0061] The molecular property output module of the molecular oxidation-reduction property prediction model is used for extracting HOMO prediction values, LUMO prediction values, VIP prediction values, VEA prediction values, AIP prediction values, AEA prediction values, oxidation potential prediction values and reduction potential prediction values from the received HOMO / LUMO prediction vector, the VIP / VEA prediction vector, the AIP / AEA prediction vector and the oxidation / reduction potential prediction vector to form a corresponding oxidation-reduction property prediction sequence and output the oxidation-reduction property prediction sequence.
[0062] Step 2, model training is performed on the molecular oxidation-reduction property prediction model based on a preset first data set;
[0063] Specifically, step 21, the first data set is randomly divided to obtain a corresponding first test set and a first evaluation set based on a preset first test evaluation ratio;
[0064] The first data set of the embodiment of the application is a model training data set constructed in advance by big data collection, which is composed of a plurality of first data records; the first data record includes a first training molecular structure type, first training molecular structure information, and a first property label sequence; the first training molecular structure type includes one-dimensional, two-dimensional, and three-dimensional types; when the first training molecular structure type is one-dimensional, the corresponding first training molecular structure information is a one-dimensional molecular sequence; when the first training molecular structure type is two-dimensional, the corresponding first training molecular structure information is a two-dimensional molecular topology graph; and when the first training molecular structure type is three-dimensional, the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first property label sequence includes a first HOMO label value, a first LUMO label value, a first VIP label value, a first VEA label value, a first AIP label value, a first AEA label value, a first oxidation potential label value, and a first reduction potential label value;
[0065] The first test evaluation ratio of the embodiment of the application is a pre-set ratio value, for example, 8:2;
[0066] The first test set and the first evaluation set of the embodiment of the application are both composed of a plurality of first data records; and the ratio of the total number of records of the first test set to the total number of records of the first evaluation set meets the first test evaluation ratio;
[0067] Step 22, setting the corresponding first debugging object as a nonlinear regression model; and extracting the first data record of the first test set as the corresponding current test record;
[0068] The first debugging object includes a nonlinear regression model and a structure optimization model.
[0069] Step 23, composing a corresponding first training formula component information from the first training molecular structure type and the first training molecular structure information of the current test record;
[0070] Step 24, inputting the first training formula component information into the molecular oxidation-reduction property prediction model to perform four-group oxidation-reduction property prediction and obtain a corresponding first oxidation-reduction property prediction sequence;
[0071] The first oxidation-reduction property prediction sequence includes a first HOMO prediction value, a first LUMO prediction value, a first VIP prediction value, a first VEA prediction value, a first AIP prediction value, a first AEA prediction value, a first oxidation potential prediction value, and a first reduction potential prediction value.
[0072] Step 25, bringing the first oxidation-reduction property prediction sequence and the first property label sequence of the current test record into a pre-set first model loss function to calculate a corresponding first loss value;
[0073] wherein the first model loss function comprises at least an L1 loss function, an L2 loss function and a cross-entropy loss function;
[0074] Step 26, whether the first loss value satisfies a preset first loss value range is identified; if the first loss value satisfies the first loss value range, whether the current test record is the last first data record of the first test set is identified, if yes, step 27 is transferred to, if no, the next first data record of the first test set is extracted as a new current test record and step 23 is returned; if the first loss value does not satisfy the first loss value range, the first debugging object is identified, if the first debugging object is a nonlinear regression model, the model parameters of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module are modulated in one round based on a preset first model parameter optimizer towards the direction of making the first model loss function reach a minimum value on the premise that the model parameters of the molecular structure optimization module remain unchanged and step 24 is returned at the end of the round of parameter modulation, if the first debugging object is a structure optimization model, the model parameters of the molecular structure optimization module are fine-tuned in one round based on a preset second model parameter optimizer towards the direction of making the first model loss function reach a minimum value on the premise that the model parameters of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module remain unchanged and step 24 is returned at the end of the round of parameter fine-tuning;
[0075] wherein the first loss value range is a preset loss value range; the first model parameter optimizer comprises at least an SGD optimizer, an ADAM optimizer; the second model parameter optimizer comprises at least an RMSprop optimizer, an AdamW optimizer and an ADAM optimizer;
[0076] Step 27, whether the first debugging object is a structure optimization model is identified; if yes, step 28 is transferred to; if no, the first debugging object is set as a structure optimization model, the first data record of the first test set is extracted as a new current test record and step 23 is returned;
[0077] Here, as can be seen from the current step, the embodiment of the application will adopt a two-stage training method to reduce the training difficulty when training the molecular redox property prediction model; 1) In the first stage, set the first debugging object = nonlinear regression model. During this stage of training, only the parameters of the four nonlinear regression models (HOMO / LUMO prediction module, VIP / VEA prediction module, AIP / AEA prediction module, and oxidation / reduction potential prediction module) are modulated, and the model parameters of the Uni-Mol model (molecular structure optimization module) are not processed; 2) In the second stage, set the first debugging object = structure optimization model. During this stage of training, the parameters of the four nonlinear regression models (HOMO / LUMO prediction module, VIP / VEA prediction module, AIP / AEA prediction module, and oxidation / reduction potential prediction module) are not modulated, and only the model parameters of the Uni-Mol model (molecular structure optimization module) are fine-tuned.
[0078] Step 28, traverse all first data records of the first evaluation set; and in the traversal process, take the currently traversed first data record as a corresponding current evaluation record; and form a corresponding second training formula component information from the first training molecular structure type and the first training molecular structure information of the current evaluation record; and input the second training formula component information into the molecular redox property prediction model to perform four-group redox property prediction and obtain a corresponding second redox property prediction sequence; and form a corresponding prediction-label pair from the second redox property prediction sequence and the first property label sequence of the current test record; and at the end of the traversal, bring all obtained prediction-label pairs into the preset first model evaluation function to obtain a corresponding first evaluation value;
[0079] Among them, the first model evaluation function at least includes the RMSE function;
[0080] Step 29, identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 21 to continue training; if yes, stop training and confirm that the model training is completed; wherein the first evaluation value range is a pre-set evaluation value range.
[0081] Step 3, after the model training is completed, receive the user input of any electrolyte formula as a corresponding first formula; and use the molecular redox property prediction model to analyze the oxidation and reduction properties of the formula components of the first formula to obtain a corresponding first property analysis report to feedback to the user;
[0082] Specifically includes: step 31, after the model training is completed, receive the user input of any electrolyte formula as a corresponding first formula;
[0083] The first formula includes a plurality of first formula component information; the first formula component information includes a first molecular structure type and first molecular structure information; the first molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the first molecular structure type corresponds to a one-dimensional molecular sequence when being a one-dimensional type, a two-dimensional molecular topology map when being a two-dimensional type, and a three-dimensional molecular conformation when being a three-dimensional type;
[0084] In step 32, the redox properties of the formula components of the first formula are analyzed using the molecular redox property prediction model to obtain a corresponding first property analysis report, which is fed back to the user;
[0085] The first property analysis report includes a plurality of first component analysis records; the first component analysis record includes first formula component information and a corresponding first redox property analysis sequence;
[0086] Specifically, each first formula component information of the first formula is input into the molecular redox property prediction model to perform four-group redox property prediction, and the obtained redox property prediction sequence is taken as the corresponding first redox property analysis sequence; each first formula component information and the corresponding first redox property analysis sequence form a corresponding first component analysis record; and all obtained first component analysis records form a corresponding first property analysis report, which is fed back to the user.
[0087] Figure 3 A module structure diagram of a processing device for predicting the redox properties of electrolyte formula components is provided for the second embodiment of the present application. The device is a terminal device or a server for implementing the foregoing method embodiments, or a device that can enable the foregoing terminal device or server to implement the foregoing method embodiments, such as a device or chip system of the foregoing terminal device or server. As shown in the figure, the device includes a model construction module 201, a model training module 202 and a model application module 203. Figure 3
[0088] The model construction module 201 is configured to construct a molecular redox property prediction model; the molecular redox property prediction model is configured to perform four-group redox property prediction according to inputted formula ingredient information and output a corresponding redox property prediction sequence; wherein the formula ingredient information comprises a molecular structure type and molecular structure information; the molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the molecular structure type is one-dimensional, the corresponding molecular structure information is a one-dimensional molecular sequence; when the molecular structure type is two-dimensional, the corresponding molecular structure information is a two-dimensional molecular topology graph; when the molecular structure type is three-dimensional, the corresponding molecular structure information is a three-dimensional molecular conformation; the redox property prediction sequence comprises a HOMO prediction value, a LUMO prediction value, a VIP prediction value, a VEA prediction value, an AIP prediction value, an AEA prediction value, an oxidation potential prediction value and a reduction potential prediction value.
[0089] The model training module 202 is configured to perform model training on the molecular redox property prediction model based on a preset first data set; wherein the first data set comprises a plurality of first data records; the first data record comprises a first training molecular structure type, first training molecular structure information and a first property label sequence; the first training molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the first training molecular structure type is one-dimensional, the corresponding first training molecular structure information is a one-dimensional molecular sequence; when the first training molecular structure type is two-dimensional, the corresponding first training molecular structure information is a two-dimensional molecular topology graph; when the first training molecular structure type is three-dimensional, the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first property label sequence comprises a first HOMO label value, a first LUMO label value, a first VIP label value, a first VEA label value, a first AIP label value, a first AEA label value, a first oxidation potential label value and a first reduction potential label value.
[0090] The model application module 203 is configured to receive a user inputted arbitrary electrolyte formula as a corresponding first formula after the model training is completed, and analyze the redox properties of the formula components of the first formula by using the molecular redox property prediction model to obtain a corresponding first property analysis report and feed back to the user; wherein the first formula comprises a plurality of first formula component information; the first formula component information comprises a first molecular structure type and first molecular structure information; the first molecular structure type comprises one-dimensional, two-dimensional and three-dimensional types; when the first molecular structure type is one-dimensional type, the corresponding first molecular structure information is a one-dimensional molecular sequence, when the first molecular structure type is two-dimensional type, the corresponding first molecular structure information is a two-dimensional molecular topological graph, and when the first molecular structure type is three-dimensional type, the corresponding first molecular structure information is a three-dimensional molecular conformation; the first property analysis report comprises a plurality of first component analysis records; the first component analysis record comprises the first formula component information and a corresponding first redox property analysis sequence.
[0091] The processing device for predicting the redox properties of electrolyte formula components provided by the embodiment of the present application can execute the method steps in the method embodiments described above, and has similar implementation principles and technical effects, which will not be described here again.
[0092] It should be noted that the division of each module of the above device is only a logical division of functions, and all or part of the modules can be integrated into one physical entity, or physically separated. And these modules can all be implemented in the form of software called by the processing element; all can be implemented in the form of hardware; some modules can be implemented in the form of software called by the processing element, and some modules can be implemented in the form of hardware. For example, the model construction module can be a separate processing element, or can be integrated into a chip of the above device, in addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above determination module is called and executed by a processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or can be independently implemented. The processing element described herein can be an integrated circuit with signal processing capability. In the implementation process, each step of the above method or each module can be completed by the integrated logic circuit of hardware or the instruction of software in the processing element.
[0093] For example, the above modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), or one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs), etc. For another example, when a certain module above is implemented in the form of a processing element scheduling code, the processing element can be a general purpose processor, such as a Central Processing Unit (CPU) or other processor that can invoke code. For another example, the modules can be integrated together to implement in the form of a System-on-a-chip (SOC).
[0094] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When loaded and executed by a computer, the computer instructions generate all or part of the processes or functions described in the above method embodiments. The computer can be a general purpose computer, a special purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.) manner. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0095] Figure 4 A structural schematic diagram of an electronic device is provided for Embodiment Three of the present application. The electronic device can be a terminal device or a server implementing the method of the above embodiments, or a terminal device or a server connected to the terminal device or the server implementing the method of the above embodiments. As shown in FIG. 3, the electronic device includes a processor 301, a memory 302, a transceiver 303 and an antenna 304. The processor 301, the memory 302, the transceiver 303 and the antenna 304 can be connected to each other through a bus or other suitable connection means. The processor 301 can be configured to implement the method of the above embodiments. The memory 302 can be configured to store the computer instructions of the processor 301. The transceiver 303 can be configured to transmit and receive signals. The antenna 304 can be configured to transmit and receive signals. Figure 4As shown, the electronic device can include a processor 301 (such as a CPU), a memory 302, a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiving action of the transceiver 303. The memory 302 can store various instructions for completing various processing functions and implementing the processing steps described in the foregoing embodiment method description. Preferably, the electronic device related to the embodiments of the present application further includes a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize the communication connection between elements. The above-mentioned communication port 306 is used for connection communication between the electronic device and other peripherals.
[0096] In Figure 4 The system bus 305 mentioned in the above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to realize the communication between the database access device and other devices (such as the client, the read-write library and the read-only library). The memory can contain a Random Access Memory (RAM), and can also include a Non-Volatile Memory, such as at least one disk memory.
[0097] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0098] It should be noted that the embodiments of the present application also provide a computer readable storage medium, which stores instructions, when running on a computer, causes the computer to execute the method and processing procedure provided in the above embodiments.
[0099] The embodiment of the present application provides a processing method and device for predicting the oxidation-reduction property of electrolyte formula components, electronic equipment and computer readable storage medium. According to the above content, the embodiment of the present application first constructs a molecular oxidation-reduction property prediction model based on a Uni-Mol model and four nonlinear regression models, then makes the model trained to be capable of predicting four groups of oxidation-reduction properties (HOMO / LUMO, VIP / VEA, AIP / AEA, oxidation / reduction potential) of one / two / three-dimensional molecular structure and outputting the corresponding oxidation-reduction property prediction sequence, and finally uses the molecular oxidation-reduction property prediction model to analyze the four groups of oxidation-reduction properties of all formula components of any electrolyte formula input by a user after the model training is completed and feeds back the obtained property analysis report to the user. Through the oxidation-reduction property analysis of the formula components of the electrolyte formula by the embodiment of the present application, the analysis complexity is reduced, the analysis period is shortened, and the analysis efficiency is improved.
[0100] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0101] The above detailed description is further detailed for the purpose of the present application, technical solutions and beneficial effects, and it should be understood that the above detailed description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A process for predicting the redox properties of a component of an electrolyte formulation, characterized by, The method comprises: constructing a molecular redox property prediction model; the molecular redox property prediction model is used for four-group redox property prediction and output of corresponding redox property prediction sequence according to input formula component information; wherein, the formula component information includes molecular structure type and molecular structure information; the molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the molecular structure type is one-dimensional type, and the corresponding molecular structure information is a one-dimensional molecular sequence; the molecular structure type is two-dimensional type, and the corresponding molecular structure information is a two-dimensional molecular topological graph; the molecular structure type is three-dimensional type, and the corresponding molecular structure information is a three-dimensional molecular conformation; the redox property prediction sequence includes HOMO prediction value, LUMO prediction value, VIP prediction value, VEA prediction value, AIP prediction value, AEA prediction value, oxidation potential prediction value and reduction potential prediction value; model training is performed on the molecular redox property prediction model based on a preset first data set; wherein, the first data set includes a plurality of first data records; the first data record includes first training molecular structure type, first training molecular structure information and first property label sequence; the first training molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the first training molecular structure type is one-dimensional type, and the corresponding first training molecular structure information is a one-dimensional molecular sequence; the first training molecular structure type is two-dimensional type, and the corresponding first training molecular structure information is a two-dimensional molecular topological graph; the first training molecular structure type is three-dimensional type, and the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first property label sequence includes first HOMO label value, first LUMO label value, first VIP label value, first VEA label value, first AIP label value, first AEA label value, first oxidation potential label value and first reduction potential label value; after the model training is completed, any electrolyte formula input by a user is recorded as a corresponding first formula; and the molecular redox property prediction model is used to analyze the redox property of formula components of the first formula to obtain a corresponding first property analysis report for user feedback; wherein, the first formula includes a plurality of first formula component information; the first formula component information includes first molecular structure type and first molecular structure information; the first molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; the first molecular structure type is one-dimensional type, and the corresponding first molecular structure information is a one-dimensional molecular sequence; the first molecular structure type is two-dimensional type, and the corresponding first molecular structure information is a two-dimensional molecular topological graph; the first molecular structure type is three-dimensional type, and the corresponding first molecular structure information is a three-dimensional molecular conformation; the first property analysis report includes a plurality of first component analysis records; the first component analysis record includes the first formula component information and a corresponding first redox property analysis sequence.
2. The processing method of claim 1, wherein the molecular redox property prediction model is configured to receive the inputted formula component information at a model input end and output a corresponding redox property prediction sequence at a model output end. The molecular redox property prediction model comprises a molecular structure initialization module, a molecular structure optimization module, a HOMO / LUMO prediction module, a VIP / VEA prediction module, an AIP / AEA prediction module, an oxidation / reduction potential prediction module, and a molecular property output module. The input end of the molecular structure initialization module is connected to the model input end, and the output end is connected to the input end of the molecular structure optimization module. The output end of the molecular structure optimization module is connected to the input end of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module, and the oxidation / reduction potential prediction module. The four input ends of the molecular property output module are connected to the output end of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module, and the oxidation / reduction potential prediction module. The molecular structure initialization module is configured to extract the molecular structure type and the molecular structure information from the formula component information as current type and current structure information, and identify the current type. If the current type is a one-dimensional type, the current structure information is taken as a first molecular sequence, and a one-dimensional sequence processing interface provided by a preset chemoinformatics tool is used to perform three-dimensional molecular structure modeling processing according to the first molecular sequence to obtain an initial molecular conformation. If the current type is a two-dimensional type, the current structure information is taken as a first molecular topology graph, and a two-dimensional topology graph processing interface provided by the chemoinformatics tool is used to perform three-dimensional molecular structure modeling processing according to the first molecular topology graph to obtain the initial molecular conformation. If the current type is a three-dimensional type, the current structure information is taken as the initial molecular conformation. The obtained initial molecular conformation is sent to the molecular structure optimization module. The chemoinformatics tool at least comprises OpenBabel and RDKit. The one-dimensional sequence processing interface is a processing interface provided by a chemoinformatics tool, which is used to convert a one-dimensional molecular sequence inputted by an interface into a three-dimensional molecular conformation and output the three-dimensional molecular conformation as interface output data. The two-dimensional topology graph processing interface is a processing interface provided by a chemoinformatics tool, which is used to convert a two-dimensional molecular topology graph inputted by an interface into a three-dimensional molecular conformation and output the three-dimensional molecular conformation as interface output data. The molecular structure optimization module is implemented based on a Uni-Mol model that has completed model pre-training; the molecular structure optimization module is configured to perform three-dimensional structure optimization processing on the initialized molecular conformation to obtain a corresponding optimized molecular conformation; and the optimized molecular conformation is sent to the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module, and the oxidation / reduction potential prediction module, respectively; The HOMO / LUMO prediction module is implemented based on a first nonlinear regression model; the HOMO / LUMO prediction module is configured to perform prediction processing on the highest occupied molecular orbital and the lowest unoccupied molecular orbital of the optimized molecular conformation to obtain a corresponding HOMO prediction value and a corresponding LUMO prediction value, thereby forming a corresponding HOMO / LUMO prediction vector, which is sent to the molecular property output module; the first nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model; The VIP / VEA prediction module is implemented based on a second nonlinear regression model; the VIP / VEA prediction module is configured to perform prediction processing on the vertical ionization energy and the vertical electron affinity energy of the optimized molecular conformation to obtain a corresponding VIP prediction value and a corresponding VEA prediction value, thereby forming a corresponding VIP / VEA prediction vector, which is sent to the molecular property output module; the second nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model; The AIP / AEA prediction module is implemented based on a third nonlinear regression model; the AIP / AEA prediction module is configured to perform prediction processing on the adiabatic ionization energy and the adiabatic electron affinity energy of the optimized molecular conformation to obtain a corresponding AIP prediction value and a corresponding AEA prediction value, thereby forming a corresponding AIP / AEA prediction vector, which is sent to the molecular property output module; the third nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model; The oxidation / reduction potential prediction module is implemented based on a fourth nonlinear regression model; the oxidation / reduction potential prediction module is configured to perform prediction processing on the oxidation potential and the reduction potential of the optimized molecular conformation to obtain a corresponding oxidation potential prediction value and a corresponding reduction potential prediction value, thereby forming a corresponding oxidation / reduction potential prediction vector, which is sent to the molecular property output module; the fourth nonlinear regression model at least includes an XGBoost model, a GBDT model, a Random Forest model, and an MLP model; The molecular property output module is used to extract the HOMO prediction value, the LUMO prediction value, the VIP prediction value, the VEA prediction value, the AIP prediction value, the AEA prediction value, the oxidation potential prediction value and the reduction potential prediction value from the received HOMO / LUMO prediction vector, the VIP / VEA prediction vector, the AIP / AEA prediction vector and the oxidation / reduction potential prediction vector to form a corresponding redox property prediction sequence and output.
3. The method of claim 2, wherein the method is used to predict the redox properties of a component of an electrolyte formulation. The model training of the molecular redox property prediction model based on the preset first data set specifically includes: Step 31, randomly dividing the first data set based on a preset first test evaluation ratio to obtain a corresponding first test set and a first evaluation set; Wherein, the first test set and the first evaluation set are both composed of a plurality of first data records; the ratio of the total number of records of the first test set to the total number of records of the first evaluation set satisfies the first test evaluation ratio; Step 32, setting a corresponding first debugging object as a nonlinear regression model; and extracting a first first data record of the first test set as a corresponding current test record; Wherein, the first debugging object includes a nonlinear regression model and a structure optimization model; Step 33, a corresponding first training formula component information is composed of the first training molecular structure type of the current test record and the first training molecular structure information; Step 34, inputting the first training formula component information into the molecular redox property prediction model to perform four-group redox property prediction and obtaining a corresponding first redox property prediction sequence; Wherein, the first redox property prediction sequence includes a first HOMO prediction value, a first LUMO prediction value, a first VIP prediction value, a first VEA prediction value, a first AIP prediction value, a first AEA prediction value, a first oxidation potential prediction value and a first reduction potential prediction value; Step 35, bringing the first redox property prediction sequence and the first property label sequence of the current test record into a preset first model loss function to calculate a corresponding first loss value; Wherein, the first model loss function at least includes L1 loss function, L2 loss function and cross-entropy loss function; Step 36, identifying whether the first loss value meets a preset first loss value range; if the first loss value meets the first loss value range, identifying whether the current test record is the last first data record of the first test set, if yes, going to step 37, if no, extracting the next first data record of the first test set as a new current test record and returning to step 33; if the first loss value does not meet the first loss value range, identifying the first debugging object, if the first debugging object is a nonlinear regression model, performing one round of model parameter modulation on the model parameters of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module based on a preset first model parameter optimizer towards the direction of making the first model loss function reach a minimum value on the premise that the model parameters of the molecular structure optimization module remain unchanged and returning to step 34 at the end of the current round of parameter modulation, if the first debugging object is a structure optimization model, performing one round of model parameter fine-tuning on the model parameters of the molecular structure optimization module based on a preset second model parameter optimizer towards the direction of making the first model loss function reach a minimum value on the premise that the model parameters of the HOMO / LUMO prediction module, the VIP / VEA prediction module, the AIP / AEA prediction module and the oxidation / reduction potential prediction module remain unchanged and returning to step 34 at the end of the current round of parameter fine-tuning; Wherein, the first model parameter optimizer at least includes an SGD optimizer, an ADAM optimizer; the second model parameter optimizer at least includes an RMSprop optimizer, an AdamW optimizer and an ADAM optimizer; Step 37, identifying whether the first debugging object is a structure optimization model; if yes, going to step 38; if no, setting the first debugging object as a structure optimization model, extracting the first first data record of the first test set as a new current test record and returning to step 33; Step 38, traversing all the first data records of the first evaluation set; and in the traversal process, taking the first data record currently traversed as a corresponding current evaluation record; and taking the first training molecular structure type and the first training molecular structure information of the current evaluation record to form a corresponding second training formula component information; and inputting the second training formula component information into the molecular oxidation-reduction property prediction model to perform four-group oxidation-reduction property prediction and obtain a corresponding second oxidation-reduction property prediction sequence; and taking the second oxidation-reduction property prediction sequence and the first property label sequence of the current test record to form a corresponding prediction-label pair; and at the end of the traversal, taking all the prediction-label pairs obtained into a preset first model evaluation function to obtain a corresponding first evaluation value; The first model evaluation function at least includes an RMSE function; Step 39, whether the first evaluation value meets the preset first evaluation value range is identified; if not, return to step 31 to continue training; if yes, stop training and confirm that the model training is completed.
4. The method of claim 1, wherein the process is for predicting the redox properties of a component of an electrolyte formulation. The first property analysis report corresponding to the analysis of the redox properties of the first formula ingredients by using the molecular redox property prediction model is fed back to the user, specifically including: The first formula ingredient information of the first formula is input into the molecular redox property prediction model to perform four-group redox property prediction, and the obtained redox property prediction sequence is taken as the first redox property analysis sequence corresponding thereto; each first formula ingredient information and the first redox property analysis sequence corresponding thereto form a corresponding first component analysis record; and all the first component analysis records obtained form the first property analysis report corresponding thereto to feed back to the user.
5. An apparatus for performing the process of predicting the redox properties of components of an electrolyte formulation according to any one of claims 1-4, characterized in that, The device comprises a model construction module, a model training module and a model application module; The model construction module is used to construct a molecular redox property prediction model; the molecular redox property prediction model is used to perform four-group redox property prediction according to input formula ingredient information and output a corresponding redox property prediction sequence; wherein the formula ingredient information includes molecular structure type and molecular structure information; the molecular structure type includes one-dimensional, two-dimensional and three-dimensional types; when the molecular structure type is one-dimensional type, the corresponding molecular structure information is a one-dimensional molecular sequence, when the molecular structure type is two-dimensional type, the corresponding molecular structure information is a two-dimensional molecular topology graph, and when the molecular structure type is three-dimensional type, the corresponding molecular structure information is a three-dimensional molecular conformation; the redox property prediction sequence includes HOMO prediction value, LUMO prediction value, VIP prediction value, VEA prediction value, AIP prediction value, AEA prediction value, oxidation potential prediction value and reduction potential prediction value; The first model evaluation function at least includes an RMSE function; The model training module is configured to train the molecular redox property prediction model based on a preset first data set; the first data set includes a plurality of first data records; the first data record includes a first training molecular structure type, first training molecular structure information, and a first property label sequence; the first training molecular structure type includes one-dimensional, two-dimensional, and three-dimensional types; when the first training molecular structure type is one-dimensional, the corresponding first training molecular structure information is a one-dimensional molecular sequence; when the first training molecular structure type is two-dimensional, the corresponding first training molecular structure information is a two-dimensional molecular topology graph; when the first training molecular structure type is three-dimensional, the corresponding first training molecular structure information is a three-dimensional molecular conformation; the first property label sequence includes a first HOMO label value, a first LUMO label value, a first VIP label value, a first VEA label value, a first AIP label value, a first AEA label value, a first oxidation potential label value, and a first reduction potential label value; The model application module is configured to, after the model training is completed, receive a user inputted arbitrary electrolyte formula as a corresponding first formula; and analyze the redox properties of the formula components of the first formula using the molecular redox property prediction model to obtain a corresponding first property analysis report for user feedback; the first formula includes a plurality of first formula component information; the first formula component information includes a first molecular structure type and first molecular structure information; the first molecular structure type includes one-dimensional, two-dimensional, and three-dimensional types; when the first molecular structure type is one-dimensional, the corresponding first molecular structure information is a one-dimensional molecular sequence; when the first molecular structure type is two-dimensional, the corresponding first molecular structure information is a two-dimensional molecular topology graph; when the first molecular structure type is three-dimensional, the corresponding first molecular structure information is a three-dimensional molecular conformation; the first property analysis report includes a plurality of first component analysis records; the first component analysis record includes the first formula component information and a corresponding first redox property analysis sequence.
6. An electronic device, comprising: comprise: a memory, a processor, and a transceiver; the processor is configured to couple with the memory, read and execute instructions in the memory to implement the method of any one of claims 1-4; the transceiver is coupled with the processor, and is controlled by the processor to perform message transceiving.
7. A computer readable storage medium characterized in that, The computer readable storage medium stores computer instructions, when the computer instructions are executed by a computer, the computer executes the method of any one of claims 1-4.
Citation Information
Patent Citations
Workflow task processing system for simulating lithium battery electrolyte
CN116700820A
Prediction method for electrochemical performance of electrolyte
CN117437999A