A processing method and device for structural optimization of electrolyte formula components
By constructing a Uni-Mol model and performing pre-training and model tuning, the problem of the quality of molecular structure modeling of electrolyte formulas being affected by human factors was solved, and more efficient three-dimensional structure optimization and modeling quality improvement were achieved.
Patent Information
- Application Number
- CN202411715293.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-11-27
AI Technical Summary
In the existing technology, the quality of molecular structure modeling of electrolyte formulas is greatly affected by human factors, which increases the difficulty and cost of design and development.
A structural optimization model based on the Uni-Mol model was constructed. A pre-training dataset was constructed through big data collection, and a fine-tuning dataset was constructed through molecular dynamics simulation. The model was pre-trained and tuned to achieve three-dimensional structural optimization of the electrolyte formula components.
It reduces the impact of human factors on the quality of molecular structure modeling, improves modeling quality and efficiency, and can output optimized three-dimensional molecular structures end-to-end.
Smart Images

Figure CN119560057B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a processing method and device for optimizing the structure of electrolyte formula components. BACKGROUND
[0002] An electrolyte formula is composed of three components, namely solvents, electrolytes and additives, and each component corresponds to a molecular structure. When designing and developing an electrolyte formula, the physical and chemical properties of each component corresponding to the formula need to be analyzed. This analysis process generally includes the following steps: first, the molecular structure of each component of the electrolyte formula is modeled in three dimensions through a chemical informatics tool (such as OpenBabel, RDKit, etc.), then the formula component structure output by the modeling is used to simulate molecular motion and chemical reactions based on a molecular dynamics (MD) simulation tool (such as LAMMPS, GROMACS, openMM, etc.) to obtain the corresponding solvation structure, and finally, the solvation structure is used to calculate the properties using a quantum chemistry calculation tool (such as Gaussian, GAMESS, etc.). In practice, we found that the worse the quality of the molecular structure modeling, the lower the accuracy of the corresponding property analysis. The quality of the molecular structure modeling is closely related to the professional knowledge and industry background of the modeling personnel, i.e., the quality of the molecular structure modeling is affected by human factors, which increases the difficulty and cost of design and development. SUMMARY
[0003] The present application is aimed at overcoming the shortcomings of the prior art and provides a processing method and device for optimizing the structure of electrolyte formula components, electronic equipment and computer readable storage medium. The present application constructs a structure optimization model based on the Uni-Mol model, which is used to perform three-dimensional structure optimization processing on the original molecular structure input by the model and output the corresponding optimized molecular structure. A pre-training data set, i.e., a first data set, is constructed through big data collection, and a fine-tuning data set, i.e., a second data set, is constructed through molecular dynamics simulation. The structure optimization model is pre-trained based on the first data set, enabling the structure optimization model to fully learn the optimal solution of the compound molecular structure. After the pre-training is completed, the structure optimization model is fine-tuned based on the second data set, enabling the structure optimization model to further fully learn the optimal solution of the molecular structure of each electrolyte component. After the model fine-tuning is completed, the molecular structure of the formula components of any electrolyte formula is optimized based on the structure optimization model. The structure optimization model of the present application can output the optimized three-dimensional molecular structure end-to-end. If the structure optimization model of the present application is used as an optimization tool for regular molecular structure modeling work, it can not only reduce the impact of human factors on the quality of molecular structure modeling, but also improve the modeling quality and efficiency.
[0004] To achieve the above object, the first aspect of the embodiment of the present application provides a processing method for structural optimization of electrolyte formula components, comprising:
[0005] constructing a structural optimization model; the structural optimization model is realized based on a Uni-Mol model; the structural optimization model is used for three-dimensional structural optimization processing of original molecular structures input by the model and output of corresponding optimized molecular structures;
[0006] constructing a corresponding first data set through big data collection; and constructing a corresponding second data set through molecular dynamics simulation;
[0007] pre-training the structural optimization model based on the first data set; and after the pre-training ends, model tuning of the structural optimization model based on the second data set;
[0008] After the model tuning ends, input of a user as a corresponding current electrolyte formula; and using a preset chemoinformatics tool to perform corresponding three-dimensional molecular structure modeling according to molecular sequences of each component of the current electrolyte formula to obtain corresponding component modeling structures; and performing three-dimensional structural optimization processing on each of the component modeling structures based on the structural optimization model to obtain corresponding component optimized structures; and each of the component molecular sequences, the corresponding component type and the component optimized structure form a corresponding formula component structure optimization record; and all the obtained formula component structure optimization records form a corresponding formula structure optimization report to feedback to the user; the current electrolyte formula includes multiple formula components; each of the formula components includes a component type and the component molecular sequence; the component type includes a solvent, an electrolyte and an additive; the component molecular sequence is a one-dimensional SMILES sequence of a solvent molecule, an electrolyte molecule or an additive molecule corresponding to the component type; and the chemoinformatics tool at least includes OpenBabel, RDKit.
[0009] Preferably, the original and optimized molecular structures input and output by the structural optimization model are composed of corresponding atomic tensors and chemical bond tensors;
[0010] The atomic tensor is composed of multiple atomic vectors; each of the atomic vectors corresponds to an atom; each of the atomic vectors at least includes an atomic identifier, an atomic element type and an atomic coordinate; the atomic coordinate is a three-dimensional coordinate; and the atomic vectors in the original and optimized molecular structures are one-to-one aligned based on the atomic identifier;
[0011] The chemical bond tensor is composed of a plurality of chemical bond vectors; each of the chemical bond vectors corresponds to a chemical bond; each of the chemical bond vectors at least includes a chemical bond identifier, a chemical bond bond length and two corresponding atom identifiers; the chemical bond vectors in the original, optimized molecular structure are one-to-one aligned based on the chemical bond identifier;
[0012] The first data set includes a plurality of first data records; the first data record includes a first training molecular structure and a first label molecular structure; each of the first data records corresponds to a three-dimensional molecular structure of a compound;
[0013] The second data set includes a plurality of second data records; the second data record includes a second training molecular structure and a second label molecular structure; each of the second data records corresponds to a three-dimensional molecular structure of an electrolyte solvent, electrolyte or additive.
[0014] Preferably, the corresponding first data set is constructed by big data collection, specifically including:
[0015] A large number of three-dimensional molecular structures of compounds are collected through a plurality of public channels to form a corresponding first collection data set; the plurality of public channels at least include a public compound molecular structure database and a public technical literature / scientific journal / paper; the first collection data set includes a plurality of first three-dimensional molecular structures; the first three-dimensional molecular structure is composed of a corresponding atom tensor and a chemical bond tensor;
[0016] All of the first three-dimensional molecular structures of the first collection data set are traversed; and during traversal, the first three-dimensional molecular structure currently traversed is taken as a corresponding current molecular structure; and the current molecular structure is taken as a corresponding first label molecular structure; then one or more atom vectors are randomly selected from the atom tensor of the current molecular structure as corresponding to-be-masked vectors; and the atom coordinates of each to-be-masked vector are modified based on a preset coordinate perturbation rule, and the atom element type of each to-be-masked vector is replaced with a preset empty element type; and the modified current molecular structure is taken as a corresponding first training molecular structure; and the first training molecular structure and the first label molecular structure corresponding to the current molecular structure form a corresponding first data record; and at the end of traversal, all of the first data records obtained form a corresponding first data set;
[0017] The coordinate perturbation rule is to take the atomic coordinates before modification as the center of a sphere, take a preset perturbation radius r as the radius of the sphere, and make a corresponding spherical space, which is called the corresponding current perturbation space, and select a point coordinate in the current perturbation space as the atomic coordinates after modification; the preset perturbation radius r is less than or equal to
[0018] Preferably, the constructing a corresponding second data set through molecular dynamics simulation specifically comprises:
[0019] A large number of electrolyte formula compositions are collected through various public channels to obtain a corresponding second collection data set; the various public channels at least include a public electrolyte formula database and a public technical literature / scientific journal / paper; the second collection data set includes a plurality of first electrolyte formulas; the first electrolyte formula includes a plurality of first formula components; each first formula component includes a first component type and a first component molecular sequence; the first component type includes a solvent, an electrolyte, and an additive; the first component molecular sequence is a one-dimensional SMILES sequence of a solvent molecule, an electrolyte molecule, or an additive molecule corresponding to the first component type;
[0020] and using the chemical informatics tool to perform corresponding three-dimensional molecular structure modeling according to each first component molecular sequence to obtain a corresponding first modeling structure; the first modeling structure is composed of a corresponding atomic tensor and a chemical bond tensor;
[0021] and using a preset molecular dynamics simulation tool to perform A rounds of iterative molecular motion simulation on each first modeling structure according to a preset first total number of iterations A; and after the number of iterations exceeds a preset first iteration number B, each time an iteration simulation is completed, the current simulation trajectory file is saved as a corresponding first backup trajectory file, and 0 < B < A; and at the end of the iteration simulation, three-dimensional molecular structure recognition is performed according to all the first backup trajectory files to obtain a corresponding first recognition structure; the molecular dynamics simulation tool at least includes LAMMPS, GROMACS, and openMM; the first recognition structure is composed of a corresponding atomic tensor and a chemical bond tensor; the atomic vectors in the first modeling structure and the first recognition structure are one-to-one aligned based on the atomic identifiers; the chemical bond vectors in the first modeling structure and the first recognition structure are one-to-one aligned based on the chemical bond identifiers;
[0022] and the first modeling structure and the first recognition structure corresponding to each first component molecular sequence form a corresponding second training molecular structure and a second label molecular structure to form a corresponding second data record; and all the second data records obtained form a corresponding second data set.
[0023] Further, performing three-dimensional molecular structure recognition on all the obtained first backup trajectory files to obtain corresponding first recognition structures specifically includes:
[0024] Performing atomic alignment on all the first backup trajectory files and the corresponding first modeling structures; and after completing the atomic alignment, taking each atom in the first first backup trajectory file as a corresponding first atom; and setting the atom identifier and the atom element type corresponding to each first atom as the atom identifier and the atom element type of the aligned atom vector in the first modeling structure; and calculating the mean value of the trajectory coordinates of each first atom in all the first backup trajectory files and taking the calculation result as the corresponding atom coordinate; and forming a new atom vector from the atom identifier, the atom element type, and the atom coordinate corresponding to each first atom; and forming a new atom tensor from all the atom vectors corresponding to all the first atoms;
[0025] And copying the chemical bond tensor of the first modeling structure to obtain a new chemical bond tensor; and traversing each chemical bond vector of the new chemical bond tensor; and during the traversal, taking the currently traversed chemical bond vector as the corresponding current vector; and calculating the straight-line distance between the two atom coordinates corresponding to the two atom identifiers of the current vector to obtain a corresponding first distance, and resetting the chemical bond length of the current vector based on the first distance;
[0026] And after the traversal ends, forming the corresponding first recognition structure from the newly generated atom tensor and chemical bond tensor this time.
[0027] Preferably, pre-training the structure optimization model based on the first data set specifically includes:
[0028] Step 601, randomly splitting the first data set into two data sets denoted as corresponding first training set and first evaluation set based on a preset first training evaluation ratio;
[0029] Among them, both the first training set and the first evaluation set include multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first training evaluation ratio;
[0030] Step 602, taking the first first data record in the first training set as the corresponding current test record;
[0031] Step 603, extract the atom identifier and the atom element type of each atom vector in the atomic tensor of the first label molecular structure of the current test record to form a corresponding first type label vector, and extract the atom identifier and the atom coordinates of each atom vector to form a corresponding first coordinate label vector; and form a corresponding first type label tensor from all the first type label vectors obtained, and form a corresponding first coordinate label tensor from all the first coordinate label vectors obtained; and extract the bond identifier and the bond bond length of each bond vector in the bond tensor of the first label molecular structure to form a corresponding first bond length label vector, and form a corresponding first bond length label tensor from all the first bond length label vectors obtained;
[0032] Step 604, input the first training molecular structure of the current test record into the structure optimization model to perform three-dimensional structure optimization processing to obtain a corresponding first predicted molecular structure;
[0033] Wherein, the first predicted molecular structure includes a corresponding atomic tensor and a bond tensor;
[0034] Step 605, extract the atom identifier and the atom element type of each atom vector in the atomic tensor of the first predicted molecular structure as a corresponding first type predicted vector, and extract the atom identifier and the atom coordinates of each atom vector as a corresponding first coordinate predicted vector; and form a corresponding first type predicted tensor from all the first type predicted vectors obtained, and form a corresponding first coordinate predicted tensor from all the first coordinate predicted vectors obtained; and extract the bond identifier and the bond bond length of each bond vector in the bond tensor of the first predicted molecular structure to form a corresponding first bond length predicted vector, and form a corresponding first bond length predicted tensor from all the first bond length predicted vectors obtained;
[0035] Step 606, form a corresponding first prediction-label pair from the first type predicted tensor obtained and the corresponding first type label tensor; and form a corresponding second prediction-label pair from the first coordinate predicted tensor obtained and the corresponding first coordinate label tensor; and form a corresponding third prediction-label pair from the first bond length predicted tensor obtained and the corresponding first bond length label tensor;
[0036] Step 607, the first, second and third prediction-label pairs are brought into preset first, second and third loss functions to calculate corresponding first, second and third loss values; and the first, second and third loss values are weighted and summed to calculate the calculation result as the corresponding first total loss value;
[0037] The first loss function at least includes an L1 loss function, an L2 loss function and a multi-class cross-entropy loss function; the second loss function at least includes an L1 loss function, an L2 loss function and an MSE loss function; and the third loss function at least includes an L1 loss function, an L2 loss function and an MSE loss function.
[0038] Step 608, whether the first loss value meets a preset first loss value range is identified; if yes, step 609 is turned to; if no, a first model parameter optimizer is used to perform a round of parameter optimization on the structure optimization model in a direction of making the first loss function reach a minimum value, and step 604 is returned at the end of the round of parameter optimization;
[0039] The first model parameter optimizer at least includes an SGD optimizer and an ADAM optimizer.
[0040] Step 609, whether the second loss value meets a preset second loss value range is identified; if yes, step 610 is turned to; if no, a second model parameter optimizer is used to perform a round of parameter optimization on the structure optimization model in a direction of making the second loss function reach a minimum value, and step 604 is returned at the end of the round of parameter optimization;
[0041] The second model parameter optimizer at least includes an SGD optimizer and an ADAM optimizer.
[0042] Step 610, whether the third loss value meets a preset third loss value range is identified; if yes, step 611 is turned to; if no, a third model parameter optimizer is used to perform a round of parameter optimization on the structure optimization model in a direction of making the third loss function reach a minimum value, and step 604 is returned at the end of the round of parameter optimization;
[0043] The third model parameter optimizer at least includes an SGD optimizer and an ADAM optimizer.
[0044] Step 611, identify whether the first total loss value meets a preset fourth loss value range; if yes, go to step 612; if not, perform a round of parameter optimization on the structure optimization model based on a preset fourth model parameter optimizer towards a direction of minimizing the first total loss value, and return to step 604 at the end of the round of parameter optimization;
[0045] The fourth model parameter optimizer at least includes an SGD optimizer and an ADAM optimizer.
[0046] Step 612, identify whether the current test record is the last first data record in the first training set; if yes, go to step 613; if not, take the next first data record in the first training set as a new current test record and return to step 603;
[0047] Step 613, count the total number of the first data records in the first evaluation set to obtain a corresponding record total number N; and traverse all the first data records in the first evaluation set; and in the traversal, take the first data record being currently traversed as a corresponding current evaluation record; and input the first training molecular structure of the current evaluation record into the structure optimization model to perform three-dimensional structure optimization processing to obtain a corresponding first sample predicted structure p,i , 1≤index i≤N; and take the first label molecular structure of the current evaluation record as a corresponding first sample label structure t,i ; and form a corresponding first sample predicted-label pair by the first sample predicted structure p,i and the first sample label structure t,i corresponding to the current evaluation record.
[0048] Step 614, at the end of the traversal of all the first data records in the first evaluation set, bring all the obtained first sample predicted-label pairs into a preset first evaluation function to obtain a corresponding first evaluation value;
[0049] The first evaluation function is implemented based on an RMSD function, specifically:
[0050]
[0051] Step 615, identify whether the first evaluation value meets a preset first evaluation value range; if not, return to step 601 to continue training; if yes, stop continuing training and confirm that the pre-training is ended.
[0052] Preferably, the model tuning of the structure optimization model based on the second data set specifically includes:
[0053] Step 701, randomly split the second data set into two data sets based on a preset second training evaluation ratio, denoted as a corresponding second training set and a second evaluation set;
[0054] Wherein, the second training set and the second evaluation set both include a plurality of second data records; the ratio of the total number of records in the second training set to the total number of records in the second evaluation set meets the second training evaluation ratio;
[0055] Step 702, take the first second data record in the second training set as a corresponding current test record;
[0056] Step 703, extract the atom identifier and the atom element type of each atom vector in the atom tensor of the second label molecular structure of the current test record to form a corresponding second type label vector, and extract the atom identifier and the atom coordinates of each atom vector to form a corresponding second coordinate label vector; and form a corresponding second type label tensor from all the obtained second type label vectors, and form a corresponding second coordinate label tensor from all the obtained second coordinate label vectors; and extract the bond identifier and the bond bond length of each bond vector in the bond tensor of the second label molecular structure to form a corresponding second bond length label vector, and form a corresponding second bond length label tensor from all the obtained second bond length label vectors;
[0057] Step 704, input the second training molecular structure of the current test record into the structure optimization model to perform three-dimensional structure optimization processing to obtain a corresponding second predicted molecular structure;
[0058] Wherein, the second predicted molecular structure includes a corresponding atom tensor and a bond tensor;
[0059] Step 705, extract the atom identifier and the atom element type of each atom vector in the atom tensor of the second predicted molecular structure as a corresponding second type predicted vector, and extract the atom identifier and the atom coordinates of each atom vector as a corresponding second coordinate predicted vector; and form a corresponding second type predicted tensor from all the obtained second type predicted vectors, and form a corresponding second coordinate predicted tensor from all the obtained second coordinate predicted vectors; and extract the bond identifier and the bond bond length of each bond vector in the bond tensor of the second predicted molecular structure to form a corresponding second bond length predicted vector, and form a corresponding second bond length predicted tensor from all the obtained second bond length predicted vectors;
[0060] Step 706, a corresponding fourth prediction-label pair is formed by the obtained second type prediction tensor and the corresponding second type label tensor; and a corresponding fifth prediction-label pair is formed by the obtained second coordinate prediction tensor and the corresponding second coordinate label tensor; and a corresponding sixth prediction-label pair is formed by the obtained second bond length prediction tensor and the corresponding second bond length label tensor;
[0061] Step 707, the fourth, fifth and sixth prediction-label pairs are brought into pre-set fourth, fifth and sixth loss functions to calculate corresponding fourth, fifth and sixth loss values; and the fourth, fifth and sixth loss values are weighted and summed to calculate the calculation result as a corresponding second total loss value;
[0062] Among them, the fourth loss function at least includes L1 loss function, L2 loss function and multi-class cross-entropy loss function; the fifth loss function at least includes L1 loss function, L2 loss function and MSE loss function; the sixth loss function at least includes L1 loss function, L2 loss function and MSE loss function;
[0063] Step 708, whether the fourth loss value meets the pre-set fifth loss value range is identified; if it meets, it goes to step 709; if it does not meet, a round of parameter fine-tuning is performed on the structure optimization model based on the pre-set fifth model parameter optimizer towards the direction of making the fourth loss function reach the minimum value, and returns to step 704 at the end of this round of parameter fine-tuning;
[0064] Among them, the fifth model parameter optimizer at least includes RMSprop optimizer, ADAM optimizer, AdamW optimizer;
[0065] Step 709, whether the fifth loss value meets the pre-set sixth loss value range is identified; if it meets, it goes to step 710; if it does not meet, a round of parameter fine-tuning is performed on the structure optimization model based on the pre-set sixth model parameter optimizer towards the direction of making the fifth loss function reach the minimum value, and returns to step 704 at the end of this round of parameter fine-tuning;
[0066] Among them, the sixth model parameter optimizer at least includes RMSprop optimizer, ADAM optimizer, AdamW optimizer;
[0067] Step 710, whether the sixth loss value meets the pre-set seventh loss value range is identified; if it meets, it goes to step 711; if it does not meet, a round of parameter fine-tuning is performed on the structure optimization model based on the pre-set seventh model parameter optimizer towards the direction of making the sixth loss function reach the minimum value, and returns to step 704 at the end of this round of parameter fine-tuning;
[0068] wherein the seventh model parameter optimizer comprises at least an RMSprop optimizer, an ADAM optimizer, an AdamW optimizer;
[0069] Step 711, identify whether the second total loss value meets a preset eighth loss value range; if yes, go to step 712; if no, based on a preset eighth model parameter optimizer, perform a round of parameter fine-tuning on the structure optimization model in a direction of minimizing the second total loss value, and return to step 704 at the end of the round of parameter fine-tuning;
[0070] wherein the eighth model parameter optimizer comprises at least an RMSprop optimizer, an ADAM optimizer, an AdamW optimizer;
[0071] Step 712, identify whether the current test record is the last second data record in the second training set; if yes, go to step 713; if no, take the next second data record in the second training set as a new current test record and return to step 703;
[0072] Step 713, count the total number of the second data records in the second evaluation set to obtain a corresponding record total number M; and traverse all the second data records in the second evaluation set; and in the traversal, take the currently traversed second data record as a corresponding current evaluation record; and input the second training molecular structure of the current evaluation record into the structure optimization model to obtain a corresponding second sample predicted structure x p,j ,1≤indexj≤M; and take the second label molecular structure of the current evaluation record as a corresponding second sample label structure x t,j ; and form a corresponding second sample predicted-label pair by the second sample predicted structure x p,j and the second sample label structure x t,j of the current evaluation record;
[0073] Step 714, at the end of the traversal of all the second data records in the second evaluation set, bring all the obtained second sample predicted-label pairs into a preset second evaluation function to obtain a corresponding second evaluation value;
[0074] wherein the second evaluation function is implemented based on an RMSD function, specifically:
[0075]
[0076] Step 715, whether the second evaluation value meets the preset second evaluation value range is identified; if not, return to step 701 to continue training; if yes, stop continuing training and confirm that the model optimization is completed.
[0077] The second aspect of the embodiment of the application provides a device for implementing the processing method for optimizing the components of the electrolyte formula structure as described in the first aspect, and the device comprises a model construction module, a data preparation module, a model training module and a model application module.
[0078] The model construction module is used for constructing a structure optimization model; the structure optimization model is realized based on a Uni-Mol model; and the structure optimization model is used for performing three-dimensional structure optimization processing on the original molecular structure input by the model and outputting the corresponding optimized molecular structure.
[0079] The data preparation module is used for constructing a corresponding first data set through big data acquisition and constructing a corresponding second data set through molecular dynamics simulation.
[0080] The model training module is used for pre-training the structure optimization model based on the first data set; and after the pre-training is completed, model optimization is performed on the structure optimization model based on the second data set.
[0081] The model application module is used for, after the model optimization is completed, inputting an electrolyte formula input by a user as a corresponding current electrolyte formula; using a preset chemoinformatics tool to perform corresponding three-dimensional molecular structure modeling according to the molecular sequences of each component of the current electrolyte formula to obtain corresponding component modeling structures; performing three-dimensional structure optimization processing on each component modeling structure based on the structure optimization model to obtain corresponding component optimized structures; and composing a corresponding formula component structure optimization record by each component molecular sequence, the corresponding component type and the component optimized structure; and composing a corresponding formula structure optimization report by all the obtained formula component structure optimization records to feed back to the user; the current electrolyte formula comprises a plurality of formula components; each formula component comprises a component type and a component molecular sequence; the component type comprises a solvent, an electrolyte and an additive; the component molecular sequence is a one-dimensional SMILES sequence of a solvent molecule, an electrolyte molecule or an additive molecule corresponding to the component type; and the chemoinformatics tool at least comprises OpenBabel and RDKit.
[0082] The third aspect of the embodiment of the application provides an electronic device, which comprises a memory, a processor and a transceiver.
[0083] The processor is used for coupling with the memory, reading and executing instructions in the memory to implement the method steps of the first aspect.
[0084] The transceiver is coupled with the processor, and the transceiver is controlled by the processor to perform message transmission and reception.
[0085] The fourth aspect of the embodiment of the present application provides a computer readable storage medium, which stores computer instructions, and the computer instructions make the computer execute the method of the first aspect when executed by the computer.
[0086] The embodiment of the present application provides a processing method and device for optimizing the structure of electrolyte formula components, electronic equipment and computer readable storage medium. From the above content, it can be known that the embodiment of the present application constructs a structure optimization model based on a Uni-Mol model, the structure optimization model is used for performing three-dimensional structure optimization processing on an original molecular structure input by the model and outputting a corresponding optimized molecular structure; and a pre-training data set, that is, a first data set, is constructed through big data collection, a fine-tuning data set, that is, a second data set, is constructed through molecular dynamics simulation; the structure optimization model is pre-trained based on the first data set, so that the structure optimization model fully learns the optimal solution of the compound molecular structure; after the pre-training is completed, the structure optimization model is model tuned based on the second data set, so that the structure optimization model can further fully learn the optimal solution of the molecular structure of each electrolyte component; and after the model tuning is completed, the molecular structure of the formula component of any electrolyte formula is optimized based on the structure optimization model. The structure optimization model of the embodiment of the present application can output an optimized three-dimensional molecular structure end to end; if the structure optimization model of the embodiment of the present application is used as an optimization tool for conventional molecular structure modeling work, not only the influence of human factors on the modeling quality of the molecular structure is reduced, but also the modeling quality and the modeling efficiency are improved. BRIEF DESCRIPTION OF DRAWINGS
[0087] Figure 1 A processing method for optimizing the structure of electrolyte formula components is provided for the embodiment one of the present application;
[0088] Figure 2 A module structure diagram of a processing device for optimizing the structure of electrolyte formula components is provided for the embodiment two of the present application;
[0089] Figure 3 A structure schematic diagram of electronic equipment is provided for the embodiment three of the present application. DETAILED DESCRIPTION
[0090] In order to make the objects, technical solutions and advantages of the present application clearer, the following will further describe the present application in detail with reference to the drawings. Obviously, the described embodiments are only a part but not all of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present application.
[0091] The embodiment one of the present application provides a processing method for structural optimization of electrolyte formula components, as shown in a schematic diagram of the processing method. Figure 1 The embodiment one of the present application provides a processing method for structural optimization of electrolyte formula components, as shown in a schematic diagram of the processing method. The method mainly includes the following steps:
[0092] Step 1, constructing a structural optimization model;
[0093] Here, the structural optimization model of the embodiment of the present application is realized based on the Uni-Mol model, which is used to perform three-dimensional structural optimization processing on the original molecular structure input by the model and output the corresponding optimized molecular structure.
[0094] The detailed information of the Uni-Mol model can be understood based on the public technical document “UNI-MOL:A UNIVERSAL 3D MOLECULAR REPRESENTATION LEARNING FRAMEWORK”, which will not be described further here. As known from the public document, the Uni-Mol model can be combined with different task head modules to construct various business models based on a backbone network, wherein one business model is an end-to-end three-dimensional structural optimization model: connected by a backbone network and three parallel prediction task head modules, the structural optimization model of the embodiment of the present application is realized by referring to this end-to-end three-dimensional structural optimization model. In addition, the backbone network of the Uni-Mol model is realized based on the model structure of the transformer model, and the three parallel prediction task head modules connected with the backbone network are respectively an atom type prediction task head (Atom Type Head) module, an atom coordinate prediction task head (Coord Head) module and a chemical bond length prediction task head (Pair-dist Head) module.
[0095] The original and optimized molecular structures input and output by the structural optimization model of the embodiment of the present application are composed of corresponding atomic tensors and chemical bond tensors; wherein,
[0096] The atomic tensor is composed of multiple atomic vectors; each atomic vector corresponds to an atom; each atomic vector at least includes an atomic identifier, an atomic element type and an atomic coordinate; the atomic coordinate is a three-dimensional coordinate; it should be noted that the atomic vectors in the original and optimized molecular structures of the embodiment of the present application are one-to-one aligned based on the atomic identifier.
[0097] The chemical bond tensor is composed of a plurality of chemical bond vectors; each chemical bond vector corresponds to a chemical bond; each chemical bond vector at least includes a chemical bond identifier, a chemical bond bond length and two corresponding atom identifiers; it should be noted that the chemical bond vectors in the original, optimized molecular structure of the embodiment of the present application are aligned one by one based on the chemical bond identifier.
[0098] Step 2, constructing a corresponding first data set through big data collection; and constructing a corresponding second data set through molecular dynamics simulation;
[0099] Specifically includes: step 21, constructing a corresponding first data set through big data collection;
[0100] The first data set includes a plurality of first data records; the first data record includes a first training molecular structure and a first label molecular structure; each first data record corresponds to a three-dimensional molecular structure of a compound;
[0101] Specifically includes: step 211, collecting a large number of three-dimensional molecular structures of compounds through a plurality of public channels to form a corresponding first collection data set;
[0102] The plurality of public channels at least include a public compound molecular structure database and a public technical literature / scientific journal / paper; the first collection data set includes a plurality of first three-dimensional molecular structures; the first three-dimensional molecular structure is composed of a corresponding atomic tensor and a chemical bond tensor;
[0103] Here, a large number of three-dimensional molecular structures can also be generated through chemical informatics tools (such as OpenBabel, RDKit) to add data items to the first collection data set;
[0104] Step 212, traversing all first three-dimensional molecular structures of the first collection data set; and during traversal, taking the currently traversed first three-dimensional molecular structure as a corresponding current molecular structure; and taking the current molecular structure as a corresponding first label molecular structure; then randomly selecting one or more atomic vectors from the atomic tensor of the current molecular structure as corresponding to-be-masked vectors; and modifying the atomic coordinates of each to-be-masked vector based on a preset coordinate perturbation rule, and replacing the atomic element type of each to-be-masked vector with a preset empty element type; and taking the modified current molecular structure as a corresponding first training molecular structure; and taking the first training molecular structure and the first label molecular structure corresponding to the current molecular structure to form a corresponding first data record; and at the end of traversal, taking all the obtained first data records to form a corresponding first data set;
[0105] Here, the coordinate perturbation rule of the embodiment of the application is to make a corresponding spherical space with the atomic coordinates before modification as the center of the sphere and a preset perturbation radius r as the radius of the sphere, denoted as a corresponding current perturbation space, and to select a point coordinate in the current perturbation space as the atomic coordinates after modification; the empty element type is a preset integer, which is set as 0 or a negative number by default; the preset perturbation radius r is less than or equal to
[0106] Step 22, and constructing a corresponding second data set through molecular dynamics simulation;
[0107] The second data set includes a plurality of second data records; each second data record corresponds to a three-dimensional molecular structure of an electrolyte solvent, electrolyte or additive; and the second data record includes a second training molecular structure and a second label molecular structure.
[0108] Specifically, step 221 includes: collecting a large number of electrolyte formulae through a plurality of public channels to obtain a corresponding second collection data set; and step 222 includes: using a preset cheminformatics tool to model a three-dimensional molecular structure according to each first component molecular sequence to obtain a corresponding first modeling structure.
[0109] The plurality of public channels at least include a public electrolyte formula database and public technical literature / scientific journals / papers; the second collection data set includes a plurality of first electrolyte formulae; each first electrolyte formula includes a plurality of first formula components; each first formula component includes a first component type and a first component molecular sequence; the first component type includes a solvent, an electrolyte and an additive; and the first component molecular sequence is a one-dimensional SMILES sequence of a solvent molecule, an electrolyte molecule or an additive molecule corresponding to the first component type.
[0110] Step 222 includes: using a preset cheminformatics tool to model a three-dimensional molecular structure according to each first component molecular sequence to obtain a corresponding first modeling structure.
[0111] The first modeling structure is composed of a corresponding atomic tensor and a chemical bond tensor.
[0112] Here, the cheminformatics tool of the embodiment of the application at least includes OpenBabel and RDKit.
[0113] Step 223 includes: using a preset molecular dynamics simulation tool to perform A rounds of iterative molecular motion simulation on each first modeling structure according to a preset first total number of iterations A; after the number of iterations exceeds a preset first iteration number B, saving the current simulation trajectory file as a corresponding first backup trajectory file after completing each iteration simulation; and when the iteration simulation ends, identifying a three-dimensional molecular structure according to all the first backup trajectory files to obtain a corresponding first identified structure.
[0114] Here, the molecular dynamics simulation tool of the embodiment of the application at least includes LAMMPS, GROMACS, openMM; the first total number of iterations A and the first number of iterations B are two pre-set integers, 0 < B < A; the first identified structure is composed of the corresponding atomic tensor and the bond tensor; the atomic vectors in the first modeling structure and the first identified structure are one-to-one aligned based on the atomic identification; the bond vectors in the first modeling structure and the first identified structure are one-to-one aligned based on the bond identification;
[0115] Wherein, the corresponding first identified structure is obtained by three-dimensional molecular structure identification according to all obtained first backup trajectory files, specifically including:
[0116] Step A1, aligning atoms for all first backup trajectory files and corresponding first modeling structures; and after completing the atomic alignment, taking each atom in the first first backup trajectory file as the corresponding first atom; and setting the atomic identification and atomic element type of each first atom as the atomic identification and atomic element type of the aligned atomic vector in the first modeling structure; and performing mean value calculation on the trajectory coordinates of each first atom in all first backup trajectory files and taking the calculation result as the corresponding atomic coordinates; and composing a new atomic vector from the atomic identification, atomic element type and atomic coordinates corresponding to each first atom; and composing a new atomic tensor from all atomic vectors corresponding to all first atoms;
[0117] Step A2, copying the bond tensor of the first modeling structure to obtain a new bond tensor; and traversing each bond vector of the new bond tensor; and in the traversal process, taking the currently traversed bond vector as the corresponding current vector; and calculating the straight line distance between the atomic coordinates of the two first atoms corresponding to the two atomic identifications of the current vector to obtain the corresponding first distance, and resetting the bond length of the current vector based on the first distance;
[0118] Step A3, and after the traversal ends, composing the corresponding first identified structure from the newly generated atomic tensor and bond tensor this time;
[0119] Step 224, and composing a corresponding second data record from the first modeling structure and the first identified structure corresponding to each first component molecular sequence as the corresponding second training molecular structure and the second labeled molecular structure; and composing a corresponding second data set from all obtained second data records.
[0120] Step 3, pre-training the structure optimization model based on the first data set; and after the pre-training ends, model tuning of the structure optimization model based on the second data set;
[0121] Specifically comprising: step 31, pre-training the structure optimization model based on the first data set;
[0122] Specifically comprising: step 3101, randomly splitting the first data set into two data sets based on a preset first training evaluation ratio, denoted as a corresponding first training set and a first evaluation set;
[0123] Here, the first training evaluation ratio is a preset ratio, for example, 8:2; the first training set and the first evaluation set both include a plurality of first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first training evaluation ratio;
[0124] Step 3102, taking the first first data record in the first training set as a corresponding current test record;
[0125] Step 3103, extracting the atom identifier and atom element type of each atom vector in the atom tensor of the first label molecular structure of the current test record to form a corresponding first type label vector, and extracting the atom identifier and atom coordinates of each atom vector to form a corresponding first coordinate label vector; and all first type label vectors form a corresponding first type label tensor, and all first coordinate label vectors form a corresponding first coordinate label tensor; and extracting the bond identifier and bond bond length of each bond vector in the bond tensor of the first label molecular structure to form a corresponding first bond length label vector, and all first bond length label vectors form a corresponding first bond length label tensor;
[0126] Step 3104, inputting the first training molecular structure of the current test record into the structure optimization model to obtain a corresponding first predicted molecular structure through three-dimensional structure optimization processing;
[0127] Wherein, the first predicted molecular structure includes a corresponding atom tensor and a bond tensor;
[0128] Step 3105, extracting the atom identifier and atom element type of each atom vector in the atom tensor of the first predicted molecular structure as a corresponding first type prediction vector, and extracting the atom identifier and atom coordinates of each atom vector as a corresponding first coordinate prediction vector; and all first type prediction vectors form a corresponding first type prediction tensor, and all first coordinate prediction vectors form a corresponding first coordinate prediction tensor; and extracting the bond identifier and bond bond length of each bond vector in the bond tensor of the first predicted molecular structure to form a corresponding first bond length prediction vector, and all first bond length prediction vectors form a corresponding first bond length prediction tensor;
[0129] Step 3106, a corresponding first prediction-label pair is formed by the obtained first type prediction tensor and the corresponding first type label tensor; and a corresponding second prediction-label pair is formed by the obtained first coordinate prediction tensor and the corresponding first coordinate label tensor; and a corresponding third prediction-label pair is formed by the obtained first bond length prediction tensor and the corresponding first bond length label tensor;
[0130] Step 3107, the first, second and third prediction-label pairs are brought into the preset first, second and third loss functions to calculate the corresponding first, second and third loss values; and the first, second and third loss values are weighted and summed to calculate the calculation result as the corresponding first total loss value;
[0131] Among them, the first loss function at least includes L1 loss function, L2 loss function and multi-class cross-entropy loss function; the second loss function at least includes L1 loss function, L2 loss function and MSE loss function; the third loss function at least includes L1 loss function, L2 loss function and MSE loss function;
[0132] Step 3108, whether the first loss value meets the preset first loss value range is identified; if it meets, it is turned to step 3109; if it does not meet, a round of parameter optimization is performed on the structure optimization model based on the preset first model parameter optimizer in the direction of making the first loss function reach the minimum value, and returns to step 3104 at the end of this round of parameter optimization;
[0133] Here, the first loss value range is a pre-set loss value range, which is 0 by default; the first model parameter optimizer at least includes SGD optimizer, ADAM optimizer;
[0134] Step 3109, whether the second loss value meets the preset second loss value range is identified; if it meets, it is turned to step 3110; if it does not meet, a round of parameter optimization is performed on the structure optimization model based on the preset second model parameter optimizer in the direction of making the second loss function reach the minimum value, and returns to step 3104 at the end of this round of parameter optimization;
[0135] Here, the second loss value range is a pre-set loss value range; the second model parameter optimizer at least includes SGD optimizer, ADAM optimizer;
[0136] Step 3110, whether the third loss value meets the preset third loss value range is identified; if it meets, it is turned to step 3111; if it does not meet, a round of parameter optimization is performed on the structure optimization model based on the preset third model parameter optimizer in the direction of making the third loss function reach the minimum value, and returns to step 3104 at the end of this round of parameter optimization;
[0137] Here, the third loss value range is a pre-set loss value range; and the third model parameter optimizer at least includes an SGD optimizer and an ADAM optimizer.
[0138] Step 3111, whether the first total loss value meets a pre-set fourth loss value range is identified; if yes, step 3112 is reached; if no, a round of parameter optimization is performed on the structure optimization model based on a pre-set fourth model parameter optimizer towards a direction of making the first total loss value reach a minimum value, and step 3104 is returned at the end of the round of parameter optimization;
[0139] Here, the fourth loss value range is a pre-set loss value range; and the fourth model parameter optimizer at least includes an SGD optimizer and an ADAM optimizer.
[0140] Step 3112, whether the current test record is the last first data record in the first training set is identified; if yes, step 3113 is reached; if no, the next first data record in the first training set is taken as a new current test record and step 3103 is returned;
[0141] Step 3113, the total number of the first data records in the first evaluation set is counted to obtain a corresponding record total number N; and all the first data records in the first evaluation set are traversed; and in the traversal, the currently traversed first data record is taken as a corresponding current evaluation record; and the first training molecular structure of the current evaluation record is input into the structure optimization model to obtain a corresponding first sample predicted structure x p,i ,1≤i≤N; and the first label molecular structure of the current evaluation record is taken as a corresponding first sample label structure x t,i ; and a corresponding first sample predicted-label pair is formed by the first sample predicted structure x p,i and the first sample label structure x t,i of the current evaluation record.
[0142] Step 3114, at the end of the traversal of all the first data records in the first evaluation set, all the obtained first sample predicted-label pairs are brought into a pre-set first evaluation function to obtain a corresponding first evaluation value;
[0143] Here, the first evaluation function is realized based on an RMSD function, and specifically is:
[0144]
[0145] Step 3115, whether the first evaluation value meets a pre-set first evaluation value range is identified; if no, step 3101 is returned to continue training; if yes, the training is stopped and it is confirmed that the pre-training is ended.
[0146] Here, the first evaluation value range is a pre-set evaluation value range;
[0147] Step 32, and after the pre-training is completed, the model tuning is performed on the structure optimization model based on the second data set;
[0148] Specifically, step 3201, the second data set is randomly divided into two data sets based on a pre-set second training evaluation ratio, denoted as a corresponding second training set and a second evaluation set;
[0149] Here, the second training evaluation ratio is a pre-set ratio, for example, 7:3; the second training set and the second evaluation set both include a plurality of second data records; and the ratio of the total number of records in the second training set to the total number of records in the second evaluation set satisfies the second training evaluation ratio;
[0150] Step 3202, the first second data record in the second training set is taken as a corresponding current test record;
[0151] Step 3203, the atomic identifier and the atomic element type of each atomic vector in the atomic tensor of the second label molecular structure of the current test record are extracted to form a corresponding second type label vector, and the atomic identifier and the atomic coordinates of each atomic vector are extracted to form a corresponding second coordinate label vector; all the second type label vectors are combined to form a corresponding second type label tensor, and all the second coordinate label vectors are combined to form a corresponding second coordinate label tensor; and the bond identifier and the bond length of each bond vector in the bond tensor of the second label molecular structure are extracted to form a corresponding second bond length label vector, and all the second bond length label vectors are combined to form a corresponding second bond length label tensor;
[0152] Step 3204, the second training molecular structure of the current test record is input into the structure optimization model to perform three-dimensional structure optimization processing to obtain a corresponding second predicted molecular structure;
[0153] The second predicted molecular structure includes a corresponding atomic tensor and a bond tensor;
[0154] Step 3205, extract the atom identifier and atom element type of each atom vector in the atomic tensor of the second predicted molecular structure as the corresponding second type prediction vector, and extract the atom identifier and atom coordinates of each atom vector as the corresponding second coordinate prediction vector; and form the corresponding second type prediction tensor from all the obtained second type prediction vectors, and form the corresponding second coordinate prediction tensor from all the obtained second coordinate prediction vectors; and extract the bond identifier and bond bond length of each bond vector in the bond tensor of the second predicted molecular structure to form the corresponding second bond length prediction vector, and form the corresponding second bond length prediction tensor from all the obtained second bond length prediction vectors;
[0155] Step 3206, form a corresponding fourth prediction-label pair from the obtained second type prediction tensor and the corresponding second type label tensor; and form a corresponding fifth prediction-label pair from the obtained second coordinate prediction tensor and the corresponding second coordinate label tensor; and form a corresponding sixth prediction-label pair from the obtained second bond length prediction tensor and the corresponding second bond length label tensor;
[0156] Step 3207, bring the fourth, fifth and sixth prediction-label pairs into the pre-set fourth, fifth and sixth loss functions to calculate the corresponding fourth, fifth and sixth loss values; and calculate the weighted sum of the fourth, fifth and sixth loss values and take the calculation result as the corresponding second total loss value;
[0157] Among them, the fourth loss function at least includes L1 loss function, L2 loss function and multi-class cross-entropy loss function; the fifth loss function at least includes L1 loss function, L2 loss function and MSE loss function; the sixth loss function at least includes L1 loss function, L2 loss function and MSE loss function;
[0158] Step 3208, identify whether the fourth loss value meets the pre-set fifth loss value range; if it meets, go to step 3209; if it does not meet, based on the pre-set fifth model parameter optimizer, adjust the parameters of the structure optimization model in the direction of making the fourth loss function reach the minimum value for one round of parameter fine-tuning, and return to step 3204 when the round of parameter fine-tuning is over;
[0159] Here, the first loss value range is a pre-set loss value range, which is 0 by default; the fifth model parameter optimizer at least includes RMSprop optimizer, ADAM optimizer, AdamW optimizer;
[0160] Step 3209, whether the fifth loss value meets a preset sixth loss value range is identified; if yes, go to step 3210; if no, based on a preset sixth model parameter optimizer, a round of parameter fine-tuning is performed on the structure optimization model in the direction of making the fifth loss function reach a minimum value, and step 3204 is returned at the end of the round of parameter fine-tuning;
[0161] Here, the sixth loss value range is a preset loss value range; the sixth model parameter optimizer at least includes an RMSprop optimizer, an ADAM optimizer and an AdamW optimizer;
[0162] Step 3210, whether the sixth loss value meets a preset seventh loss value range is identified; if yes, go to step 3211; if no, based on a preset seventh model parameter optimizer, a round of parameter fine-tuning is performed on the structure optimization model in the direction of making the sixth loss function reach a minimum value, and step 3204 is returned at the end of the round of parameter fine-tuning;
[0163] Here, the seventh loss value range is a preset loss value range; the seventh model parameter optimizer at least includes an RMSprop optimizer, an ADAM optimizer and an AdamW optimizer;
[0164] Step 3211, whether the second total loss value meets a preset eighth loss value range is identified; if yes, go to step 3212; if no, based on a preset eighth model parameter optimizer, a round of parameter fine-tuning is performed on the structure optimization model in the direction of making the second total loss value reach a minimum value, and step 3204 is returned at the end of the round of parameter fine-tuning;
[0165] Here, the eighth loss value range is a preset loss value range; the eighth model parameter optimizer at least includes an RMSprop optimizer, an ADAM optimizer and an AdamW optimizer;
[0166] Step 3212, whether the current test record is the last second data record in the second training set is identified; if yes, go to step 3213; if no, the next second data record in the second training set is taken as a new current test record and step 3203 is returned;
[0167] Step 3213, the total number of second data records in the second evaluation set is counted to obtain a corresponding record total number M; all second data records in the second evaluation set are traversed; when traversing, the currently traversed second data record is taken as a corresponding current evaluation record; and the second training molecular structure of the current evaluation record is input into the structure optimization model to obtain a corresponding second sample predicted structure x p,j, 1≤index j≤M; and the second label molecular structure of the current evaluation record as a corresponding second sample label structure x t,j ; and the second sample prediction structure x p,j and the second sample label structure x t,j comprise a corresponding second sample prediction-label pair;
[0168] Step 3214, at the end of the traversal of all second data records of the second evaluation set, all obtained second sample prediction-label pairs are brought into a preset second evaluation function for calculation to obtain a corresponding second evaluation value;
[0169] Wherein, the second evaluation function is implemented based on the RMSD function, specifically:
[0170]
[0171] Step 3215, whether the second evaluation value satisfies a preset second evaluation value range is identified; if not, return to step 3201 for continuous training; if yes, stop continuous training and confirm that the model tuning is ended.
[0172] Here, the second evaluation value range is a preset evaluation value range.
[0173] Step 4, after the model tuning is ended, the electrolyte formula input by the user is taken as a corresponding current electrolyte formula; and a preset cheminformatics tool is used to perform corresponding three-dimensional molecular structure modeling on the molecular sequences of each component of the current electrolyte formula to obtain corresponding component modeling structures; and three-dimensional structure optimization processing is performed on each component modeling structure based on the structure optimization model to obtain corresponding component optimized structures; and each component molecular sequence, the corresponding component type and the component optimized structure comprise a corresponding formula component structure optimization record; and all obtained formula component structure optimization records comprise a corresponding formula structure optimization report to feedback to the user;
[0174] Wherein, the current electrolyte formula comprises multiple formula components; each formula component comprises a component type and a component molecular sequence; the component type comprises a solvent, an electrolyte and an additive; and the component molecular sequence is a one-dimensional SMILES sequence of a solvent molecule, an electrolyte molecule or an additive molecule corresponding to the component type.
[0175] Figure 2 A module structure diagram of a processing device for structure optimization of electrolyte formula components provided by the second embodiment of the present application, the device is a terminal device or a server for implementing the foregoing method embodiments, and can also be a device capable of enabling the foregoing terminal device or server to implement the foregoing method embodiments, for example, the device can be a device or a chip system of the foregoing terminal device or server. As Figure 2As shown, the device comprises a model construction module 201, a data preparation module 202, a model training module 203 and a model application module 204.
[0176] The model construction module 201 is configured to construct a structure optimization model; the structure optimization model is implemented based on a Uni-Mol model; and the structure optimization model is configured to perform three-dimensional structure optimization processing on a model input original molecular structure and output a corresponding optimized molecular structure.
[0177] The data preparation module 202 is configured to construct a corresponding first data set through big data acquisition; and construct a corresponding second data set through molecular dynamics simulation.
[0178] The model training module 203 is configured to pre-train the structure optimization model based on the first data set; and after the pre-training ends, perform model tuning on the structure optimization model based on the second data set.
[0179] The model application module 204 is configured to, after the model tuning ends, input a user input electrolyte formula as a corresponding current electrolyte formula; use a preset chemoinformatics tool to perform corresponding three-dimensional molecular structure modeling according to each component molecular sequence of the current electrolyte formula to obtain a corresponding component modeling structure; perform three-dimensional structure optimization processing on each component modeling structure based on the structure optimization model to obtain a corresponding component optimized structure; and form a corresponding formula component structure optimization record by each component molecular sequence, a corresponding component type and a component optimized structure; and form a corresponding formula structure optimization report by all obtained formula component structure optimization records to feedback to the user; the current electrolyte formula comprises a plurality of formula components; each formula component comprises a component type and a component molecular sequence; the component type comprises a solvent, an electrolyte and an additive; the component molecular sequence is a one-dimensional SMILES sequence of a solvent molecule, an electrolyte molecule or an additive molecule corresponding to the component type; and the chemoinformatics tool at least comprises OpenBabel and RDKit.
[0180] The processing device for performing structure optimization on electrolyte formula components provided by the embodiment of the present application can execute the method steps in the method embodiment, and has similar implementation principles and technical effects, which will not be described here.
[0181] It should be noted that the division of the various modules of the above apparatus is only a logical functional division, and in actual implementation, all or part of them can be integrated into one physical entity, or can be physically separated. These modules can all be implemented in the form of software invoked by a processing element; all can be implemented in the form of hardware; or some modules can be implemented in the form of software invoked by a processing element, and some modules can be implemented in the form of hardware. For example, the model construction module can be a separately established processing element, or can be integrated in a certain chip of the above apparatus, in addition, it can also be stored in the form of program code in the memory of the above apparatus, and the function of the above determination module is invoked and executed by a certain processing element of the above apparatus. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or can be independently implemented. The processing element described herein can be an integrated circuit with signal processing capability. In the implementation process, each step of the above method or each module can be completed by the integrated logic circuit of the hardware in the processor element or the instructions in the form of software.
[0182] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), etc. For another example, when a certain module above is implemented in the form of program code invoked by a processing element, the processing element can be a general purpose processor, such as a central processing unit (CPU) or other processor that can invoke program code. For another example, these modules can be integrated together to implement in the form of system on a chip (SOC).
[0183] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When loaded and executed by a computer, all or part of the computer program instructions generate the processes or functions described in the foregoing method embodiments. The computer can be a general purpose computer, a special purpose computer, a computer network or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (solid state disk, SSD)) and the like.
[0184] Figure 3 A structural schematic diagram of an electronic device is provided for the third embodiment of the present application. The electronic device can be a terminal device or a server for implementing the method of the foregoing embodiments, or a terminal device or a server connected with the foregoing terminal device or server for implementing the method of the foregoing embodiments. As shown in the figure, the electronic device can include a processor 301 (such as a CPU), a memory 302, a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiving action of the transceiver 303. The memory 302 can store various instructions for completing various processing functions and implementing the processing steps described in the foregoing embodiment methods. Preferably, the electronic device related to the embodiments of the present application further includes a power supply 304, a system bus 305 and a communication port 306. The system bus 305 is used to realize the communication connection between elements. The communication port 306 is used for the connection and communication between the electronic device and other external devices. Figure 3 In
[0185] Figure 3 The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus. The communication interface is used to realize the communication between the database access device and other devices (such as a client, a read-write library and a read-only library). The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory, such as at least one disk memory.
[0186] The processor mentioned above can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; can also be a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0187] It should be noted that the embodiments of the present application also provide a computer readable storage medium, which stores instructions, and when the instructions are run on a computer, the computer executes the method and process provided in the above embodiments.
[0188] The embodiment of the present application provides a processing method and device for structural optimization of electrolyte formula components, electronic equipment and computer readable storage medium. As known from the above, the embodiment of the present application constructs a structural optimization model based on a Uni-Mol model, the structural optimization model is used for performing three-dimensional structural optimization processing on original molecular structures input by the model and outputting corresponding optimized molecular structures; a pre-training data set, that is, a first data set, is constructed through big data collection, a fine-tuning data set, that is, a second data set, is constructed through molecular dynamics simulation; the structural optimization model is pre-trained based on the first data set, so that the structural optimization model fully learns the optimal solution of the compound molecular structure; after the pre-training is completed, the structural optimization model is model-tuned based on the second data set, so that the structural optimization model can further fully learn the optimal solution of the molecular structure of each electrolyte component; and after the model tuning is completed, the formula component molecular structure of any electrolyte formula is optimized based on the structural optimization model. The structural optimization model of the embodiment of the present application can output the optimized three-dimensional molecular structure end to end; if the structural optimization model of the embodiment of the present application is used as an optimization tool for conventional molecular structure modeling work, the influence of human factors on the modeling quality of the molecular structure is reduced, and the modeling quality and efficiency are improved.
[0189] Those skilled in the art should further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0190] The steps of the method or algorithm described in combination with the embodiments disclosed herein can be implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0191] The above specific embodiments further illustrate the purpose, technical solutions and advantages of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A process for the structural optimization of components of an electrolyte formulation, characterized in that The method comprises: constructing a structure optimization model; the structure optimization model is realized based on a Uni-Mol model; the structure optimization model is used for three-dimensional structure optimization processing of a model input original molecular structure and outputs a corresponding optimized molecular structure; constructing a corresponding first data set through big data collection; and constructing a corresponding second data set through molecular dynamics simulation; pre-training the structure optimization model based on the first data set; and after the pre-training ends, model tuning of the structure optimization model based on the second data set; after the model tuning ends, inputting an electrolyte formula input by a user as a corresponding current electrolyte formula; using a preset chemoinformatics tool to perform corresponding three-dimensional molecular structure modeling according to molecular sequences of each component of the current electrolyte formula to obtain corresponding component modeling structures; performing three-dimensional structure optimization processing on each component modeling structure based on the structure optimization model to obtain corresponding component optimized structures; and each component molecular sequence, a corresponding component type, and the component optimized structure form a corresponding formula component structure optimization record; and all obtained formula component structure optimization records form a corresponding formula structure optimization report to feedback to the user; the current electrolyte formula comprises a plurality of formula components; each formula component comprises a component type and a component molecular sequence; the component type comprises a solvent, an electrolyte, and an additive; the component molecular sequence is a one-dimensional SMILES sequence of a solvent molecule, an electrolyte molecule, or an additive molecule corresponding to the component type; and the chemoinformatics tool at least comprises OpenBabel and RDKit; wherein the construction of the corresponding second data set through molecular dynamics simulation specifically comprises: collecting a large number of electrolyte formulas through multiple public channels to form a corresponding second collection data set; the second collection data set comprises a plurality of first electrolyte formulas; the first electrolyte formula comprises a plurality of first formula components; each first formula component comprises a first component type and a first component molecular sequence; using the chemoinformatics tool to perform corresponding three-dimensional molecular structure modeling according to each first component molecular sequence to obtain corresponding first modeling structures; using a preset molecular dynamics simulation tool to perform A rounds of iteration of molecular motion simulation on each first modeling structure according to a preset first total number of iterations A; and after the number of iteration simulations exceeds a preset first iteration number B, each time an iteration simulation is completed, a current simulation trajectory file is taken as a corresponding first backup trajectory file and saved, 0 < B < A; and at the end of the iteration simulation, three-dimensional molecular structure recognition is performed according to all obtained first backup trajectory files to obtain corresponding first recognition structures; each first component molecular sequence corresponding first modeling structure and the first recognition structure as a corresponding second training molecular structure and a second label molecular structure form a corresponding second data record; and all obtained second data records form a corresponding second data set.
2. The method of claim 1, wherein, the input and output of the structure optimization model are composed of atomic tensors and bond tensors; the atomic tensors are composed of atomic vectors; each atomic vector corresponds to an atom; each atomic vector includes at least an atomic identifier, an atomic element type, and an atomic coordinate; the atomic coordinate is a three-dimensional coordinate; the atomic vectors in the original and optimized molecular structures are aligned one by one based on the atomic identifiers; the bond tensors are composed of bond vectors; each bond vector corresponds to a bond; each bond vector includes at least a bond identifier, a bond length, and two corresponding atomic identifiers; the bond vectors in the original and optimized molecular structures are aligned one by one based on the bond identifiers; the first data set includes a plurality of first data records; each first data record includes a first training molecular structure and a first labeled molecular structure; each first data record corresponds to a three-dimensional molecular structure of a compound; the second data set includes a plurality of second data records; each second data record includes a second training molecular structure and a second labeled molecular structure; each second data record corresponds to a three-dimensional molecular structure of an electrolyte solvent, electrolyte, or additive.
3. The process for structural optimization of electrolyte formulation components according to claim 2, characterized in that, the first data set is constructed by big data collection, specifically including: a large number of three-dimensional molecular structures of compounds are collected through a plurality of public channels to form a corresponding first collection data set; the plurality of public channels include at least public compound molecular structure databases and public technical literature / scientific journals / papers; the first collection data set includes a plurality of first three-dimensional molecular structures; each first three-dimensional molecular structure is composed of corresponding atomic tensors and bond tensors; all first three-dimensional molecular structures in the first collection data set are traversed; during traversal, the currently traversed first three-dimensional molecular structure is taken as a corresponding current molecular structure; the current molecular structure is taken as a corresponding first labeled molecular structure; one or more atomic vectors are randomly selected from the atomic tensors of the current molecular structure as corresponding to-be-masked vectors; the atomic coordinates of each to-be-masked vector are modified based on a pre-set coordinate perturbation rule, and the atomic element types of each to-be-masked vector are replaced with a pre-set empty element type; the modified current molecular structure is taken as a corresponding first training molecular structure; the first training molecular structure and the first labeled molecular structure corresponding to the current molecular structure form a corresponding first data record; when the traversal is completed, all the first data records obtained form a corresponding first data set; wherein the coordinate perturbation rule is to make a corresponding spherical space with the atomic coordinate before modification as the center of the sphere and a preset perturbation radius r as the radius of the sphere, denoted as a corresponding current perturbation space, and to select an optional point coordinate in the current perturbation space as the atomic coordinate after modification; the preset perturbation radius r is less than or equal to 4. The method of claim 2, wherein, the plurality of public channels include at least public electrolyte formula databases and public technical literature / scientific journals / papers; The first component type includes solvent, electrolyte and additive; the first component molecular sequence is a one-dimensional SMILES sequence of a solvent molecule, an electrolyte molecule or an additive molecule corresponding to the first component type; The first modeling structure is composed of the corresponding atomic tensor and the bond tensor; The molecular dynamics simulation tool at least includes LAMMPS, GROMACS and openMM; The first identification structure is composed of the corresponding atomic tensor and the bond tensor; The atomic vectors in the first modeling structure and the first identification structure are one-to-one aligned based on the atomic identification; The bond vectors in the first modeling structure and the first identification structure are one-to-one aligned based on the bond identification.
5. The method for optimizing the structure of electrolyte formulation according to claim 4, characterized in that: The first identification structure is obtained by performing three-dimensional molecular structure identification on all the obtained first backup trajectory files, specifically including: Aligning atoms in all the first backup trajectory files and the corresponding first modeling structure; after completing the atomic alignment, taking each atom in the first first backup trajectory file as a corresponding first atom; setting the atomic identification and the atomic element type corresponding to each first atom as the atomic identification and the atomic element type of the aligned atomic vector in the first modeling structure; performing mean value calculation on the trajectory coordinates of each first atom in all the first backup trajectory files and taking the calculation result as the corresponding atomic coordinates; and composing a new atomic vector from the atomic identification, the atomic element type and the atomic coordinates corresponding to each first atom; and composing a new atomic tensor from all the atomic vectors corresponding to all the first atoms; Copying the bond tensor of the first modeling structure to obtain a new bond tensor; traversing each bond vector of the new bond tensor; in the traversal process, taking the current traversed bond vector as a corresponding current vector; calculating the straight line distance between the atomic coordinates of the two first atoms corresponding to the two atomic identifications of the current vector to obtain a corresponding first distance, and resetting the bond length of the current vector based on the first distance; After the traversal ends, the new generated atomic tensor and the bond tensor compose a corresponding first identification structure.
6. The method for optimizing the structure of electrolyte formulation according to claim 2, characterized in that: The structure optimization model is pre-trained based on the first data set, specifically including: Step 601, based on a preset first training evaluation ratio, the first data set is randomly divided into two data sets, which are recorded as a corresponding first training set and a first evaluation set; The first training set and the first evaluation set both include a plurality of first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set meets the first training evaluation ratio; Step 602, taking the first first data record in the first training set as a corresponding current test record; Step 603: extract the atomic identifier and the atomic element type of each atomic vector in the atomic tensor of the first label molecular structure of the current test record to form a corresponding first type label vector, and extract the atomic identifier and the atomic coordinate of each atomic vector to form a corresponding first coordinate label vector; and form a corresponding first type label tensor with all the obtained first type label vectors, and form a corresponding first coordinate label tensor with all the obtained first coordinate label vectors; and extract the chemical bond identifier and the chemical bond length of each chemical bond vector in the chemical bond tensor of the first label molecular structure to form a corresponding first bond length label vector, and form a corresponding first bond length label tensor with all the obtained first bond length label vectors; Step 604: Input the first training molecular structure of the current test record into the structure optimization model to perform three-dimensional structure optimization processing to obtain a corresponding first predicted molecular structure; Wherein, the first predicted molecular structure includes the corresponding atom tensor and the chemical bond tensor; Step 605: extract the atom identifier and the atomic element type of each atom vector in the atom tensor of the first predicted molecular structure as a corresponding first type prediction vector, and extract the atom identifier and the atomic coordinate of each atom vector as a corresponding first coordinate prediction vector; and form a corresponding first type prediction tensor from all the obtained first type prediction vectors, and form a corresponding first coordinate prediction tensor from all the obtained first coordinate prediction vectors; and extract the chemical bond identifier and the chemical bond length of each chemical bond vector in the chemical bond tensor of the first predicted molecular structure to form a corresponding first bond length prediction vector, and form a corresponding first bond length prediction tensor from all the obtained first bond length prediction vectors; Step 606: The obtained first type prediction tensor and the corresponding first type label tensor form a corresponding first prediction-label pair; the obtained first coordinate prediction tensor and the corresponding first coordinate label tensor form a corresponding second prediction-label pair; and the obtained first bond length prediction tensor and the corresponding first bond length label tensor form a corresponding third prediction-label pair; Step 607: Substitute the first, second, and third prediction-label pairs into the preset first, second, and third loss functions to calculate corresponding first, second, and third loss values; and perform a weighted sum calculation on the first, second, and third loss values and use the calculated result as the corresponding first total loss value; The first loss function includes at least an L1 loss function, an L2 loss function, and a multi-classification cross entropy loss function; the second loss function includes at least an L1 loss function, an L2 loss function, and an MSE loss function; the third loss function includes at least an L1 loss function, an L2 loss function, and an MSE loss function; Step 608, whether the first loss value meets the preset first loss value range is identified; if it meets, it is turned to step 609; if it does not meet, a round of parameter optimization is carried out on the structure optimization model based on the preset first model parameter optimizer towards the direction of making the first loss function reach the minimum value, and it is returned to step 604 at the end of this round of parameter optimization; Wherein, the first model parameter optimizer at least includes SGD optimizer, ADAM optimizer; Step 609, whether the second loss value meets the preset second loss value range is identified; if it meets, it is turned to step 610; if it does not meet, a round of parameter optimization is carried out on the structure optimization model based on the preset second model parameter optimizer towards the direction of making the second loss function reach the minimum value, and it is returned to step 604 at the end of this round of parameter optimization; Wherein, the second model parameter optimizer at least includes SGD optimizer, ADAM optimizer; Step 610, whether the third loss value meets the preset third loss value range is identified; if it meets, it is turned to step 611; if it does not meet, a round of parameter optimization is carried out on the structure optimization model based on the preset third model parameter optimizer towards the direction of making the third loss function reach the minimum value, and it is returned to step 604 at the end of this round of parameter optimization; Wherein, the third model parameter optimizer at least includes SGD optimizer, ADAM optimizer; Step 611, whether the first total loss value meets the preset fourth loss value range is identified; if it meets, it is turned to step 612; if it does not meet, a round of parameter optimization is carried out on the structure optimization model based on the preset fourth model parameter optimizer towards the direction of making the first total loss value reach the minimum value, and it is returned to step 604 at the end of this round of parameter optimization; Wherein, the fourth model parameter optimizer at least includes SGD optimizer, ADAM optimizer; Step 612, whether the current test record is the last first data record in the first training set is identified; if it is, it is turned to step 613; if it is not, the next first data record in the first training set is taken as a new current test record and it is returned to step 603; Step 613, count the total number of the first data records of the first evaluation set to obtain a corresponding record total number N; and traverse all the first data records of the first evaluation set; and when traversing, take the first data record currently traversed as a corresponding current evaluation record; and input the first training molecular structure of the current evaluation record into the structure optimization model to obtain a corresponding first sample predicted structure x p,i , 1≤index i≤N; and take the first label molecular structure of the current evaluation record as a corresponding first sample label structure x t,i ; and form a corresponding first sample predicted-label pair composed of the first sample predicted structure x p,i and the first sample label structure x t,i corresponding to the current evaluation record; Step 614, when the traversal of all the first data records of the first evaluation set ends, all the first sample prediction-label pairs obtained are brought into the preset first evaluation function to calculate the corresponding first evaluation value; Wherein, the first evaluation function is realized based on RMSD function, specifically: Step 615, whether the first evaluation value meets the preset first evaluation value range is identified; if it does not meet, it is returned to step 601 to continue training; if it meets, it stops to continue training and confirms that the pre-training ends.
7. The method for optimizing the structure of electrolyte formulation according to claim 2, characterized in that: The model tuning of the structure optimization model based on the second data set specifically includes: Step 701, the second data set is randomly divided into two data sets based on the preset second training evaluation ratio, which are recorded as corresponding second training set and second evaluation set; The second training set and the second evaluation set both include a plurality of second data records; the ratio of the total number of records in the second training set to the total number of records in the second evaluation set meets the second training evaluation ratio; Step 702, taking the first second data record in the second training set as a corresponding current test record; Step 703, extracting the atom identifier and the atom element type of each atom vector in the atom tensor of the second label molecular structure of the current test record to form a corresponding second type label vector, and extracting the atom identifier and the atom coordinates of each atom vector to form a corresponding second coordinate label vector; and forming a corresponding second type label tensor from all the second type label vectors obtained, and forming a corresponding second coordinate label tensor from all the second coordinate label vectors obtained; and extracting the bond identifier and the bond bond length of each bond vector in the bond tensor of the second label molecular structure to form a corresponding second bond length label vector, and forming a corresponding second bond length label tensor from all the second bond length label vectors obtained; Step 704, inputting the second training molecular structure of the current test record into the structure optimization model to obtain a corresponding second predicted molecular structure through three-dimensional structure optimization processing; The second predicted molecular structure includes a corresponding atom tensor and a bond tensor; Step 705, extracting the atom identifier and the atom element type of each atom vector in the atom tensor of the second predicted molecular structure as a corresponding second type predicted vector, and extracting the atom identifier and the atom coordinates of each atom vector as a corresponding second coordinate predicted vector; and forming a corresponding second type predicted tensor from all the second type predicted vectors obtained, and forming a corresponding second coordinate predicted tensor from all the second coordinate predicted vectors obtained; and extracting the bond identifier and the bond bond length of each bond vector in the bond tensor of the second predicted molecular structure to form a corresponding second bond length predicted vector, and forming a corresponding second bond length predicted tensor from all the second bond length predicted vectors obtained; Step 706, forming a corresponding fourth prediction-label pair from the second type predicted tensor obtained and the corresponding second type label tensor; and forming a corresponding fifth prediction-label pair from the second coordinate predicted tensor obtained and the corresponding second coordinate label tensor; and forming a corresponding sixth prediction-label pair from the second bond length predicted tensor obtained and the corresponding second bond length label tensor; Step 707, inputting the fourth, fifth and sixth prediction-label pairs into the preset fourth, fifth and sixth loss functions to obtain corresponding fourth, fifth and sixth loss values; and performing weighted summation calculation on the fourth, fifth and sixth loss values and taking the calculation result as a corresponding second total loss value; The fourth loss function at least includes an L1 loss function, an L2 loss function and a multi-class cross-entropy loss function; the fifth loss function at least includes an L1 loss function, an L2 loss function and an MSE loss function; and the sixth loss function at least includes an L1 loss function, an L2 loss function and an MSE loss function. In step 708, whether the fourth loss value meets a preset fifth loss value range is identified; if yes, step 709 is performed; if no, a round of parameter fine-tuning is performed on the structure optimization model in a direction of minimizing the fourth loss function based on a preset fifth model parameter optimizer, and step 704 is returned when the round of parameter fine-tuning ends. The fifth model parameter optimizer at least includes an RMSprop optimizer, an ADAM optimizer and an AdamW optimizer. In step 709, whether the fifth loss value meets a preset sixth loss value range is identified; if yes, step 710 is performed; if no, a round of parameter fine-tuning is performed on the structure optimization model in a direction of minimizing the fifth loss function based on a preset sixth model parameter optimizer, and step 704 is returned when the round of parameter fine-tuning ends. The sixth model parameter optimizer at least includes an RMSprop optimizer, an ADAM optimizer and an AdamW optimizer. In step 710, whether the sixth loss value meets a preset seventh loss value range is identified; if yes, step 711 is performed; if no, a round of parameter fine-tuning is performed on the structure optimization model in a direction of minimizing the sixth loss function based on a preset seventh model parameter optimizer, and step 704 is returned when the round of parameter fine-tuning ends. The seventh model parameter optimizer at least includes an RMSprop optimizer, an ADAM optimizer and an AdamW optimizer. In step 711, whether the second total loss value meets a preset eighth loss value range is identified; if yes, step 712 is performed; if no, a round of parameter fine-tuning is performed on the structure optimization model in a direction of minimizing the second total loss value based on a preset eighth model parameter optimizer, and step 704 is returned when the round of parameter fine-tuning ends. The eighth model parameter optimizer at least includes an RMSprop optimizer, an ADAM optimizer and an AdamW optimizer. In step 712, whether the current test record is the last second data record in the second training set is identified; if yes, step 713 is performed; if no, the next second data record in the second training set is taken as a new current test record, and step 703 is returned. Step 713, count the total number of the second data records of the second evaluation set to obtain a corresponding record total number M; and traverse all the second data records of the second evaluation set; and when traversing, take the second data record currently traversed as a corresponding current evaluation record; and input the second training molecular structure of the current evaluation record into the structure optimization model to obtain a corresponding second sample predicted structure x p,j , 1≤index j≤M; and take the second label molecular structure of the current evaluation record as a corresponding second sample label structure x t,j ; and form a corresponding second sample predicted-label pair composed of the second sample predicted structure x p,j and the second sample label structure x t,j corresponding to the current evaluation record; In step 714, when the traversal of all the second data records in the second evaluation set ends, all the second sample prediction-label pairs obtained are taken into a preset second evaluation function to obtain a corresponding second evaluation value. The second evaluation function is implemented based on an RMSD function, and specifically is: Step 715, identify whether the second evaluation value meets the preset second evaluation value range; if not, return to step 701 for continuous training; if yes, stop continuous training and confirm that the model optimization is completed.
8. An apparatus for performing the process of structurally optimizing components of an electrolyte formulation according to any one of claims 1-7, characterized in that, The device comprises a model construction module, a data preparation module, a model training module and a model application module. The model construction module is configured to construct a structure optimization model; the structure optimization model is realized based on a Uni-Mol model; the structure optimization model is configured to perform three-dimensional structure optimization processing on the original molecular structure input by the model and output the corresponding optimized molecular structure. The data preparation module is configured to construct a corresponding first data set through big data acquisition; and construct a corresponding second data set through molecular dynamics simulation. The model training module is configured to pre-train the structure optimization model based on the first data set; and after the pre-training is completed, model optimization is performed on the structure optimization model based on the second data set. The model application module is configured to, after the model optimization is completed, input the electrolyte formula input by the user as a corresponding current electrolyte formula; and use a preset chemoinformatics tool to perform corresponding three-dimensional molecular structure modeling according to the molecular sequence of each component of the current electrolyte formula to obtain a corresponding component modeling structure; and perform three-dimensional structure optimization processing on each component modeling structure based on the structure optimization model to obtain a corresponding component optimized structure; and each component molecular sequence, the corresponding component type and the component optimized structure form a corresponding formula component structure optimization record; and all the obtained formula component structure optimization records form a corresponding formula structure optimization report to feedback to the user; the current electrolyte formula comprises a plurality of formula components; each formula component comprises a component type and a component molecular sequence; the component type comprises a solvent, an electrolyte and an additive; the component molecular sequence is a one-dimensional SMILES sequence of a solvent molecule, an electrolyte molecule or an additive molecule corresponding to the component type; and the chemoinformatics tool at least comprises OpenBabel, RDK i t.
9. An electronic device, comprising: Comprise: a memory, a processor and a transceiver; the processor is configured to be coupled with the memory, read and execute the instructions in the memory to realize the method in any one of claims 1-7; the transceiver is coupled with the processor, and the transceiver is controlled by the processor to perform message transmission and reception.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, when the computer instructions are executed by the computer, the computer executes the method in any one of claims 1-7. The computer readable storage medium stores computer instructions, when the computer instructions are executed by the computer, the computer executes the method in any one of claims 1-7.
Citation Information
Patent Citations
Method and device for processing polycyclic conjugated system molecule generation model
CN118866165A