Method for optimizing electrolyte formula by utilizing artificial intelligence

By pre-processing and feature engineering of the electrolyte formula data, combined with optimization algorithms, the deviation problems caused by data errors and missing in the existing technology are solved, and efficient optimization of the electrolyte formula is achieved, which significantly improves the development efficiency and result accuracy.

CN119943176AActive Publication Date: 2025-05-06WUHAN INSTITUTES OF ADVANCED TECHNOLOGY CHINESE ACADEMY OF SCIENCES
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411946958.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-06
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

When using artificial intelligence to optimize electrolyte formulas, the prior art faces the deviation problems caused by data errors and missing, and the model training efficiency is low, resulting in low results accuracy and efficiency.

Method used

By collecting electrolyte formula data from laboratory and Internet databases, data preprocessing, feature engineering and model training are carried out, including data cleaning, missing value filling, feature selection and model evaluation, and global search is carried out in combination with optimization algorithms to output the optimized electrolyte formula.

Benefits of technology

The abnormal data is effectively removed, the accuracy and completeness of the data is improved, and more accurate complex interactions between the electrolyte components are obtained, which significantly improves the efficiency of electrolyte formulation development and reduces the number of experiments and time costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943176A_ABST
    Figure CN119943176A_ABST
Patent Text Reader

Abstract

According to the method, data are well processed, abnormal data can be effectively removed, a large number of effective data sources are obtained, then a proper machine algorithm is selected for machine model training through the large number of processed effective data, and then the obtained model is effectively corrected by optimizing an obtained model correction program. The complex interaction among the electrolyte components can be more accurately captured, and the formula of the electrolyte can be better optimized. In addition, the method provided by the invention can also greatly improve the development efficiency of the electrolyte formula, more excellent electrolyte formula can be provided more quickly according to actual application scene requirements, and the experiment frequency and time cost can be remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of battery electrolyte preparation, and in particular to a method for optimizing electrolyte formula by using artificial intelligence. Background Art

[0002] As a key component of the battery, the formula of the electrolyte has a significant impact on the battery performance (such as capacity, cycle life, safety, etc.). The traditional research and development of electrolyte formulas relies on a lot of experiments and experience accumulation, which is not only inefficient but also costly. Using artificial intelligence to optimize the electrolyte formula can obviously greatly improve efficiency and reduce R&D costs. In recent years, as the research on artificial intelligence has made certain breakthroughs, some technologies have emerged that use artificial intelligence to optimize material formulas. However, there are still some problems with this type of technology.

[0003] For example, the use of artificial intelligence optimization requires the use of a large amount of data to train the model. Internet databases can provide a large amount of training data. Generally speaking, the larger the amount of training data, the more conducive it is to obtain a more accurate model. However, there are some erroneous data in the Internet data, which cannot be well identified and eliminated by the machine, which will cause deviations in the fitting process and reduce accuracy. Manual processing can increase accuracy, but it will reduce efficiency and increase complexity. In the industry, in order to reduce the workload of data processing, a small amount of data is usually selected to train the model, which will lead to low accuracy of the model obtained through training, increase the verification process, and then lead to a longer running process and poor final results. In addition to the deviation caused by erroneous data, data loss is also a problem that interferes with the fitting results. Traditional practices will abandon this part of the data. However, this method has a large loss for niche fields with less data. In addition, artificial intelligence is still in the research and development stage in the field of material formulation optimization due to reasons such as model selection, training, calibration and unreasonable program design, and cannot be well applied in actual production.

[0004] The electrolyte formula involves multiple components during the development process. The type and addition amount of each component have a great influence on the performance of the electrolyte, and also bring certain difficulties to the development of the electrolyte formula. Therefore, the use of artificial intelligence to assist in the development of electrolyte formula has great prospects. However, the existing artificial intelligence cannot be well applied in the optimization of electrolyte formula. There is still a certain gap between the obtained results and the actual application effect, and there are still few existing technologies for optimizing electrolyte formulas using artificial intelligence. Therefore, the technology of optimizing electrolyte formulas using artificial intelligence still needs a lot of development for a long time in the future. In view of this, the present invention provides a new method for optimizing electrolyte formulas using artificial intelligence. Summary of the invention

[0005] The present invention aims to provide an efficient method for optimizing electrolyte formulation using artificial intelligence to improve R&D efficiency and optimization effect.

[0006] The specific technical solutions of the present invention are as follows:

[0007] In the first aspect, an efficient method for optimizing electrolyte formulation using artificial intelligence is provided, comprising:

[0008] Step 1, data collection and screening: Collect the electrolyte formula and its corresponding performance data in the laboratory, remove outliers and data with obvious errors, and obtain data 1; collect the electrolyte formula and corresponding performance data from the Internet database to obtain data 2;

[0009] Step 2, data preprocessing: use data 1 to clean data 2 and supplement missing values ​​to obtain data 3, and standardize the obtained data 3 to a unified numerical range to eliminate dimensional differences;

[0010] Step 3, feature engineering: extract multiple features including conventional types, concentrations, and physical and chemical properties of each component in data 3 electrolyte, introduce descriptors based on molecular structure and features obtained through quantum chemical calculations, extract these features and their corresponding properties to form a feature library, and use feature selection and construction techniques to automatically determine the most relevant and representative feature combination;

[0011] Step 4: Artificial intelligence algorithm selection and model training: Try multiple machine learning algorithms and select the appropriate machine learning algorithm by comparing the evaluation results through the cross-validation method; use part of data 3 as the training set to train the model and use the remaining data 3 as the validation set to validate the model;

[0012] Step 5: Model evaluation and optimization: Use the validation set to evaluate the model obtained in step 4 to determine whether the model is overfitting or underfitting; if there is overfitting or underfitting, calibrate the obtained model until the obtained model is no longer overfitting or underfitting;

[0013] Step 6, combined with optimization algorithm: input the electrolyte formula constraint conditions, automatically combine the feature library to use the optimization algorithm to perform a global search on the electrolyte formula, output the optimized electrolyte formula, and finally experimentally verify the obtained electrolyte formula.

[0014] Furthermore, the Internet database in step 1 includes public databases of research institutions, academic journal databases, and databases formed by their combination;

[0015] Furthermore, the amount of data 1 in step 1 is not limited here. Within a certain range, the larger the amount of data, the better the effect of processing data 2. However, in order to reduce the workload, the amount of database 1 should not be too large, and can be selected to be 10-30% of data 2.

[0016] Furthermore, the specific method of cleaning data 2 by inputting data 1 in step 2 is as follows:

[0017] (1) Data fusion: Integrate data 1 and data 2 to compare the same features and indicators;

[0018] (2) Anomaly detection: using statistical methods or machine learning algorithms to identify outliers in the Internet database that have a large distribution difference from the laboratory data;

[0019] (3) Data filling: Missing values ​​in the Internet database can be filled using appropriate interpolation methods based on the characteristics and relationships of laboratory data.

[0020] Furthermore, the data fusion in step (1) is specifically as follows:

[0021] 1) Data alignment: ensuring that data from different sources are consistent in terms of measurement units and feature definitions;

[0022] 2) Feature matching: identifying features in two data sources that represent the same concept or attribute and associating them;

[0023] 3) Data merging: merging the aligned and matched laboratory data with the data in the Internet database;

[0024] 4) Data standardization and normalization: standardize or normalize the merged data to eliminate data differences between different data sources;

[0025] 5) Duplicate data processing: Check and process possible duplicate data records to ensure the uniqueness and accuracy of the data.

[0026] After the above processing, the final data 3 includes data 1. In addition, through data fusion, the advantages of data from different sources can be fully utilized to provide a richer and more comprehensive information basis for subsequent data cleaning, analysis and modeling.

[0027] Furthermore, the machine learning algorithm described in step (2) includes one of a Z-score algorithm, an IQR (interquartile range) method, an isolation forest algorithm, a local outlier factor (LOF) algorithm, and a One-Class SVM (one-class support vector machine); this type of algorithm can better identify outliers.

[0028] Furthermore, the interpolation method in step (3) includes one of mean interpolation, median interpolation, mode interpolation, regression interpolation, multiple interpolation, and nearest neighbor interpolation.

[0029] Furthermore, the descriptor in step 3 includes at least one of a topological descriptor, a geometric descriptor, and an electronic descriptor.

[0030] Molecular structure-based descriptors are numerical representations used to quantitatively describe the molecular structure characteristics of each component of the electrolyte. These descriptors can capture information such as the size, shape, topology, and electronic properties of each component molecule.

[0031] Furthermore, the quantum calculation described in step 3 is implemented through Gaussian, ORCA, and VASP, and the calculation results can be properly normalized or standardized.

[0032] In the optimization of electrolyte formula, quantum chemical calculations can be used to obtain characteristics such as energy and charge distribution related to the molecules of each component of the electrolyte, providing valuable information for feature engineering.

[0033] Furthermore, the specific method for determining the most relevant and representative feature combination may include at least one of a filtering method, a wrapping method, an embedded method, and a clustering-based method;

[0034] The filter methods of the present invention can calculate the correlation between each feature (such as the type, amount, structure, molecular weight, etc. of each substance in the electrolyte) and the target variable (such as electrolyte performance: cycle performance, capacity, rate performance, flame retardant performance, etc.), such as Pearson correlation coefficient, mutual information, etc. A threshold is set according to the correlation score, and features with scores higher than the threshold are selected as relevant features;

[0035] Wrapper Methods of the present invention: Use machine learning models (such as random forests, support vector machines) to recursively add or delete features (such as the type, amount, structure, molecular weight, etc. of each substance in the electrolyte) to evaluate the electrolyte performance corresponding to different feature combinations. For example, starting from an empty feature set, gradually add features, and evaluate the model performance after each addition until the electrolyte performance no longer improves;

[0036] The embedded methods of the present invention: automatically select key features during model training. For example, Lasso regression can achieve feature selection by compressing key feature coefficients and compressing some coefficients to zero.

[0037] The clustering method of the present invention is to cluster the features (such as the type, amount, structure, molecular weight, etc. of each substance in the electrolyte) and group the features with similar features. Then, a representative feature is selected from each cluster for combination to reduce the number of key features and maintain the diversity of information.

[0038] Furthermore, the data used as the training set in step 4 can be randomly selected from 75-90% of the data in data 3, and the remaining data can be used as the validation set; or data 1 can be selected as the validation set, and the corresponding remaining part of data 3 can be used as the training set;

[0039] Furthermore, the machine learning algorithm described in step 4 includes random forest, support vector machine, neural network or convolutional neural network; the machine learning algorithm in this step is used for prediction and modeling, and pays more attention to learning and fitting of data.

[0040] Furthermore, when evaluating the model in step 4, if overfitting is found, regularization technology can be used for adjustment; if underfitting exists, feature engineering can be re-examined or the amount of training data can be increased; according to actual needs, when the above methods cannot solve the problem better or faster, such as overfitting and underfitting occur more than 4 times in total, the algorithm selection can be re-examined and other algorithms can be selected for retraining.

[0041] Furthermore, in step 5, the obtained model is evaluated by calculating at least two of the accuracy, recall, and mean square error indicators.

[0042] Furthermore, when the accuracy of the training set is above 95%, the recall rate is above 90%, or the mean square error is below 0.05, and the accuracy of the validation set is below 75%, the recall rate is below 70%, or the mean square error is above 0.15, it indicates overfitting;

[0043] Furthermore, when the accuracy of the training set is less than 75%, the recall rate is less than 60%, or the mean square error is more than 0.25, and the accuracy of the validation set is more than 80%, the recall rate is more than 80%, or the mean square error is less than 0.2, it indicates underfitting.

[0044] Furthermore, the optimization algorithm described in step 6 includes a genetic algorithm, a simulated annealing algorithm, and a particle swarm optimization algorithm; the optimization algorithm in this step is selected to search for the optimal solution according to the constraints of the problem and the optimization goal.

[0045] When there are many features affecting the performance of the electrolyte to be analyzed and the solution space of the problem is very large and complex, the genetic algorithm is given priority. When there is a greater probability of jumping out of the local performance optimum of the electrolyte during the optimization process, the simulated annealing algorithm is given priority. When it is hoped to quickly obtain a better solution and there are fewer features, the particle swarm optimization algorithm is given priority.

[0046] Furthermore, the electrolyte formula constraint conditions in step 6 are the electrolyte formula to be optimized or the partial electrolyte formula and the electrolyte target performance parameters, or directly the electrolyte target performance parameters.

[0047] Further, the optimized electrolyte formula outputted in step 6 is one or more;

[0048] The number of constraints input in step 6 of the present invention affects the number of optimized electrolyte formulas output; generally, the more constraints there are, the fewer the output results are; therefore, when the output results of the electrolyte formula are large and it brings inconvenience to the test, constraints can be added to reduce the output results.

[0049] Furthermore, the electrolyte formula includes but is not limited to a lithium ion battery electrolyte formula and a sodium ion battery electrolyte formula.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] 1. The method of the present invention can effectively remove abnormal data and fill missing values ​​to obtain a large number of valid data sources through data processing, and then use the processed large amount of valid data for model training, which is conducive to more accurately capturing the complex interactions between electrolyte components and facilitating the acquisition of a better training model;

[0052] 2. The method of the present invention selects, trains and corrects the model well, and reasonably selects and sets the program for optimizing the electrolyte formula by artificial intelligence, so that the final model can better match the development of the electrolyte, and the obtained program can better achieve the optimization of the electrolyte formula;

[0053] 3. The method of the present invention uses artificial intelligence to assist in the development of electrolyte formulas, greatly improving the efficiency of electrolyte formula development and significantly reducing the number of experiments and time costs. In addition, the method of the present invention also takes into account production process and cost constraints to ensure that the optimized formula has practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 A flow chart of the present invention using artificial intelligence to optimize the electrolyte formula;

[0055] Figure 2 This is a test diagram of the discharge specific capacity of the electrolyte in Example 1. DETAILED DESCRIPTION

[0056] The technical solution of the present application will be clearly and completely described below in conjunction with the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0057] The present invention provides a method for optimizing electrolyte formula using artificial intelligence, and its flow chart is shown in Figure 1 , the method comprising:

[0058] Step 1, data collection and screening: Collect the electrolyte formula and its corresponding performance data in the laboratory, remove outliers and data with obvious errors, and obtain data 1; collect the electrolyte formula and corresponding performance data from the Internet database to obtain data 2;

[0059] Step 2, data preprocessing: use data 1 to clean data 2 and supplement missing values ​​to obtain data 3, and standardize the obtained data 3 to a unified numerical range to eliminate dimensional differences;

[0060] Step 3, feature engineering: extract multiple features including conventional types, concentrations, and physical and chemical properties of each component in data 3 electrolyte, introduce descriptors based on molecular structure and features obtained through quantum chemical calculations, extract these features and their corresponding properties to form a feature library, and use feature selection and construction techniques to automatically determine the most relevant and representative feature combination;

[0061] Step 4: Artificial intelligence algorithm selection and model training: Try multiple machine learning algorithms and select the appropriate machine learning algorithm by comparing the evaluation results through the cross-validation method; use part of data 3 as the training set to train the model and obtain the model, while retaining the remaining data of data 3 as the validation set to validate the model;

[0062] Step 5: Model evaluation and optimization: Use the validation set to evaluate the model obtained in step 4 to determine whether the model is overfitting or underfitting; if there is overfitting or underfitting, calibrate the obtained model until the obtained model is no longer overfitting or underfitting;

[0063] Step 6, combined with optimization algorithm: input the electrolyte formula constraint conditions, automatically combine the feature library to use the optimization algorithm to perform a global search on the electrolyte formula, output the optimized electrolyte formula, and finally experimentally verify the obtained electrolyte formula.

[0064] The Internet database described in step 1 of this embodiment includes public databases of research institutions, academic journal databases, and databases formed by their combination.

[0065] The performance data corresponding to the electrolyte formula collected by the laboratory in step 1 of this embodiment must be authentic and valid to ensure that the results used to process data 2 have a better effect; for example, when collecting this part of the data, data from authoritative laboratory sources can be selected; and based on the evaluation of this part of the data by technicians with senior experience in this field, data with obvious problems are eliminated.

[0066] The amount of data 1 in step 1 of this embodiment can be selected to be 10-30% of data 2, for example, 15%, 20%, 25% of the amount of data 2. Within a certain range, the larger the amount of data 1, the better the effect of data 2, the fewer subsequent corrections, and the higher the operating efficiency; but in order to reduce the workload of manual screening, the amount of data 1 should not be too large; of course, the amount of data 1 should not be too small. When data 1 is too small, data 2 cannot be cleaned according to data 1, and data 2 may be over-cleaned, resulting in large data loss, which may also lead to large deviation in the results.

[0067] The specific method of cleaning data 2 by using data 1 in step 2 of this embodiment is as follows:

[0068] (1) Data fusion: Integrate data 1 and data 2 to compare the same features and indicators;

[0069] (2) Anomaly detection: using statistical methods or machine learning algorithms to identify outliers in the Internet database that have a large distribution difference from the laboratory data;

[0070] (3) Data filling: Missing values ​​in the Internet database can be filled using appropriate interpolation methods based on the characteristics and relationships of laboratory data;

[0071] Specifically, the data fusion in step (1) is as follows:

[0072] 1) Data alignment: ensure that data from different sources are consistent in terms of measurement units and feature definitions; 2) Feature matching: determine the features that represent the same concept or attribute in the two data sources and associate them; 3) Data merging: merge the aligned and matched laboratory data with the data in the Internet database; 4) Data standardization and normalization: standardize or normalize the merged data to eliminate data differences between different data sources; 5) Duplicate data processing: check and process possible duplicate data records to ensure the uniqueness and accuracy of the data; Through data fusion, the present invention can make full use of the advantages of data from different sources and provide a richer and more comprehensive information basis for subsequent data cleaning, analysis and modeling.

[0073] More specifically, the outliers described in (2) are identified by one of the following methods: a Z-score algorithm, an IQR (interquartile range) method, an isolation forest algorithm, a local outlier factor (LOF) algorithm, or a One-ClassSVM (one-class support vector machine);

[0074] More specifically, the interpolation methods described in (3) include mean interpolation, median interpolation, mode interpolation, regression interpolation, multiple interpolation, and nearest neighbor interpolation.

[0075] As an example, it can be listed that: for the same electrolyte formula, data 1 shows that in the same battery, the 100-cycle retention rate is 85%, in data 2, data a shows that the 100-cycle retention rate is between 70-90%, and data b shows that the 100-cycle retention rate is below 70% or above 90%, then the system marks data a as a normal value, marks data b as an abnormal value and removes it from data 2; it is almost impossible to occur, but when all the results of data 2 appear or a larger part shows that the 100-cycle retention rate is below 70% or above 90%, then the system marks data 2 as a normal value; the normal value is the cleaned data, which is used to continue the process of subsequent steps.

[0076] As an example, it can be listed that: Data 1 shows that the 100-cycle retention rate of the electrolyte containing 0.8-1.5 mol / L lithium bis(fluorosulfonyl)imide is 85-90%, and the 500-cycle retention rate is 80-85%. The concentration of lithium bis(fluorosulfonyl)imide in the electrolyte with the same formula as in Data 1 in Data 2 is also 1.0 mol / L, but its performance test data only shows that the 100-cycle retention rate is 85-90%, and the 500-cycle retention rate is not measured. In order to better obtain the relationship between long-cycle stability and lithium salt or lithium bis(fluorosulfonyl)imide concentration in subsequent feature engineering and increase the weight of valid data, the system selects the middle value of the 500-cycle retention rate in Data 1 (i.e., 82.5%) to fill in the missing values ​​in Data 2 according to the result of the 500-cycle retention rate in Data 1.

[0077] The specific method of determining the most relevant and representative feature combination in this embodiment may include: at least one of a filtering method, a wrapping method, an embedded method, and a clustering-based method;

[0078] Specifically, filter methods can calculate the correlation between each feature (such as the type, amount, structure, molecular weight, etc. of each substance in the electrolyte) and the target variable (such as electrolyte performance: cycle performance, capacity, rate performance, flame retardant performance, etc.), such as Pearson correlation coefficient, mutual information, etc., and set a threshold based on the correlation score to filter out features with scores higher than the threshold as relevant features;

[0079] Specifically, wrapper methods use machine learning models (such as random forests and support vector machines) to evaluate the electrolyte performance corresponding to different feature combinations by recursively adding or deleting features (such as the type, amount, structure, molecular weight, etc. of each substance in the electrolyte). For example, starting from an empty feature set, gradually adding features, and evaluating the model performance after each addition until the electrolyte performance no longer improves;

[0080] Specifically, the embedded methods of the present invention automatically select key features during the model training process. For example, Lasso regression can achieve feature selection by compressing key feature coefficients and compressing some coefficients to zero.

[0081] Specifically, the clustering method is used to cluster the features (such as the type, amount, structure, molecular weight, etc. of each substance in the electrolyte) and group the features with similar features together. Then, a representative feature is selected from each cluster for combination to reduce the number of key features and maintain the diversity of information.

[0082] The molecular structure-based descriptor in step 3 of this embodiment includes at least one of a topological descriptor, a geometric descriptor, and an electronic descriptor. The molecular structure-based descriptor is a numerical representation used to quantitatively describe the characteristics of a molecular structure. These descriptors can capture information about the size, shape, topological structure, electronic properties, etc. of a molecule.

[0083] The quantum calculation described in step 3 of this embodiment is implemented by at least one of Gaussian, ORCA, and VASP, and the calculation results can be properly normalized or standardized. In the optimization of electrolyte formula, the energy, charge distribution and other characteristics related to electrolyte molecules can be obtained through quantum chemical calculation, providing valuable information for feature engineering.

[0084] Specifically, Gaussian is widely used to calculate various properties of molecules, such as molecular orbitals, energy, charge distribution, vibration frequency, etc. By inputting the structure and related parameters of each component in the electrolyte, Gaussian can perform complex calculations and simulations based on the principles of quantum mechanics, and provide predictions and analyses of thermodynamic properties, the difficulty and rate of reactions, etc.; ORCA can better consider the influence of the solvent environment on molecular properties through quantum chemical calculations, and more realistically simulate the molecular behavior under actual reaction conditions. It can also provide an evaluation of the kinetics of the reactions of each component in the electrolyte, the stability of the molecules, and the thermodynamics of the reactions; VASP is an ab initio quantum mechanics calculation program based on density functional theory, which can provide performance predictions of the mechanical properties (elastic constants, hardness, etc.), thermal properties (thermal conductivity, heat capacity, etc.), and optical properties of each component of the electrolyte.

[0085] The machine learning algorithm described in step 4 of this embodiment includes random forest, support vector machine, neural network or convolutional neural network; the selection of the machine learning algorithm described in step 4 of the present invention can be combined with the results of feature engineering;

[0086] As an example, it can be listed that: when the amount of data 3 is small (for example, the amount of data is within 1000) and the feature combination is relatively simple (for example, the number of features is within 8), the system gives priority to support vector machines, which is conducive to quickly obtaining results; as an example, it can be listed that: if the amount of data 3 is large (for example, greater than 5000), the feature combination is complex (for example, the number of features is more than 16) and there is a certain internal structure, neural networks or convolutional neural networks are given priority, which can automatically learn complex patterns from large amounts of data and are more conducive to obtaining more accurate results.

[0087] In step 5 of this embodiment, the obtained model is evaluated by calculating at least two of the accuracy, recall, and mean square error indicators.

[0088] Specifically, when the accuracy of the training set reaches above 95%, the recall rate is above 90%, or the mean square error is below 0.05, and the accuracy of the validation set is lower than 75%, the recall rate is lower than 70%, or the mean square error is higher than 0.15, it indicates overfitting; when the accuracy of the training set is lower than 75%, the recall rate is lower than 60%, or the mean square error is higher than 0.25, and the accuracy of the validation set reaches above 80%, the recall rate is higher than 80%, or the mean square error is lower than 0.2, it indicates underfitting.

[0089] When verifying and evaluating the obtained model in step 5 of this embodiment, if overfitting is found, regularization technology can be used for adjustment; the implementation of regularization technology is relatively direct and applicable to many types of models. In addition to regularization technology, other methods can also be selected for adjustment; as an example, it is listed: increasing the amount of data 1, which can enrich the learning content of the model and reduce overfitting of existing data; as an example, it is listed: removing some unimportant or redundant key factors in the feature combination, thereby reducing the complexity of the model. If it is underfitting, re-examine the feature engineering in step 3 or increase the amount of training set data; according to actual needs, when the above method cannot solve the problem better or faster, such as overfitting and underfitting occur more than 4 times in total, you can also reselect the machine learning algorithm in step 4 and retrain.

[0090] In step 6 of this embodiment, random disturbances and diversified initial solutions may be introduced during the optimization process to avoid falling into a local optimum.

[0091] Specifically, random perturbation refers to applying some random changes to the current solution or search direction during the operation of the algorithm. For example, in a certain iteration step, the proportion or parameter value of a certain component in the electrolyte formula is slightly randomly changed. This is conducive to giving the algorithm the opportunity to explore other areas near the current solution, rather than just being limited to the current local optimal solution. Diversified initial solutions are when the algorithm starts, not just starting from a fixed initial point, but generating multiple different, random initial electrolyte formulas. This can increase the starting point of the algorithm search, because different initial solutions may lead the algorithm to different search paths. Local optimality means that in the entire solution space, the solution in a certain area is better than the solution in its neighboring area, but there may be better solutions in the entire solution space. If the algorithm always starts from the same initial solution and there is not enough randomness and diversity in the search process, it is easy to be trapped in such a local optimal solution and cannot find the global optimal solution. Therefore, by introducing random perturbations and diversified initial solutions, the possibility of the algorithm jumping out of the local optimal solution and finding a better global solution can be increased.

[0092] The electrolyte formula constraint condition in step 6 of this embodiment is the electrolyte formula to be optimized or the partial electrolyte formula and the electrolyte target performance, or the electrolyte target performance;

[0093] As an example, it can be listed that: when the electrolyte formula to be optimized and the target performance parameters are input, the system substitutes the corresponding features of the electrolyte formula to be optimized into the obtained model to calculate the deviation, and the optimization algorithm preferentially adjusts the electrolyte formula around the electrolyte formula to be optimized according to the target performance parameters and in combination with the features of the feature library and the corresponding relationship between the performance, and finally outputs the electrolyte formula that meets the model. The optimized electrolyte formula may be one or more;

[0094] As an example, it can be listed that: when the target performance parameters of the electrolyte are input, the optimization algorithm selects the components and their corresponding features in combination with the features of the feature library and the performance correspondence, and substitutes them into the obtained model to calculate until the output is the electrolyte formula that satisfies the model. The obtained electrolyte formula may be one or more;

[0095] As an example, it can be listed that: when the partial components of the electrolyte and the target performance parameters of the electrolyte are input, the system substitutes the target performance parameters and the corresponding features of the partial components of the electrolyte into the model for calculation, and the optimization algorithm combines the features of the feature library and the performance correspondence to give priority to selecting the partial components of the electrolyte as fixed components and continue to select components around the partial components of the electrolyte and substitute them into the model to calculate the electrolyte formula that meets the model. The resulting electrolyte formula may be one or more.

[0096] The optimization algorithm in step 6 of this embodiment includes a genetic algorithm, a simulated annealing algorithm, and a particle swarm optimization algorithm;

[0097] Specifically, when there are many parameters and the solution space of the problem is very large and complex, the genetic algorithm is given priority; when there is a greater probability of jumping out of the local optimum during the optimization process, the simulated annealing algorithm can be set to select; when it is hoped to quickly obtain a better solution and there are fewer parameters, the particle swarm optimization algorithm can be selected; the complex electrolyte formula of the present invention involves a variety of solutes and solvents. The types and proportions, each component has a variety of optional concentration ranges. Different combinations of these components and concentrations constitute the solution to the problem. When there are many electrolyte components and each component has a wide range of options, the number of possible combinations will increase exponentially, forming a huge solution space; in addition, the interaction between these combinations and the impact on the final performance is not a simple and direct relationship. There may be a variety of nonlinear and interactive effects, making it difficult to determine the optimal solution through intuitive analysis or simple calculations. For example, a slight change in the concentration of a solute in the electrolyte may have a complex and unpredictable impact on multiple performance indicators of the electrolyte due to interactions with other solutes and solvents. This makes it extremely challenging to find the optimal or near-optimal electrolyte formula in such a huge solution space. The genetic algorithm has strong global search capabilities, parallel processing capabilities, and the ability to handle complex relationships. It does not require derivative information of the problem and has good flexibility and adaptability, which makes it uniquely advantageous in dealing with problems with large and complex solution spaces. For complex problems such as electrolyte formulation optimization, there may be many local optimal solutions in the solution space. For example, although the electrolyte formulation satisfies the model as a whole, it has excellent high temperature resistance but poor cycle performance. Traditional optimization algorithms may be easily trapped in a local optimal solution. The simulated annealing algorithm increases the possibility of exploring new areas by accepting poor solutions in a timely manner during the optimization process, making it more likely to get rid of the local optimality of electrolyte performance and find a better global solution or a solution close to the global optimal electrolyte formulation. When the number of features or parameters in the feature combination of step 3 is small and there are certain requirements for the solution speed, the particle swarm optimization algorithm can achieve fast calculation and quickly give the optimization results.

[0098] The technical solution of the present invention is further described below in conjunction with specific embodiments.

[0099] Example 1 Lithium-ion battery electrolyte formulation optimization:

[0100] Step 1, data collection and screening: 300 sets of lithium-ion battery electrolyte formulas and their corresponding capacity, cycle performance, and flame retardant performance data were collected from the laboratory. These data were reviewed, and 28 sets of data with obvious errors were eliminated to obtain 272 sets of valid data (data 1); then 3000 sets of lithium-ion battery electrolyte formulas and their corresponding capacity, cycle performance, and flame retardant performance data were obtained from the Internet database (data 2);

[0101] Step 2, data preprocessing: make the measurement units and feature definitions of data 1 and data 2 the same, associate the features of the same concept or attribute, merge data 1 and data 2, and then perform normalization, compare the same features and indicators, use One-Class SVM to identify 17 groups of outliers in data 2 that are significantly different from those in data 1, and remove the outliers in data 2; according to the characteristics and relationships of data 1, use the mean interpolation method to fill in the missing values, and finally achieve 41 missing values ​​in data 2 to obtain data 3, and standardize the data to the [0,1] interval;

[0102] Step 3, feature engineering: the types, addition amounts, ionic conductivity and decomposition voltage characteristics of organic solvents, lithium salts and additives in the electrolyte were extracted. At the same time, the molecular size and shape characteristics obtained by the geometric descriptor and the thermodynamic properties, reaction difficulty and rate characteristics obtained by quantum chemical calculations by Gaussian, as well as the corresponding capacity, cycle performance, flame retardant performance and the differences in capacity, cycle performance and flame retardant performance corresponding to different characteristics were used to form a feature library corresponding to the characteristics and performance. The correlation measure between the characteristics and the target performance was calculated by the filtering method to determine the feature selection and construction, and the 9 most relevant and representative features (solvent type and concentration, lithium salt concentration and addition amount, additive type and concentration, electrolyte thermodynamic parameters, ionic conductivity and decomposition voltage) were determined.

[0103] Step 4: Selection of artificial intelligence algorithm and model training: Try three machine learning algorithms: random forest, support vector machine and neural network. Compare and evaluate the results through 10-fold cross validation. Finally, select the random forest algorithm as the artificial intelligence algorithm. Randomly extract 85% of data 3 for model training, and keep the remaining 15% of data 3 as the validation set.

[0104] Step 5, model evaluation and optimization: The trained random forest model is evaluated using the 15% validation set mentioned above. The training set calculation accuracy is 95%, the recall rate is 91%, and the mean square error is 0.05. It is found that the model is slightly overfitted, and L2 regularization is used for adjustment; the model is evaluated again, and it is found that the model calculation accuracy is 88%, the recall rate is 87%, and the mean square error is 0.08, which meets the requirements;

[0105] Step 6: Combined with the optimization algorithm: Use the simulated annealing algorithm to input the electrolyte formula to be optimized and the 25°C discharge capacity (10 cycles of 0.2C discharge capacity greater than 220mAh g -1 ), cycle performance (500 cycle retention rate is above 75%), and flame retardant performance (self-extinguishing time is within 5s) as constraints, a global search was conducted on the electrolyte formula, and 13 optimized lithium-ion electrolyte formulas were obtained.

[0106] Randomly select one of the above 13 electrolyte formulas to perform performance tests and comparisons with the electrolyte before optimization under the same conditions. The self-extinguishing time of the electrolyte before optimization was measured to be 4s, and the 500-cycle retention rate was 71%, while the self-extinguishing time after optimization was 5s, and the 500-cycle retention rate was 78%. Figure 2 It can be seen that the capacity of the optimized electrolyte formula after 10 cycles at 0.2C is greater than 220mAh g -1 .

[0107] It can be seen from the above results that the method for optimizing the electrolyte formula using artificial intelligence provided by the present invention can optimize the electrolyte formula well, so that the optimized electrolyte formula meets the set requirements in terms of capacity, 500-cycle cycle retention rate, and flame retardant performance, and the overall performance is significantly improved compared to before optimization.

Claims

1. A method for optimizing electrolyte formula using artificial intelligence, characterized in that: include: Step 1, data collection and screening: collect the electrolyte formula and its corresponding performance data in the laboratory, remove outliers and data with obvious errors, and obtain data 1; Collect electrolyte formulas and corresponding performance data from the Internet database to obtain data 2; Step 2, data preprocessing: use data 1 to clean data 2 and supplement missing values ​​to obtain data 3, and standardize the obtained data 3 to a unified numerical range to eliminate dimensional differences; Step 3, feature engineering: extract multiple features including conventional types, concentrations, and physical and chemical properties of each component in data 3 electrolyte, introduce descriptors based on molecular structure and features obtained through quantum chemical calculations, extract these features and their corresponding properties to form a feature library, and use feature selection and construction techniques to automatically determine the most relevant and representative feature combination; Step 4: Artificial intelligence algorithm selection and model training: Try multiple machine learning algorithms and select the appropriate machine learning algorithm by comparing the evaluation results through the cross-validation method; use part of data 3 as the training set to train the model and use the remaining data 3 as the validation set to validate the obtained model; Step 5: Model evaluation and optimization: Use the validation set to evaluate the model obtained in step 4 to determine whether the model is overfitting or underfitting; if there is overfitting or underfitting, calibrate the model until the model is no longer overfitting or underfitting; Step 6, combined with optimization algorithm: input the electrolyte formula constraint conditions, automatically combine the feature library to use the optimization algorithm to perform a global search on the electrolyte formula, output the optimized electrolyte formula, and finally experimentally verify the obtained electrolyte formula.

2. The method for optimizing electrolyte formula using artificial intelligence according to claim 1, characterized in that: The specific methods in step 2 include: (1) Data fusion: Integrate data 1 and data 2 to compare the same features and indicators; (2) Anomaly detection: Use statistical methods or machine learning algorithms to identify outliers in data 2 that are significantly different from the distribution of data 1; (3) Data filling: For missing values ​​of data 2, they can be filled using appropriate interpolation methods based on the characteristics and relationships of data 1.

3. The method for optimizing electrolyte formula using artificial intelligence according to claim 1, characterized in that: The training set in step 4 may be 75-90% of the randomly selected data 3, and the remaining data 3 may be used as a validation set; or data 1 may be selected as a validation set, and the corresponding remaining data 3 may be used as a training set.

4. The method for optimizing electrolyte formula using artificial intelligence according to claim 1, characterized in that: The specific method for determining the feature combination in step 3 includes at least one of a filtering method, a wrapping method, an embedded method, and a clustering-based method.

5. The method for optimizing electrolyte formula using artificial intelligence according to claim 1, characterized in that: The descriptor in step 3 includes at least one of a topological descriptor, a geometric descriptor, and an electronic descriptor; the quantum computing in step 3 is at least one of Gaussian, ORCA, and VASP.

6. The method for optimizing electrolyte formula using artificial intelligence according to claim 1, characterized in that: The machine learning algorithm in step 4 includes but is not limited to at least one of a random forest, a support vector machine, a neural network or a convolutional neural network.

7. The method for optimizing electrolyte formula using artificial intelligence according to claim 1, characterized in that: When evaluating the model in step 4, if overfitting is found, regularization technology can be used to correct it; if underfitting is found, feature engineering should be reviewed or the amount of training set data should be increased for correction.

8. The method for optimizing electrolyte formula using artificial intelligence according to claim 1, characterized in that: In step 5, the obtained model is evaluated by calculating at least two of the precision, recall, and mean square error indicators.

9. The method for optimizing electrolyte formula using artificial intelligence according to claim 1, characterized in that: The optimization algorithm in step 6 includes but is not limited to at least one of a genetic algorithm, a simulated annealing algorithm, and a particle swarm optimization algorithm.

10. The method for optimizing electrolyte formula using artificial intelligence according to claim 1, characterized in that: The electrolyte formula constraint conditions in step 6 are the electrolyte formula to be optimized or the partial electrolyte formula and the target performance of the electrolyte, or the target performance of the electrolyte; the electrolyte formula includes but is not limited to one of the lithium ion battery electrolyte formula and the sodium ion battery electrolyte formula.

Citation Information

Patent Citations

  • Method for screening and optimizing methanation nickel-based catalyst formula based on ANN-NSGA-II

    CN112131785A

  • Electrolyte formula recommendation method based on lithium battery performance simulation environment

    CN113312807A

  • Electrolyte performance data determination method and device based on machine learning

    CN118136147A

  • Specific electrolyte formula research method based on neural network

    CN118230845A

  • Method and system for identifying electrolyte composition for optimal battery performance

    US20230420757A1