A method, system, and equipment for predicting the performance and verifying the synthesis of organic cathode materials for lithium batteries.

CN122575593APending Publication Date: 2026-08-14NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

但该方法仅针对含能材料体系,未涵盖将有机分子应用于正极所涉及的核心电化学指标;CN114818948A公开了一种图神经网络的数据-机理驱动的材料属性预测方法,利用采用特征工程和描述符的组合进行预测,但该方法描述符很难获得全面分子信息,且组合过程为直接拼接,缺少优化过程,方法预测精度差;CN120808952B公开了一种融合三维结构与先验特征的分子性质预测方法和装置,其采用了分子特征和先验特征融合的方法,但新增特征需人工验证合理性,难以适配快速演进的材料研发

Benefits of technology

[0016]本发明的有益技术效果:基于Transformer架构的Uni-Mol模型,将三维分子深层隐式特征与关键分子描述符子集进行特定权重的特征拼接与融合,实现关键性能的高精度预测,大幅降低了研发成本,为锂电池新型有机正极材料的快速迭代提供了技术手段。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575593A_ABST
    Figure CN122575593A_ABST
Patent Text Reader

Abstract

This invention discloses a method, system, and device for predicting the performance and verifying the synthesis of organic cathode materials for lithium batteries. The steps include: constructing a multi-physicochemical performance prediction model based on a Uni-Mol backbone feature extraction network with a Transformer architecture; performing high-throughput batch predictions using the multi-physicochemical performance prediction model, and then selecting target organic cathode material molecules through multi-constraint screening; constructing an inverse synthesis prediction model based on message passing neural network and global reactive attention model training, and determining the optimal synthesis route for the target organic cathode material molecules through the inverse synthesis prediction model. This invention achieves accurate screening and efficient synthesis design of organic cathode materials throughout the entire process, reduces R&D costs, and accelerates the iteration of novel organic cathode materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial process control and optimization technology, specifically to a method, system, and equipment for predicting the performance of organic cathode materials for lithium batteries and verifying their synthesis. Background Technology

[0002] Lithium-ion batteries, as the mainstream electrochemical energy storage devices, are widely used in consumer electronics, new energy vehicles, and large-scale energy storage. Their core performance and cost are highly dependent on cathode materials. Although traditional inorganic transition metal oxide cathodes have been commercialized, they face inherent bottlenecks such as resource constraints, high costs, and thermal safety risks. Organic electrode materials, with their abundant elemental reserves, environmental friendliness, strong molecular structure designability, and multi-electron transfer characteristics, have become highly promising candidate systems for novel cathode materials. However, the current research and development and preparation of organic cathode materials still heavily rely on the empirical model of "design first - trial and error synthesis - electrochemical testing," resulting in a technical bottleneck in practical applications where performance prediction and synthesis verification are mutually exclusive.

[0003] In the area of ​​molecular performance prediction and high-throughput screening, CN110728047B discloses a computer-aided design system for energetic molecules based on machine learning performance prediction. This method generates candidate molecules by combining benzene parent rings with substituents, and combines molecular descriptor calculation with supervised learning algorithms to achieve rapid prediction of physicochemical properties such as density and detonation velocity. However, this method only applies to energetic material systems and does not cover the core electrochemical indicators involved in applying organic molecules to cathodes. CN114818948A discloses a data-mechanism-driven material property prediction method using graph neural networks, which uses a combination of feature engineering and descriptors for prediction. However, the descriptors in this method are difficult to obtain comprehensive molecular information, and the combination process is a direct splicing without optimization, resulting in poor prediction accuracy. CN120808952B discloses a molecular property prediction method and device that integrates three-dimensional structure and prior features. It adopts a method of fusing molecular features and prior features, but the rationality of newly added features needs to be manually verified, making it difficult to adapt to the rapidly evolving material development.

[0004] In summary, the performance prediction methods in this field are insufficient and cannot meet the needs of large-scale and rapid R&D of organic cathode materials. At the same time, there is a lack of practical methods for the entire process from molecular construction to final experimental synthesis and preparation verification. The R&D of organic cathode materials still relies heavily on manual trial and error, molecular screening is highly blind, and the failure rate during engineering transformation is high, which seriously restricts its engineering application. Summary of the Invention

[0005] To address the shortcomings of the prior art, this invention provides, in a first aspect, a method for predicting the performance of organic cathode materials for lithium batteries, and verifies the reverse synthesis of target organic cathode materials selected by this method through experimental preparation. The method specifically includes the following steps:

[0006] S1: Construct a multi-source organic energy storage material dataset, extract a subset of key molecular descriptors from the molecules in the dataset using a molecular descriptor extraction tool, and extract the three-dimensional deep implicit features of the molecules in the dataset using a backbone feature extraction network. S2: The deep implicit features of three-dimensional molecules are spliced ​​and fused with a subset of key molecular descriptors to generate a comprehensive molecular characterization vector. A multi-task prediction head is constructed with the comprehensive molecular characterization vector as input, and multiple key physicochemical properties are output to obtain a multi-physicochemical property prediction model. S3: Based on a multi-physicochemical performance prediction model, high-throughput prediction is performed on the constructed organic cathode material molecules to obtain multi-dimensional physicochemical performance parameters of the molecules. Based on preset electrochemical characteristic thresholds, step-by-step judgment and multi-condition joint constraint screening are performed to obtain the target organic cathode material molecules.

[0007] Furthermore, in step S1, the specific process includes: A multi-source organic energy storage material dataset is constructed, which includes a dedicated dataset for organic energy storage materials and a commercial pre-trained dataset. The dedicated dataset for organic energy storage materials contains molecular physicochemical property labels constructed from first-principles calculations and experimental tests. Molecular descriptor extraction tools were used to quantitatively characterize the molecular structures in the dedicated dataset, obtaining initial high-dimensional molecular descriptors including topological features, electronic features, thermodynamic features, geometric conformation features, charge features, hydrogen bond features, and aromaticity features. Based on SHAP interpretability analysis, the initial high-dimensional molecular descriptors were ranked and screened according to their contribution to obtain a subset of key molecular descriptors. The Uni-Mol model based on the Transformer architecture of equivariant spatial coding was used as the backbone feature extraction network for model pre-training and fine-tuning to obtain the deep implicit features of three-dimensional molecules.

[0008] Furthermore, in step S2, the specific process includes: performing missing value imputation and Z-score standardization on the subset of key molecular descriptors to unify the numerical format with the three-dimensional molecular deep implicit features; concatenating the three-dimensional molecular deep implicit features and the subset of key molecular descriptors after linear weighting according to the optimal weights to generate a comprehensive molecular representation vector, wherein the optimal weight α of the three-dimensional molecular deep implicit features is 0.66 during linear weighted fusion; constructing a multi-task prediction head with the comprehensive molecular representation vector; and performing supervised training with a joint loss function to obtain a multi-physicochemical performance prediction model.

[0009] Furthermore, the multi-physicochemical performance prediction model outputs multiple key physicochemical properties, including thermodynamic characteristics, electronic characteristics, and polarity characteristics, through a multi-task prediction head. The key physicochemical properties include at least: enthalpy of formation, molecular frontier orbital energy level HOMO / LUMO, melting point, boiling point, dipole moment, and molecular electrostatic potential.

[0010] Furthermore, in the pre-training stage: based on unlabeled molecules on a commercial pre-training dataset, unsupervised pre-training is performed through self-supervised tasks of masked atom prediction and 3D coordinate recovery to learn general molecular structure rules and obtain pre-training weights; Fine-tuning stage: Load pre-trained weights and perform supervised fine-tuning on a dataset specifically for organic energy storage materials to output three-dimensional molecular deep implicit features.

[0011] In step S3, preset electrochemical characteristic thresholds are used for step-by-step judgment and multi-condition joint constraint screening. The specific screening rules include: Based on the structural stability requirements of battery cycling, the molecular formation enthalpy is controlled within the preset thermodynamic stability range of -200 to -800 kJ per mole; Based on the electrolyte compatibility requirements, the HOMO and LUMO energy levels are controlled within the range of -5.2 to -6.0 electron volts, which is compatible with the lithium battery electrolyte and operating potential, relative to the vacuum energy level. According to the requirements of electrode processing and working temperature, the melting point and boiling point are controlled within the preset high temperature range, with the melting point greater than 180 degrees Celsius, the boiling point greater than 350 degrees Celsius, or the decomposition temperature greater than 250 degrees Celsius. Based on the redox activity requirements, the dipole moment was controlled within the range of 15-25 Debye, which characterizes high specific capacity. According to the solubility control requirements, the molecular electrostatic potential is controlled by a polar exponent based on the normalization of the variance of the molecular surface electrostatic potential, with a preset range of 0.15–0.35 and the difference between the extreme values ​​of positive and negative electrostatic potentials ΔV is less than 1.2 electron volts. During the screening process, molecules that do not meet any of the above performance threshold conditions are sequentially eliminated, and molecules that meet all performance constraints are retained as target organic cathode material molecules.

[0012] Furthermore, S4: Construct a standardized reaction dataset containing precise atomic mappings, extract local reaction templates, and train and construct an inverse synthesis prediction model based on a graph neural network architecture of message passing neural network and global reactive attention. S5: Input the structural information of the target organic cathode material molecule obtained in step S4 into the retrosynthesis prediction model, use the retrosynthesis prediction model to output the probability score of the local reaction template of the target organic cathode material molecule to complete the single-step retrosynthesis prediction, and then perform multi-step iterative search to generate multiple candidate synthesis routes. Select the candidate route with the highest score as the optimal synthesis route; prepare the target organic cathode material molecule based on the optimal synthesis route and perform performance testing. Feed the experimental characterization and test data back to the model training dataset to perform closed-loop iterative optimization of the retrosynthesis prediction model.

[0013] Furthermore, in step S4, the specific process includes: Construction of standardized reaction dataset: Single-step organic reaction data were obtained from the Reaxys database, and the data was cleaned and standardized. Atom mapping tools were used to clarify the atomic correspondence between reactants and products, locate the reaction center, and remove low-quality reactions to obtain a standardized reaction dataset with accurate atom mapping. Local reaction template extraction: For each reaction in the standardized reaction dataset, the structures of the mapped reactants and products are compared to locate the changes in atoms and chemical bonds at the reaction center, and local reaction templates are obtained. Local reaction templates are classified into atomic reaction templates, bond reaction templates and multiple change reaction templates according to the change type. Construction and training of the retrosynthesis prediction model: The product molecule is represented as an atomic-bond topological graph structure for feature initialization; a message-passing neural network is used to iteratively update the atomic features, and local reaction center features are extracted by fusing neighborhood chemical environment information and reconstructing and enhancing bond features; a multi-head global reactive attention mechanism is introduced to capture global structural information and output globally perceived atomic and bond features; using the globally perceived atomic and bond features as inputs, atomic template classifiers and bond template classifiers are constructed respectively, and the probability scores of each atom and each chemical bond corresponding to various local reaction templates are output; the cross-entropy loss function is used for training to obtain the retrosynthesis prediction model.

[0014] Secondly, this application provides a prediction system, including a first module, which can receive constructed organic cathode material molecules and output the key physicochemical properties of the organic cathode material molecules through a multi-physicochemical performance prediction model in the first module; a second module, which receives the key physicochemical properties of the organic cathode material molecules output by the first module and can screen target organic cathode material molecules by setting constraints in an interactive manner; and a third module, which receives the target organic cathode material molecules screened by the second module and performs single-step or multi-step optimal synthesis path planning through a retrosynthetic prediction model in the third module.

[0015] Thirdly, this application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores the prediction system of the second aspect, the prediction system being executable by the at least one processor.

[0016] The beneficial technical effects of this invention are as follows: Based on the Uni-Mol model of the Transformer architecture, the deep implicit features of three-dimensional molecules and the subset of key molecular descriptors are spliced ​​and fused with specific weights to achieve high-precision prediction of key performance, which greatly reduces the R&D cost and provides a technical means for the rapid iteration of new organic cathode materials for lithium batteries.

[0017] By using model building and experience in the field for molecular screening, a practical method for the entire process of organic cathode materials, namely "batch construction by molecular computer - performance prediction - synthesis optimization - experimental verification", has been established. Through model optimization and empirical design, a closed-loop optimization link for the entire process has been established, which can accelerate the large-scale and rapid research and development of organic cathode materials.

[0018] The retrosynthetic prediction model based on MPNN and global reactive attention effectively overcomes the bias in identifying local reaction centers, significantly improves the accuracy and feasibility of synthetic route planning, realizes the automated design of synthetic routes for target organic cathode materials and the precise optimization of reaction conditions. The selected synthetic routes have fewer steps, readily available raw materials, mild reaction conditions, significantly reduced synthesis costs, good process repeatability, and are easy to scale up for production. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the overall technical roadmap for the performance prediction and synthesis verification method of the organic cathode material for lithium batteries of the present invention. Figure 2 Ranking of feature importance based on SHAP value analysis; Figure 3 Distribution of SHAP values ​​for key molecule descriptors; Figure 4 To improve the accuracy of enthalpy prediction for the formation of organic cathode materials; Figure 5 Accuracy of orbital energy level prediction for organic cathode materials; Figure 6 To improve the accuracy of boiling point prediction for organic cathode materials; Figure 7 To improve the accuracy of melting point prediction for organic cathode materials; Figure 8 This is a scatter plot of the fitted model on the test set; Figure 9 A scatter plot showing the fit of the uni-mol model on the test set only; Figure 10 A scatter plot of the fit of the descriptor-only model on the test set; Figure 11 The organic cathode material prepared; Figure 12 XRD pattern of organic cathode material; Figure 13 FTIR spectrum of organic cathode material; Figure 14 Raman spectra of organic cathode materials; Figure 15 The charge-discharge curves are for organic cathode materials. Detailed Implementation

[0021] To facilitate understanding of the present invention, the present invention will be described more fully and in detail below with reference to the accompanying drawings and preferred embodiments, but the scope of protection of the present invention is not limited to the following specific embodiments.

[0022] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by those skilled in the art. The technical terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the invention.

[0023] Unless otherwise specified, all raw materials, reagents, instruments and equipment used in this invention can be purchased from the market or prepared by existing methods.

[0024] Example: This example obtains an organic cathode material for lithium batteries through a method for predicting the performance of organic cathode materials for lithium batteries and verifying their synthesis. The method also prepares the material and measures its physicochemical properties to verify the effectiveness of the method.

[0025] like Figure 1 This is a general technical roadmap for predicting the performance of organic cathode materials and verifying their synthesis. The method for predicting the performance of organic cathode materials includes the following steps: The S1 multi-physicochemical performance prediction model is constructed as follows: organic molecular structure and performance data from various sources are collected to build an initial dataset. Multi-dimensional molecular descriptors are constructed using various molecular descriptor extraction tools. Feature selection and dimensionality reduction are performed using algorithms to retain highly relevant molecular descriptors. These descriptors are then fused and spliced ​​with deep implicit features extracted by the pre-trained backbone feature extraction network to obtain a comprehensive molecular representation vector. Finally, a multi-task fully connected prediction head is constructed using the comprehensive molecular representation vector to predict multiple key physicochemical properties of organic cathode materials.

[0026] S1.1 Construct a dataset of multi-source organic energy storage materials; A dedicated dataset for organic energy storage materials includes experimental data on reported nitrogen-containing heterocyclic, quinone, and nitro-based organic electrode materials, calculated using first-principles calculations. After deduplication, SMILES standardization, and removal of invalid data, over 100,000 labeled, finely tuned data points were obtained. Labels cover molecular enthalpy of formation, HOMO / LUMO levels, melting point, boiling point, dipole moment, molecular electrostatic potential, and theoretical specific capacity. All molecular SMILES were standardized using RDKit 2023.09.1 ​​and divided into training, validation, and test sets in an 8:1:1 ratio.

[0027] Commercial pre-training dataset: 220 million unlabeled three-dimensional conformational data of organic molecules from the ZINC15 database were used as the base data for pre-training the Uni-Mol model.

[0028] S1.2 Molecular descriptor extraction and feature selection; Using toolkits such as RDKit, Mordred, PaDel-descriptor, and Dragon, over 7000 initial molecular descriptors were generated for prediction targets including enthalpy, HOMO / LUMO energy levels, melting point, boiling point, dipole moment, and molecular electrostatic potential. These descriptors encompassed several major categories, including topological features, electronic features, thermodynamic features, geometric conformation features, charge features, hydrogen bond features, and aromaticity features, achieving a comprehensive quantitative characterization of molecular structure information. Then, near-zero variance redundant features were eliminated using a variance thresholding method, leaving over 5000 effective descriptors. Based on dimensionality reduction using correlation analysis, and further based on SHAP interpretability analysis, the contribution of descriptors for each prediction target was ranked, and a subset of 32 key molecular descriptors with a cumulative contribution of over 95% were selected. Figure 2 The study presents a ranking of feature importance based on SHAP value analysis. Among these, mAccs_41 (MACCS molecular fingerprint feature) has the greatest average impact on the model output, followed by topological and molecular surface area-related descriptors such as SMR_VSA2 and mAccs_134. Figure 3The distribution of SHAP values ​​for key molecular descriptors is shown, where features such as maccs_41, SMR_VSA2, and maccs_134 have a significant positive impact on model output when their values ​​are high (red dots). Therefore, the optimized molecular descriptors, while retaining core structural information, reduce feature dimensions and improve model training efficiency and prediction accuracy.

[0029] S1.3 Construction of a prediction model based on Uni-Mol and feature fusion; Three-dimensional molecular deep representation learning based on Uni-Mol: Uni-Mol is used as the backbone feature extraction network, with the core being a Transformer architecture with SE (3) equivariant spatial encoding. Based on the Transformer architecture and introducing three-dimensional spatial position encoding and pairwise representation, it can efficiently capture long-range interactions and three-dimensional geometric information between atoms. The model takes atom type and atom three-dimensional coordinates as input, maintains atom representation and pairwise representation, and realizes bidirectional interaction between atom and pairwise information through attention mechanism, automatically learning deep implicit features such as molecular conjugation system, electron cloud distribution, spatial conformation, and bond interaction.

[0030] The model performs two-stage training; Pre-training phase: On the unlabeled organic molecule dataset of the ZINC15 database, general molecular structure rules are learned through self-supervised tasks of masked atom prediction and 3D coordinate recovery. Specifically, unsupervised pre-training is performed on the ZINC15 dataset, with a 6-layer Transformer encoder, 8-head attention, atom representation dimension of 512, and pair representation dimension of 128. General molecular structure rules are learned through self-supervised tasks of masked atom prediction and 3D coordinate recovery. Pre-training is performed for 50 rounds with a batch size of 128 and the optimizer is AdamW.

[0031] Fine-tuning stage: Pre-trained weights are loaded and supervised fine-tuning is performed on a dedicated dataset to output 128-dimensional molecular deep structural features. The 128-dimensional deep implicit features are intelligently concatenated and fused with a subset of 32-dimensional key molecular descriptors to form a 160-dimensional comprehensive representation vector. This comprehensive molecular representation vector takes into account both data-driven implicit features and knowledge-driven explicit features, further improving the accuracy and generalization ability of joint prediction of multiple physicochemical properties. The specific intelligent concatenation and fusion is as follows: Based on the Uni-Mol fine-tuning output of 128-dimensional fixed-dimensional three-dimensional molecular deep implicit features, missing value imputation and Z-score standardization are performed on the subset of 32-dimensional key molecular descriptors after SHAP screening to unify the numerical format and distribution of the two types of features. A linear weighted fusion mode is adopted, firstly concatenating the 128-dimensional deep implicit features and the subset of 32-dimensional key descriptors directly in the feature dimension to obtain a 160-dimensional basic fusion feature. This feature is only used as the scoring benchmark for model performance comparison and ablation experiments. The optimal weight α = 0.66 was determined through a grid search on the validation set, and a linear weighted fusion model was constructed: y^ = Concat(0.66y^ 3D 0.34y^ desc ); where y^ 3D For a 128-dimensional deep implicit feature vector, y^ desc The model is a 32-dimensional subset vector of key descriptors. `Concat(·)` represents the concatenation operation along the feature dimensions, resulting in a 160-dimensional, robust, comprehensive molecular characterization vector that balances 3D spatial conformation and 2D topology. This comprehensive molecular characterization vector is input into a multi-task fully connected prediction head, which outputs several key physicochemical properties, including: enthalpy of formation, HOMO level, LUMO level, melting point, boiling point, dipole moment, and molecular electrostatic potential—key performance parameters for screening organic cathode materials. The model employs a reasonable loss function and optimization strategy, using MSE as the joint loss function, 10-fold cross-validation to optimize hyperparameters, fine-tuning for 30 rounds, and an early stopping patience value of 5 to prevent overfitting.

[0032] Reference Figures 4 to 5 Red, blue, and pink represent the optimal value of the test set, the optimal value of the validation set, and the current test value, respectively. Figure 4 The distribution of model accuracy for predicting the enthalpy of formation of organic cathode materials under different hyperparameter combinations is shown, with the highest R² reaching 90.58%. Figure 5 The distribution of prediction accuracy of the model for the orbital energy levels of organic cathode materials under different hyperparameter combinations is shown, with the highest R² being 98.42%.

[0033] Reference Figures 6 to 7 , Figure 6 The distribution of the model's predicted values ​​for the boiling point of organic cathode materials is shown, with R² = 0.9875. Figure 7The model's predicted values ​​for the melting point of organic cathode materials are compared with experimental values, with R² = 0.9628. The dense distribution of data points along the diagonal indicates that the model has excellent prediction accuracy.

[0034] like Figure 8 The results show the model's prediction error assessment for molecular electrostatic potential, especially ESPminEV, the minimum electrostatic potential, on the test set, with a mean absolute error (MAE) of 0.0813 and a coefficient of determination (R²) of 0.7167.

[0035] The effectiveness of the fusion model of deep structural features and molecular descriptors in this application is verified by two comparative examples. Comparative example 1: ESPminEV property prediction method based solely on three-dimensional molecular representation model; The core difference between this comparative model and the present invention is that it only uses a three-dimensional molecular representation model to complete the regression prediction of ESPminEV, without introducing two-dimensional topological features such as molecular descriptors and molecular fingerprints, and without performing multi-model weighted fusion operations. The rest of the dataset division, experimental environment, and evaluation indicators are completely consistent with the present invention.

[0036] The specific implementation process of this comparative comparison is as follows: First, the original dataset is randomly divided at the sample level into training, validation, and test sets in an 8:1:1 ratio. The sample distribution of each dataset after division is completely consistent with that of this invention, ensuring fairness in the comparison. Then, the molecular SMILES expression is converted into molecular objects and hydrogen supplementation is performed. A conformation generation algorithm is used to generate corresponding three-dimensional atomic coordinates for each molecule, while simultaneously filtering out invalid samples that fail to generate conformations. Based on this, a three-dimensional molecular representation model is constructed. The model is initialized using pre-trained parameters, with atomic type and three-dimensional spatial coordinates as model inputs, and the prediction error of ESPminEV as the optimization objective. Supervised regression fine-tuning is performed on the training set. The hyperparameter settings and optimization strategies during training are consistent with those of this invention. Figure 9 After the model training was completed, the generalization performance was evaluated on the test set. The final result showed that the mean absolute error (MAE) of the model on the test set was 0.0883, and the coefficient of determination (R²) was 0.6826.

[0037] Compared to the weighted fusion model of this invention, this comparative model relies solely on the three-dimensional conformation information of molecules for prediction. It cannot effectively utilize key information such as the presence of functional groups and the distribution of substructures in the two-dimensional topology of molecules. As a result, the prediction error of ESPminEV is higher, and the ability to explain data variance is weaker. In particular, for molecular samples containing special functional groups, significant prediction bias is likely to occur. The overall prediction accuracy and robustness are clearly limited.

[0038] Comparative Example 2: ESPminEV property prediction method based solely on molecular descriptors combined with traditional ensemble tree regression; The core difference between this comparative example and the present invention is that it only uses descriptors and fingerprint features based on two-dimensional molecular structures, combined with a traditional ensemble tree regression model to complete ESPminEV prediction. It does not introduce three-dimensional molecular representation features, nor does it perform multi-model weighted fusion operations. The rest of the dataset partitioning, feature preprocessing specifications, and evaluation metrics are completely consistent with the present invention and Comparative Example 1.

[0039] The specific implementation process of this comparative example is as follows: A dataset partitioning method completely consistent with this invention is adopted to obtain a training set, a validation set, and a test set, avoiding performance interference caused by differences in data partitioning. Using the two-dimensional topological structure of the molecule as input, multi-source features are extracted using the RDKit toolkit. The feature set contains three categories: first, continuous physicochemical descriptors such as molecular weight, topological index, surface area correlation, and partial charge distribution statistics; second, Morgan circular fingerprints used to characterize the occurrence patterns of local substructures; and third, MACCS structural bonds used to capture information about the presence of common functional groups and substructures. After feature extraction, a mean-based strategy is used to fill missing values ​​in the features. Continuous features are standardized to eliminate numerical instability caused by differences in feature scales, and finally, a unified feature vector is formed. Subsequently, an ensemble tree model is used as a regressor to perform model fitting on the training set, and hyperparameter optimization is performed on the validation set. No three-dimensional conformational information is introduced during training. Figure 10 After the model training was completed, the generalization performance was evaluated on the test set. The final result showed that the mean absolute error (MAE) of the model on the test set was 0.0855, and the coefficient of determination (R²) was 0.6867.

[0040] The predictive performance of this comparative example is slightly better than the three-dimensional molecular representation model of Comparative Example 1. However, compared with the weighted fusion model of this invention, this comparative example model still has significant performance shortcomings. This method relies only on the two-dimensional topological features of molecules and cannot explicitly model the influence of the three-dimensional spatial configuration of molecules and the spatial arrangement of atoms on the electrostatic potential distribution. It cannot capture the effect of long-range spatial effects of molecules on ESPminEV, and the prediction bias is significant for molecular samples with significant differences in spatial structure. Overall, the prediction accuracy and generalization ability are weaker than the weighted fusion scheme of this invention.

[0041] S2 is based on multi-attribute prediction and multi-level threshold screening to obtain target organic cathode material molecules; Based on the trained multi-physicochemical performance prediction model, high-throughput batch predictions are performed on the constructed virtual organic cathode material molecular library, outputting key physicochemical properties and electrochemical characteristics parameters for each molecule, such as formation enthalpy, HOMO energy level, LUMO energy level, melting point, boiling point, dipole moment, and molecular electrostatic potential. Based on the requirements of redox mechanism, thermodynamic stability, and electrolyte compatibility of organic cathode material molecules in lithium-ion battery cathodes, and combined with the conventional technical experience of those skilled in the art, reasonable screening thresholds are preset for each of the above predicted performance parameters. Molecular screening is conducted using a step-by-step judgment and multi-condition joint constraint approach. The specific screening rules are as follows: Enthalpy of formation screening: The molecular enthalpy of formation is controlled within the preset thermodynamic stability range (-200~-800kJ / mol) to ensure that the molecular structure of the target organic cathode material is stable and not easily decomposed, thus meeting the structural stability requirements for long-term cycling of the battery.

[0042] HOMO / LUMO energy level screening: The HOMO and LUMO energy levels are controlled within a reasonable range (-5.2~-6.0eV (vs.vacuum)) that matches the lithium battery electrolyte and operating potential, so that the molecules have suitable redox activity, which can achieve efficient multi-electron energy storage without causing electrolyte decomposition.

[0043] Melting and boiling point screening: The melting and boiling points are controlled within the higher range that meets the electrode processing and battery operating temperature (melting point controlled at >180°C, boiling point controlled at >350°C (or decomposition temperature Td>250°C)) to ensure that the target organic cathode material molecules do not melt or volatilize under coating, drying, charging and discharging conditions, and maintain the solid structure and electrode integrity.

[0044] Dipole moment screening: The dipole moment is controlled within a reasonable range (15~25 D) that reflects good redox activity and specific capacity, so that the molecules have sufficient electron gain and loss capabilities, ensuring that the cathode material has high specific capacity and energy density.

[0045] Molecular electrostatic potential screening: to ensure that the molecular electrostatic potential meets the requirements of moderate polarity and uniform distribution, and to control the molecular polarity within a preset reasonable range (polarity index controlled at 0.15~0.35 (based on the normalization of the variance of molecular surface electrostatic potential), and the difference between the extreme values ​​of positive and negative electrostatic potential ΔV<1.2eV), thereby reducing the dissolution tendency of active molecules in the electrolyte and improving the cycle stability of the battery.

[0046] During the screening process, molecules that do not meet any threshold conditions are successively eliminated, and only organic cathode material molecules that meet all performance constraints are retained. Finally, the target organic cathode material molecule with high electrochemical performance, high thermal stability, good electrolyte compatibility and synthetic feasibility is obtained.

[0047] The method for synthesizing and verifying organic cathode materials includes the following steps: Design and optimization of retrosynthetic route for S3 target organic cathode material molecules; After obtaining the target organic cathode material molecule, the constructed retrosynthetic prediction model is used for single-step and multi-step path planning.

[0048] S3.1 Reaction Dataset Construction and Atom Mapping: Raw data extraction: Single-step organic reaction data related to organic compounds were extracted from the Reaxys database, focusing on commonly used reaction types in the synthesis of organic cathode materials such as nucleophilic substitution, condensation cyclization, nitration, coupling, and redox. A total of 620,000 raw reaction data were initially extracted.

[0049] Data cleaning and standardization: RDKit version 2023.09.1 ​​was used to standardize reaction smiles, and the following filtering operations were performed in sequence: invalid smiles, invalid reactions with the same reactant and product structures, complex reactions with multiple products, and reactions catalyzed by noble metals / heavy metals were removed; duplicate reactions, erroneous reactions with non-conservation of reaction atomic numbers, and reactions with incomplete reactant or product structures were removed; functional group representation format was standardized, valence errors were corrected, and solvent and non-reactive agent structures in the reaction were removed; after cleaning, 305,000 high-quality and valid single-step reaction data were obtained.

[0050] Atom mapping processing: The RXNMapper tool was used to perform atom mapping annotation on all cleaned reactions, clarifying the one-to-one correspondence between reactants and products, completing the precise location of reaction centers, and removing low-quality reactions with an atom mapping matching degree of less than 90%, finally obtaining a standardized reaction dataset of 305,000 records with precise atom mapping.

[0051] Dataset partitioning: The standardized reaction dataset is randomly divided into training, validation and test sets in a 90:5:5 ratio for model training, tuning and accuracy verification.

[0052] S3.2 Local reaction template extraction; Based on the reaction dataset obtained from atom mapping, local chemical transformation rules at the reaction centers are extracted, and three types of local reaction template sets are constructed. The specific process is as follows: For each reaction, by comparing the mapped reactant and product structures, the changes in atoms and chemical bonds at the reaction center are located, and templates are classified into three types according to the type of change: Atomic reaction template: Reactions in which only the type / valence state of atoms changes, without the breaking or formation of chemical bonds, such as deprotection and functional group modification reactions.

[0053] Bond reaction template: Reactions in which chemical bonds break or new chemical bonds are formed, such as carbon-carbon coupling, condensation cyclization, nitration, etc.

[0054] Multiple-change reaction template: complex reactions in which both atomic types and chemical bond connections change simultaneously.

[0055] S3.3 Model construction and training based on message passing neural network (MPNN) and global reactive attention (GRA); The inverse synthetic prediction model constructed in this embodiment adopts the LocalRetro model architecture, with its core being a graph neural network architecture consisting of a message-passing neural network (MPNN) and global reactive attention (GRA). The complete construction and training process is as follows: Molecular diagram feature initialization: The product molecule is represented as an atomic-bond topological diagram structure, and the atomic and bond features are initialized: Atomic features include atomic element type, atomicity, valence state, hybridization mode, formal charge, aromaticity, and ring affiliation, with a dimension of 64; Bond features include chemical bond type, bond order, conjugation, ring formation, and stereochemical information, with a dimension of 32.

[0056] Message Passing Neural Network (MPNN) Local Feature Learning: A 3-layer MPNN is used to iteratively update atomic features. Through message functions, aggregation functions, and update functions, each atomic feature is fully integrated with the information of the neighboring chemical environment to complete the feature extraction of the local reaction center. Based on the updated atomic features, the bond features are reconstructed and enhanced through fully connected layers to ensure that the bond features are consistent with the adjacent atomic environment.

[0057] Global Reactive Attention (GRA) Global Feature Enhancement: An 8-head multi-head self-attention mechanism is introduced to construct the GRA module, enabling each atom and each chemical bond to capture long-range interactions and global structural information within the molecule, and output globally perceived atomic and bond features, thus solving the prediction bias problem caused by focusing only on local reaction centers.

[0058] Template classifier construction: Atom template classifier and bond template classifier are constructed separately. The global enhanced atomic features and bond features are used as inputs, and the probability scores of each atom and each chemical bond corresponding to various local reaction templates are output respectively. The classifier adopts a 2-layer fully connected network structure with ReLU activation function and Softmax function for probability normalization in the output layer.

[0059] Model Training and Optimization: The loss function adopted is the cross-entropy loss function, and the total loss is the sum of the atomic template classification loss and the bond template classification loss; the optimizer adopted is the Adam optimizer, with an initial learning rate of 1e-4, a weight decay of 1e-6, and a batch size of 64; training strategy: the total number of training rounds is 50, and an early stopping strategy is adopted. Training is stopped when the validation set loss does not decrease for 6 consecutive rounds to prevent the model from overfitting; gradient clipping is used during training to avoid gradient explosion and improve the stability of model training.

[0060] Performance validation of the inverse synthesis prediction model: The accuracy of the trained inverse synthesis prediction model was validated using a reserved test set. The exact matching accuracy and the maximum fragment matching accuracy of MaxFrag were used as the core evaluation metrics. The validation results at different levels (Top-1, Top-3, Top-5, Top-10, and Top-50) are shown in Table 1 below. Table 1 TOP-K Accuracy

[0061] The verification results show that the retrosynthesis prediction model constructed in this invention can achieve a Top-10 exact matching accuracy of over 85%, which fully meets the practical application requirements of molecular synthesis route planning for organic cathode materials, and the prediction accuracy is better than that of traditional retrosynthesis prediction methods.

[0062] S3.4 Optimal synthesis route determined: Starting with the target organic cathode material molecule screened by S2, the reaction template with the highest probability score (Top K) output by the above model is used to reverse generate a legal precursor molecule, completing a single-step retrosynthesis prediction. Then, based on the constructed retrosynthesis prediction model, a multi-step iterative search is performed to generate multiple candidate synthesis routes. The model sorts the candidate routes according to the reaction template prediction score, with higher scores indicating higher route feasibility. The candidate route with the highest ranking is selected as the optimal synthesis route, completing the retrosynthesis path planning of the target organic cathode material molecule.

[0063] Based on the optimal target organic cathode material selected in the aforementioned property prediction examples, using nitroglycerin (NUDT-EEI-01) as the object, a retrosynthetic route was planned and experimentally verified. The specific process is as follows: Multi-step retrosynthetic route search: Using the SMILES expression of the optimal target organic cathode material molecule as input, the retrosynthetic prediction model trained by the model is used to perform multi-step iterative search. The maximum number of synthesis steps is set to 5. Each step outputs the Top-10 high-scoring candidate precursors, and finally 12 complete candidate synthesis routes are generated. Optimal synthetic route selection: Based on the model's predicted total score, number of synthetic steps, commercial availability of raw materials, theoretical total yield, and mild reaction conditions, the optimal synthetic route was determined to be a one-step synthesis: using commercially available glycourea as the starting material, the target product tetranitroglycourea is prepared in one step through a nitration reaction. The model predicts that the single-step reaction yield of this route is ≥70%, the raw materials are all commercially available conventional reagents, the reaction conditions are mild, there are no harsh requirements such as high temperature and high pressure, and it is easy to scale up the preparation.

[0064] Model generalization verification: To verify the universality of the retrosynthesis prediction model of this invention for organic cathode materials with different parent core structures, 20 candidate organic cathode materials, including phenazines, anthraquinones, triazoles, and imidazoles, obtained from the aforementioned property prediction examples, were selected for retrosynthesis route planning verification.

[0065] The verification results show that the retrosynthesis prediction model of the present invention can generate chemically legal and step-reasonable synthetic routes for all test molecules. The Top-5 routes all include feasible synthetic paths that have been reported in the literature or verified experimentally. The model maintains stable prediction performance for organic cathode material molecules with different core structures, without significant decrease in accuracy. It has excellent generalization ability and can be widely applied to the design of synthetic routes for organic cathode materials.

[0066] S4 synthesis verification and model iterative optimization; Based on the optimal route planned in step 3, experimental preparation of the target organic cathode material molecules and measurement of their physicochemical / electrochemical properties were conducted to verify the effectiveness of the design. Simultaneously, the experimentally obtained structural characterization and synthesis data were added to the organic reaction dataset and energy storage material dataset to perform closed-loop iterative training of the model, continuously improving the high-throughput prediction accuracy and generalization ability.

[0067] Reference Figures 11 to 14 To investigate the physicochemical properties of the target organic cathode material molecule, the synthetic route of the target organic cathode material molecule was verified: Glycourea (200 mg, 1.407 mmol, 1 eq) was added to a dry three-necked flask. After purging with nitrogen three times, 4.0 mL of 98% concentrated sulfuric acid was added. The mixture was cooled to 0°C in an ice bath, and 2.0 mL of fuming nitric acid was slowly added dropwise with stirring. After the addition was complete, the temperature was raised to 25°C, and the reaction was stirred for 6 h. After the reaction was complete, the reaction solution was slowly poured into ice water, precipitating a white solid. The filter cake was collected and washed successively with deionized water and anhydrous ethanol until the filtrate was neutral. The filtrate was then dried under vacuum at 60°C for 10 h to obtain 286 mg of the white solid product nitroglycourea, with a yield of 76.2%. After recrystallization, the purity was ≥99.5%. Structural characterization results: The XRD pattern showed characteristic diffraction peaks of the material at 17°, 23°, 24.5°, and 28.3°; the Raman spectrum showed a peak at 1120 cm⁻¹. -1and 1380cm -1 The characteristic vibrational peak corresponding to the nitro group is located at 1540 cm⁻¹ in the FTIR spectrum. - ¹ and 1320cm - The presence of characteristic peaks for both asymmetric and symmetric stretching vibrations of the nitro group at position ¹ proves that the material was successfully synthesized.

[0068] Material performance testing and verification: CR2032 coin cells were assembled in an argon atmosphere glove box with a 19mm diameter polypropylene separator (Celgard 2400) as the separator, using a positive electrode sheet prepared with nitroglycerin as the working electrode, a 15.6mm diameter lithium metal sheet as the counter / reference electrode, and an argon atmosphere glove box with water and oxygen contents both <0.1ppm. During battery assembly, the negative electrode casing, lithium metal sheet, separator, positive electrode sheet, 2025 type metal gasket, 2016 type metal spring sheet, and positive electrode casing were stacked sequentially. 50μL of electrolyte was added to wet the electrode sheet and separator, and the cells were sealed using a sealing machine before being removed. The electrolyte is a 1.5 mol / L lithium bisfluorosulfonylimide (LiFSI) + lithium bis(trifluoromethanesulfonylimide) (LiTFSI) (molar ratio 1:1) ethylene carbonate / diethyl carbonate (volume ratio 1:1) solution. The assembled battery is left to stand at 25°C for 10 hours until the electrolyte fully wets the separator and electrodes before electrochemical performance testing.

[0069] Electrochemical performance testing was performed using the A211-BTS-4S-1U charge-discharge tester from Shenzhen Xinwei Electronics Co., Ltd., with reference to... Figure 15 The battery was subjected to constant current charge-discharge tests within a voltage range of 1.5–3.8V and a current density of 50 mA / g. The test results showed that the initial charge specific capacity was 293 mAh / g, with a clear charging plateau at 3.0V; the initial discharge specific capacity was 350 mAh / g, with two significant discharge plateaus at 2.75V and 1.6V, highly consistent with the theoretical specific capacity and redox potential predicted by the model; the second charge-discharge specific capacities were 235 mAh / g and 250 mAh / g, respectively, with capacity decay mainly due to irreversible reactions and trace dissolution of active materials during the initial charge-discharge process.

Claims

1. A method for predicting the performance of organic cathode material molecules for lithium batteries, characterized in that, Includes the following steps: S1: Construct a multi-source organic energy storage material dataset, extract a subset of key molecular descriptors from the molecules in the dataset using a molecular descriptor extraction tool, and extract the three-dimensional deep implicit features of the molecules in the dataset using a backbone feature extraction network. S2: The three-dimensional molecular deep implicit features are spliced ​​and fused with the subset of key molecular descriptors to generate a comprehensive molecular characterization vector. A multi-task prediction head is constructed with the comprehensive molecular characterization vector as input, and multiple key physicochemical properties are output to obtain a multi-physicochemical property prediction model. S3: Based on the multi-physicochemical performance prediction model, high-throughput prediction is performed on the constructed organic cathode material molecule to obtain the multi-dimensional physicochemical performance parameters of the molecule. Based on the preset electrochemical characteristic threshold, step-by-step judgment and multi-condition joint constraint screening are performed to obtain the target organic cathode material molecule.

2. The method according to claim 1, characterized in that, In step S1, the specific process includes: The multi-source organic energy storage material dataset is constructed, which includes an organic energy storage material-specific dataset and a commercial pre-trained dataset. The organic energy storage material-specific dataset contains molecular physicochemical property labels constructed by first-principles calculations and experimental tests. Molecular descriptor extraction tools are used to quantitatively characterize the molecular structures in the dedicated dataset, obtaining initial high-dimensional molecular descriptors including topological features, electronic features, thermodynamic features, geometric conformation features, charge features, hydrogen bond features, and aromaticity features. Based on SHAP interpretability analysis, the initial high-dimensional molecular descriptors are ranked and screened according to their contribution to obtain the subset of key molecular descriptors. The Uni-Mol model based on the Transformer architecture of equivariant spatial coding is used as the backbone feature extraction network for model pre-training and fine-tuning to obtain the deep implicit features of the three-dimensional molecule.

3. The method according to claim 2, characterized in that, In step S2, the specific process includes: performing missing value imputation and Z-score standardization on the subset of key molecular descriptors to unify the numerical format with the three-dimensional molecular deep implicit features; concatenating the three-dimensional molecular deep implicit features and the subset of key molecular descriptors after linear weighting according to the optimal weights to generate a comprehensive molecular representation vector. During linear weighted fusion, the optimal weight α of the three-dimensional molecular deep implicit features is 0.

66. The linear weighted fusion formula is as follows: y^=Concat(0.66y^3D, 0.34y^desc); Where y^3D is a 128-dimensional deep implicit feature vector, y^desc is a 32-dimensional key descriptor subset vector, Concat(·) represents the concatenation operation along the feature dimension, and y^ is the comprehensive molecular characterization vector. A multi-task prediction head is constructed using the comprehensive molecular characterization vector, and supervised training is performed using a joint loss function to obtain the multi-physicochemical performance prediction model.

4. The method according to claim 3, characterized in that, The multi-physicochemical performance prediction model outputs multiple key physicochemical properties, including thermodynamic characteristics, electronic characteristics, and polar characteristics, through the multi-task prediction head. The key physicochemical properties include at least: molecular enthalpy of formation, molecular frontier orbital energy level HOMO / LUMO, melting point, boiling point, dipole moment, and molecular electrostatic potential.

5. The method according to claim 3, characterized in that, Pre-training phase: Based on the unlabeled molecules on the commercial pre-training dataset, unsupervised pre-training is performed through masked atom prediction and 3D coordinate recovery self-supervised tasks to learn general molecular structure rules and obtain pre-training weights; Fine-tuning stage: Load the pre-trained weights and perform supervised fine-tuning on the organic energy storage material-specific dataset to output the three-dimensional molecular deep implicit features.

6. The method according to claim 4, characterized in that, In step S3, the preset electrochemical characteristic thresholds are subjected to step-by-step judgment and multi-condition joint constraint screening. The specific screening rules include: Based on the structural stability requirements of battery cycling, the molecular formation enthalpy is controlled within the preset thermodynamic stability range of -200 to -800 kJ per mole; Based on the electrolyte compatibility requirements, the HOMO and LUMO energy levels are controlled within the range of -5.2 to -6.0 electron volts, which is compatible with the lithium battery electrolyte and operating potential, relative to the vacuum energy level. According to the requirements of electrode processing and working temperature, the melting point and boiling point are controlled within the preset high temperature range, with the melting point greater than 180 degrees Celsius, the boiling point greater than 350 degrees Celsius, or the decomposition temperature greater than 250 degrees Celsius. Based on the redox activity requirements, the dipole moment was controlled within the range of 15-25 Debye, which characterizes high specific capacity. According to the solubility control requirements, the molecular electrostatic potential is controlled by a polar exponent based on the normalization of the variance of the molecular surface electrostatic potential, with a preset range of 0.15–0.35 and the difference between the extreme values ​​of positive and negative electrostatic potentials ΔV is less than 1.2 electron volts. During the screening process, molecules that do not meet any of the above performance threshold conditions are sequentially removed, and molecules that meet all performance constraints are retained as the target organic cathode material molecules.

7. A method for synthesizing and verifying organic cathode material molecules for lithium batteries, used to synthesize and verify target organic cathode material molecules obtained by the prediction method as described in any one of claims 1 to 6, characterized in that, Includes the following steps: S1: Construct a standardized reaction dataset containing precise atom mappings, extract local reaction templates, and train and construct an inverse synthesis prediction model based on a graph neural network architecture of message passing neural network and global reactive attention. S2: Input the structural information of the target organic cathode material molecule into the retrosynthesis prediction model, use the retrosynthesis prediction model to output the probability score of the local reaction template of the target organic cathode material molecule to complete the single-step retrosynthesis prediction, and then perform multi-step iterative search to generate multiple candidate synthesis routes. Select the candidate route with the highest score as the optimal synthesis route; prepare the target organic cathode material molecule based on the optimal synthesis route and perform performance testing. Feed the experimental characterization and test data back to the model training dataset to perform closed-loop iterative optimization of the retrosynthesis prediction model.

8. The method according to claim 7, characterized in that, In step S1, the specific process includes: Construction of standardized reaction dataset: Single-step organic reaction data were obtained from the Reaxys database, and the data was cleaned and standardized. Atom mapping tools were used to clarify the atomic correspondence between reactants and products, locate the reaction center, and remove low-quality reactions to obtain a standardized reaction dataset with accurate atom mapping. Local reaction template extraction: For each reaction in the standardized reaction dataset, the structures of the mapped reactants and products are compared to locate the changes in atoms and chemical bonds at the reaction center to obtain a local reaction template. The local reaction templates are classified into atomic reaction templates, bond reaction templates and multiple change reaction templates according to the change type. Construction and training of the retrosynthesis prediction model: The product molecule is represented as an atomic-bond topological graph structure for feature initialization; the message-passing neural network is used to iteratively update the atomic features, and the local reaction center features are extracted by fusing neighborhood chemical environment information and reconstructing and enhancing bond features; a multi-head global reactive attention mechanism is introduced to capture global structural information and output globally perceived atomic and bond features; using the globally perceived atomic and bond features as inputs, atomic template classifiers and bond template classifiers are constructed respectively, and the probability scores of each atom and each chemical bond corresponding to various local reaction templates are output; the cross-entropy loss function is used for training to obtain the retrosynthesis prediction model.

9. A prediction system based on the method of claim 7 or 8, characterized in that, The system includes a first module that receives constructed organic cathode material molecules and outputs key physicochemical properties of the organic cathode material molecules through a multi-physicochemical performance prediction model in the first module; a second module that receives the key physicochemical properties of the organic cathode material molecules output by the first module and can screen target organic cathode material molecules by setting constraints in an interactive manner; and a third module that receives the target organic cathode material molecules screened by the second module and performs single-step or multi-step optimal synthesis path planning through a retrosynthesis prediction model in the third module.

10. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores the prediction system as described in claim 9, the prediction system being executed by the at least one processor.

Citation Information

Patent Citations

  • Data-mechanism driven material attribute prediction method of graph neural network

    CN114818948A

  • A method and apparatus for predicting molecular properties by integrating three-dimensional structure and prior features.

    CN120808952B