Method for predicting lowest triplet excited state energy of receptor micromolecule

Through the method of redundant factor correction descriptor, combined with machine learning algorithms and structure-effect relationship, the accuracy and efficiency of T1 prediction of acceptor small molecule in organic photovoltaic materials are solved, and a high-performance material screening framework is realized, which improves the design speed and accuracy of OPV materials.

CN120432045APending Publication Date: 2025-08-05ZHUHAI COLLEGE OF JILIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478291.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately predict the minimum triplet excited state energy of small receptor molecules in organic photovoltaic (OPV) materials. The traditional method is time-consuming and costly, and fragment analysis cannot take into account the effects of intramolecular interactions and basis group overlap, resulting in inaccurate prediction results.

Method used

The redundant factor (σexc) correction descriptor is used to predict the T1 of the molecule through machine learning algorithms, and the database is constructed using density functional theory calculations, and the molecular construction units are modified based on structure-activity relationships, including aromatic cores, side chains and end group engineering to achieve fast and accurate T1 prediction.

Benefits of technology

Fast and accurate T1 prediction based on molecular construction units are achieved, and the error is reduced to the millite level, which improves the screening efficiency and accuracy of organic photovoltaic materials and meets practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005361956160000081
    Figure BDA0005361956160000081
  • Figure BDA0005361956160000091
    Figure BDA0005361956160000091
  • Figure BDA0005361956160000101
    Figure BDA0005361956160000101
Patent Text Reader

Abstract

The invention discloses a method for predicting the lowest triplet excited state energy of receptor micromolecules, which comprises the following steps of: acquiring data and executing density functional theory (DFT) calculation to construct a database containing molecules and ET1 energy (namely descriptors) of construction fragments of the molecules; according to an ITIC series molecular modification strategy, statistical analysis is carried out on the structural features of all the construction units, and a redundant factor sigma exc is obtained; and predicting the T1 of the molecule through a machine learning algorithm by using the descriptor modified by the redundant factor sigmaexc. According to the method, the correction parameter of the redundant factor is provided and is applied to the rapid design and screening framework of the high-performance A-D-A OPV material of the ML algorithm. On the basis of the structure-function relationship, under modification of redundant factors, the photovoltaic characteristics of the molecules are coupled by means of constructing unit prediction molecules T1 and applying machine learning mathematics to quantify the data relationship, and the method has good interpretability. Through the trained ML model, direct, rapid and accurate ET1 prediction based on the molecular construction unit is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of molecular energy prediction, and in particular to a method for predicting the lowest triplet excited state energy of a receptor small molecule. Background Art

[0002] Organic photovoltaics (OPV) are considered to be a promising renewable energy technology due to their advantages such as light weight, flexibility and easy processing. Thanks to the emergence of ADA small molecule non-fullerene acceptors (ADA NFAs), the power conversion efficiency (PCE) of single junction devices has increased rapidly in recent years and is close to 20%. Despite this, the performance of OPV still lags behind that of inorganic and perovskite solar cells. First, bimolecular charge recombination is more likely to occur in organic photovoltaics (OPV), and how to further suppress bimolecular recombination, especially non-radiative recombination losses, is one of the top priorities in the development of OSCs. The formation process of the triplet state dominates this link, and T1 plays an important role in the loss mechanism. Secondly, the process of singlet splitting to obtain two carriers through one photon can indeed improve the photoelectric conversion efficiency, and is expected to break through the Shockley–Queisser limit, but the stringent requirements for this process (E S ≥2E T ) then E T1 There is a clear prediction. Researchers gradually realized the importance of T1 and T1 It is cited as an important descriptor for machine learning prediction of PCE. How to accurately predict the E of a specific molecule T1 Imminent.

[0003] However, the specific range of E T1The process of identifying NFAs is extremely difficult. Achieving breakthroughs with traditional experimental methods requires significant human and time resources. While DFT calculations free researchers from repetitive tasks such as complex organic synthesis and purification, while also providing more accurate insights into optical properties, the significant time cost still limits the application of high-throughput computation in molecular screening. Furthermore, analyzing the entire molecule solely from fragments fails to account for factors such as intramolecular interactions and basis set overlap, making it difficult to guarantee accurate predictions. Therefore, developing an intuitive structure-activity relationship (SPR)-based prediction model to pre-screen a large number of NFAs and subsequently identify more suitable potential molecules is an urgent task. In this context, computer-assisted prediction approaches are the preferred approach. Data-driven research has ushered in a new paradigm: by understanding the fundamental factors underlying the performance of materials for specific applications, using loss functions or probabilistic analysis to predict outcomes and make decisions. This is machine learning (ML). It learns from past data and assists in screening candidate candidates for laboratory experiments, even exploring unknown materials with similar properties. ML has also gradually penetrated into many sub-disciplines of materials science. Machine learning has demonstrated excellent effectiveness in accelerating the discovery of new materials, guiding the design of new materials, and exploring the quantitative structure-property relationship (QSPR) of materials.

[0004] In order to efficiently explore potential molecular design strategies, suitable building blocks are important. Based on the electron-donating and electron-withdrawing interactions in the push-pull conjugated system of ADA NFAs, they can be divided into two types of building blocks with different functions, namely donor and acceptor groups (A and D). To date, a large number of building blocks have been discovered for each type, and there are infinite ways to combine them. Due to the lack of appropriate building strategies to guide, in practice, the traditional hit-and-trial approach is inevitably used, which makes the search process both expensive and time-consuming. Summary of the Invention

[0005] Based on the technical problems existing in the background technology, the present invention proposes a method for predicting the lowest triplet excited state energy of a small receptor molecule, and proposes to use the excess factor (σ exc ) and used it as a descriptor to predict the molecular E T1 Through the trained ML model, direct, fast and accurate E T1 predict.

[0006] The present invention proposes a method for predicting the lowest triplet excited state energy of a receptor small molecule, comprising the following steps:

[0007] S1. Collect data and perform density functional theory calculations to construct a model containing the molecule and its building blocks E. T1 A database of energies, i.e. a database of descriptors;

[0008] S2. According to the ITIC series of molecular modification strategies, the structural characteristics of each building block are statistically analyzed and the redundant factor σ is obtained. exc ;

[0009] S3, using the redundant factor σ exc The modified descriptors were used to predict the T1 of the molecule using a machine learning algorithm.

[0010] Preferably, in step S1, the molecule is an ADA NFAs molecule. According to the different electron-donating and electron-withdrawing functions within the ADA NFAs molecule, the molecule is divided into two building blocks, a donor D and an acceptor A. The molecule is divided by adding a methyl group to the A or D segment to retain the effect of the CC single bond between the A and D segments.

[0011] Preferably, in step S1, the 52 building blocks and the 92 AD-ANFAs molecules composed thereof are subjected to DFT and TDDFT calculations. 1,A and T 1,D The T1 dataset composed of descriptors and molecules was divided into training set and test set in a ratio of 9:1 for model learning.

[0012] Preferably, the mean absolute error, root mean square error, coefficient of determination and Pearson correlation coefficient calculated and output by various models all adopt the values reported by fold-cross validation to offset the limitations of the segmented data set and the particularity of the samples.

[0013] Preferably, in step S2, based on the structure-activity relationship, the modification method of ADA NFAs is: aromatic core engineering, dividing the D segment structure into a basic framework of 5, 6, or 7 fused ring forms.

[0014] Preferably, in step S2, based on the structure-activity relationship, the modification method of ADA NFAs is: side chain engineering, according to the linking mode of the side chain, alkyl, alkylphenyl, and alkylthiophene are divided into two categories, namely non-conjugated side chains and conjugated side chains. For the non-conjugated side chain type, the classification is refined according to the isomerization effect and the side chain coupling strategy.

[0015] Preferably, in step S2, based on the structure-activity relationship, the modification method of ADA NFAs is: end group engineering. The modification of the A fragment mainly involves two types: one is the substitution of the three hydrogen atoms at the terminal end of the nitrile indanone, and the other is the strategy of ring expansion or replacement of the aromatic ring structure.

[0016] Preferably, in step S3, after modifying the descriptor with redundant factors, the error distribution is more concentrated, thereby reducing ΔT1 to ±0.005eV. For the GBDT model, its MAE on the test set is reduced to 0.3571E-03, RMSE is reduced to 0.519E-03, R20.9791, and r is 0.9919.

[0017] Compared with existing technologies, this invention offers the advantage of proposing a correction parameter, the redundant factor, which is applied to a rapid design and screening framework for high-performance ADA OPV materials using an ML algorithm. Starting from the structure-activity relationship, and modified by the redundant factor, the T1 of the molecule is predicted using building blocks. This method then couples the molecule's photovoltaic properties using machine learning mathematically quantified data relationships, demonstrating excellent interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is σ pat Correlation curves with ΔLS-O and ΔLC-C;

[0019] Figure 2 The trend graphs of the predicted T1 and prediction set of the five models are shown;

[0020] Figure 3 The prediction error (ΔT1) curve of different ML training models;

[0021] Figure 4 For each descriptor and E T1 Pearson correlation coefficient heatmap. DETAILED DESCRIPTION

[0022] The present invention will be further explained below with reference to specific embodiments.

[0023] Example

[0024] This example proposes a method for predicting the lowest triplet excited state energy of a small receptor molecule, which specifically includes the following three steps:

[0025] Step S1: Collect data and perform density functional theory (DFT) calculations to construct a model containing the molecule and its building blocks E T1 A database of energies (i.e., descriptors);

[0026] Step S2: Statistically analyze the structural characteristics of each building block according to the ITIC series molecular modification strategy and obtain the redundant factor σ exc ;

[0027] Step S3: Using the redundant factor σ exc The modified descriptors were used to predict the T1 of the molecule using a machine learning algorithm.

[0028] ADA NFAs are naturally highly tailorable, allowing them to be divided into two building blocks: the donor D and the acceptor A, based on their different electron-donating and electron-withdrawing functions. To preserve the influence of the CC single bond between the A and D segments, the molecules were divided by adding a methyl group to either the A or D segment for the following reasons:

[0029] 1) When synthesizing the target product experimentally (e.g., Knoevenagel condensation reaction in ITIC), the bridging carbon between fragments A and D is the aldehyde carbon at the end of fragment D. Therefore, in subsequent analysis, the single bond of the bridging carbon should be attributed to fragment D.

[0030] 2) The results of this method are closest to T1, which facilitates numerically intuitive regression prediction;

[0031] 3) In addition to the strong CC bond, there is also a relatively strong SO weak interaction between the A and D fragments, and the strong response region where the weak interaction is located is biased towards the D side. Therefore, assigning the single room to the D fragment can reduce the deviation caused by the segmentation prediction.

[0032] Sufficient data sets and easily accessible descriptors are necessary conditions for efficient screening of candidate molecules. Therefore, DFT and TDDFT calculations were performed on 52 common building blocks and 92 reported ADA NFAs composed of them. 1,A and T 1,D The dataset of T1 as descriptor and molecule was divided into training set and test set in a ratio of 9:1 for model learning. Given the small size of the dataset, in order to offset the limitations of the segmented dataset and the particularity of the samples, the mean absolute error (MASE), root mean square error (RMSE), and coefficient of determination (R 2 ) and Pearson correlation coefficient (r) were reported using fold-wise cross-validation. All calculations were performed in Gaussian 09, and IRI plots were visualized using the Multiwfn plug-in.

[0033] The molecule is broken into A and D fragments, and the weak interactions between CC single bonds and SO in the molecule are destroyed. In addition, the influence of basis set overlap (BSSE) on the excited state of the molecule makes the cumulative total value of various building blocks significantly higher than the T1 calculated based on the entire molecule (expressed as T 1,A ×2+T 1,D ).

[0034] Based on the direct machine learning prediction of A and D segments as descriptors, although the evaluation indicators MASE, RMSE, R 2The correlation coefficient r indicates that the model performs well in terms of stability and goodness of fit. However, the model exhibits large fluctuations in prediction error (T1-T1_pre), ranging from 0.09 eV to -0.03 eV. For example, the T1 value of ITIC molecules is accurate to the percentile (1.27 eV). For a series of ADA-type NFAs, the fluctuations are all at the percentile level. Given this, the current prediction accuracy is clearly insufficient. To meet the demand for T1 prediction in practical applications, it is necessary to improve the accuracy to the thousandth level.

[0035] To this end, it is proposed to use the redundant factor (σ exc ) modified descriptors. Based on the structure-activity relationship (SPR), three modification strategies of ADA NFAs were adopted, as follows:

[0036] 1) Aromatic core engineering: the D segment structure is divided into the basic framework of 5-, 6-, and 7-fused ring forms;

[0037] 2) Side chain engineering: Based on the linking mode of the side chain, we divide the structures such as alkyl, alkylphenyl, and alkylthiophene into two categories, namely SP 3 (non-conjugated side chain) and SP 2 For SP3 side chain types, we further categorize them based on isomers and side-chain conjugation strategies.

[0038] 3) End group engineering: The modification of fragment A mainly involves two categories: one is the substitution of the three hydrogen atoms at the end of the nitrile indanone group, and the other is several strategies for ring expansion or replacement of the aromatic ring structure.

[0039] The modification of molecular structure is usually based on the combination of multiple strategies under the existing modification strategy. For example, relative to the ITIC molecule, the IT end group is fluorinated while the alkyl group on the IC side chain benzene ring is isomerized (m-ITIC-2F 22 ), the IT benzene ring is replaced by a thiophene ring while the IC side chain is replaced by an alkylthiophene group (IDT6CN-Th 23 ), etc., combined the deviation statistics of several representative modification strategies and believed that a unified correction can be achieved by accumulating and combining the deviations of excited states under various molecular structures. The statistical results show that the variation of the deviations under the same type of strategy is very small and has a strong regularity. Therefore, σ exc It is broken down into two components: σ bas (aromatic nuclear engineering) and σ pat (side chain engineering, end group engineering) two parts, the σ in each category patPerform independent univariate statistical analysis, classification, and calculate the average value of each category to obtain a series of σ representing different end group transformation modes. pat When we want to modify the descriptor, we only need to modify the σ generated by the different parts of the A and D segments. bas and σσ pat When processing segment A, since its modification involves changes on both sides, half of the value should be taken for accumulation when calculating the correction value. bas and σ pat Detailed classification methods and values are shown in the table. It is noteworthy that this combination of aromatic cores, side chains, and end groups allows for predictions that are not limited to existing molecular structure types. Based on structural rationality, and following existing rules and starting with the calculation results of a single building block, a large number of low-cost, high-precision scans can be performed on a variety of combinations. Furthermore, thanks to the characteristics of machine learning, real-time response to molecular structural changes can be achieved. The specific formula is as follows:

[0040] σ exc =(T 1,A ×2+T 1,D )-T1=σ pta +∑σ pta

[0041]

[0042] D_σ=T 1,D +σ bas +∑σ pta

[0043] Figure 1 In the figure, solid marks are σ-pat, hollow marks are ΔLS-O and ΔLC-C, blue lines are fitting curves of ΔLS-O and ΔLC-C, red lines are fitting curves of σ-pat, the left and right figures are end group and side chain strategies respectively, and the upper and lower figures are ΔLS-O and ΔLC-C respectively. Figure 1 It can be seen that σ pat The change of parameters has little effect on the CC bond length at the rupture position (the range of change is less than 0.005i). pat As the value decreases, the distance between S and O decreases accordingly. Due to the conformational lock structure formed by the interaction between S and O atoms, the molecule is forced to maintain a planar rigid conformation. 24 , the reduction of SO spacing will make the conformation lock more stable, thus reducing the possibility of torsion. If the molecule has strong planarity, then the direct superposition of A and D fragments can fully reflect the microscopic characteristics of the molecule. This further confirms the use of fragments as descriptors and σ pat The rationality of the classification method.

[0044] Figure 2 The trend chart of the prediction T1 and prediction set of the five models is shown below. Figure 3 The prediction error (ΔT1) curve of different ML training models is shown in Figure 2. Figure 4 For each descriptor and E T1 The Pearson correlation coefficient heatmap and Table 1 show the predictions of T1 for the GBDT, Xgboost, LingtGBM, SVR, and (RR) models. It can be seen that after modifying the descriptors with redundant factors, the error distribution becomes more concentrated, significantly reducing ΔT1 to approximately ±0.005 eV and significantly improving the accuracy of predicted T1 energy levels, reaching the required accuracy range for practical applications. Table 1 shows that after redundant factor modification, the five ML models show significant improvements in four key evaluation metrics on both the training and test sets. In particular, for the GBDT model, its MAE on the test set dropped to 0.3571E-03, RMSE to 0.519E-03, R20.9791, and r reached 0.9919. These numerical improvements overall highlight the excellent predictive performance of the models and confirm the importance of using redundant factor modification.

[0045] Table 1 Prediction of T1 by several models

[0046]

[0047] This example proposes a correction parameter, the redundant factor. Specifically, a redundant factor comparison table, such as that shown in Table 2, is applied to a rapid design and screening framework for high-performance ADA OPV materials using an ML algorithm. Starting from the structure-activity relationship and modified by the redundant factor, the T1 of a molecule is predicted using building blocks. Machine learning mathematically quantifies data relationships to couple the molecule's photovoltaic properties, demonstrating excellent interpretability.

[0048] Table 2 Redundant Factor Comparison Table

[0049]

[0050] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for predicting the lowest triplet excited state energy of a small receptor molecule, characterized in that: The following steps are involved: S1. Collect data and perform density functional theory calculations to construct a model containing the molecule and its building blocks E. T1 A database of energies, i.e. a database of descriptors; S2. According to the ITIC series of molecular modification strategies, the structural characteristics of each building block are statistically analyzed and the redundant factor σ is obtained. exc ; S3, using the redundant factor σ exc The modified descriptors were used to predict the T1 of the molecule using a machine learning algorithm.

2. The method for predicting the lowest triplet excited state energy of a receptor small molecule according to claim 1, characterized in that: In step S1, the molecule is an ADA NFAs molecule. According to the different electron-donating and electron-withdrawing functions within the ADA NFAs molecule, the molecule is divided into two building blocks, a donor D and an acceptor A. The molecule is divided by adding a methyl group to the A or D segment to retain the effect of the CC single bond between the A and D segments.

3. The method for predicting the lowest triplet excited state energy of a receptor small molecule according to claim 1, characterized in that: In step S1, the 52 building blocks and the 92 ADA NFAs molecules composed thereof were subjected to DFT and TDDFT calculations. 1,A and T 1,D The T1 dataset composed of descriptors and molecules was divided into training set and test set in a ratio of 9:1 for model learning.

4. The method for predicting the lowest triplet excited state energy of a receptor small molecule according to claim 3, characterized in that: The mean absolute error, root mean square error, coefficient of determination, and Pearson correlation coefficient of the calculated outputs of various models all adopted the values reported by fold-cross validation to offset the limitations of the segmented dataset and the particularity of the samples.

5. The method for predicting the lowest triplet excited state energy of a receptor small molecule according to claim 1, characterized in that: In step S2, based on the structure-activity relationship, the modification method of ADA NFAs is: aromatic core engineering, which divides the D segment structure into a basic framework of 5, 6, and 7 fused ring forms.

6. The method for predicting the lowest triplet excited state energy of a receptor small molecule according to claim 1, characterized in that: In step S2, based on the structure-activity relationship, the modification method of ADA NFAs is: side chain engineering. According to the linking mode of the side chains, alkyl groups, alkylphenyl groups, and alkylthiophenes are divided into two major categories, namely non-conjugated side chains and conjugated side chains. For the non-conjugated side chain type, the classification is refined according to the isomerization effect and the side chain coupling strategy.

7. The method for predicting the lowest triplet excited state energy of a receptor small molecule according to claim 1, characterized in that: In step S2, based on the structure-activity relationship, the modification method of ADA NFAs is: end group engineering. The modification of the A fragment mainly involves two categories: one is the substitution of the three hydrogen atoms at the terminal end of the nitrile indanone, and the other is the strategy of ring expansion or replacement of the aromatic ring structure.

8. The method for predicting the lowest triplet excited state energy of a receptor small molecule according to claim 1, characterized in that: In step S3, after modifying the descriptor with redundant factors, the error distribution is more concentrated, thereby reducing ΔT1 to ±0.005eV. For the GBDT model, its MAE on the test set is reduced to 0.3571E-03, RMSE is reduced to 0.519E-03, R20.9791, and r is 0.9919.