N-type polymer thermoelectric performance prediction method based on structural decoupling and neural network

By using structural decoupling and neural network methods, the backbone and side chain features of N-type polymers are separated. Feature fusion is performed using GINE and ANN networks, which solves the problems of high prediction accuracy and high cost in existing technologies and achieves efficient prediction of the thermoelectric properties of N-type polymers.

CN122157859BActive Publication Date: 2026-08-04NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2026-05-09
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively capture the electronic transport characteristics of the conjugated backbone in N-type polymers. Side chain information interference prevents models from accurately predicting thermoelectric properties, and traditional methods are costly and time-consuming.

Method used

By employing structural decoupling and neural network methods, the SMILES sequence of N-type polymers is separated into backbone and side chains, and feature vectors are extracted from each. Feature fusion prediction is performed using GINE graph neural network and ANN artificial neural network, and thermoelectric properties are output using a fusion perceptron.

Benefits of technology

This method enables accurate prediction of the thermoelectric properties of N-type polymers, reducing experimental costs and time, and improving the model's prediction efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157859B_ABST
    Figure CN122157859B_ABST
Patent Text Reader

Abstract

The application discloses a kind of N-type high molecular thermoelectric performance prediction method based on structure decoupling and neural network, the method comprises the following steps: S1, data preprocessing: the molecular formula of N-type high molecular polymer is handled into SMILES sequence;S2, structure decoupling: the connection between the skeleton and side chain is cut off to SMILES sequence and is identified;S3, skeleton feature extraction, skeleton global graph feature vector is extracted by GINE graph neural network;S4, side chain feature extraction, side chain feature vector is generated by ANN artificial neural network;S4, isomeric fusion and prediction, the fusion feature vector is input into fusion sensor, and the output N-type high molecular thermoelectric performance prediction value is obtained by fusion sensor.The application realizes the prediction of N-type high molecular thermoelectric performance effectively by the deep learning architecture of skeleton-side chain decoupling, and greatly reduces the expensive experimental cost and risk required by traditional research and development mode such as "blind trial and error".
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of materials informatics and artificial intelligence, specifically to a method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks, applicable to predicting the thermoelectric properties of N-type polymers, such as conductivity and power factor. Background Technology

[0002] Organic thermoelectric materials, with their advantages of high flexibility, light weight, low thermal conductivity, and solution processability, have shown great promise for applications in wearable electronics, flexible thermoelectric devices, and distributed micro-energy systems. Complete thermoelectric devices require the pairing of P-type and N-type materials; however, while P-type organic thermoelectric materials are mature and perform excellently, research on N-type organic polymer materials lags behind, becoming a bottleneck restricting the large-scale application of organic thermoelectric devices.

[0003] The main reason for the lagging development of N-type organic polymers is that N-type polymers typically have extremely low lowest unoccupied molecular orbitals (LUMO) energy levels. Monomer synthesis involves multiple organic reactions, resulting in low yields, high purification difficulty, and high sensitivity to impurities, thus making synthesis challenging and costly. Furthermore, most N-type polymers are easily oxidized by water and oxygen in the doped state, leading to rapid degradation of their electrical properties. Synthesis and testing require stringent anhydrous and oxygen-free environments.

[0004] Traditional methods involve experimentation using a trial-and-error approach, but this method is time-consuming, costly, and has a low success rate, making it difficult to meet the needs of rapid screening and structural optimization. Therefore, utilizing machine learning and deep learning to achieve rapid performance prediction has become an important development direction in this field.

[0005] The closest existing technology typically employs a approach that combines global molecular characterization with an end-to-end black-box model.

[0006] Global feature representation: Using Morgan fingerprint, MACCS bond fingerprint or overall molecular map, the polymer repeating unit is directly compressed into a single global feature vector without distinguishing between backbone and side chain;

[0007] Unified modeling: Random forests, support vector machines, and traditional graph neural networks (such as graph convolutional networks (GCN) and graph attention networks (GAT)) are used to directly learn the mapping relationship between "overall molecular structure and thermoelectric properties".

[0008] However, existing technologies have obvious drawbacks:

[0009] In N-type polymers, the number of conjugated backbone atoms that dominate electron transport is far less than the number of side chain atoms. During message passing, traditional global graph neural networks will cause a large amount of side chain information to interfere with and overwhelm key backbone features, making it impossible for the model to effectively capture core information that determines conductivity, such as conjugation length, planarity, and electronic structure.

[0010] Chemical priors indicate that the N-type polymer backbone dominates electron transport and energy level structure, while the side chains dominate solubility, crystallinity, molecular packing, and processing properties; the two are functionally completely heterogeneous. Existing technologies use the same set of operators and feature processing mechanisms to treat rigid conjugated backbones and flexible aliphatic side chains in a "one-size-fits-all" manner, which is insensitive to subtle structural changes such as side chain length and branching points. Summary of the Invention

[0011] The technical problem to be solved by this invention is to provide a method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks.

[0012] The technical solution adopted by this invention to solve the above-mentioned technical problems is: a method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks, the method comprising:

[0013] S1. Data preprocessing: The molecular formula of N-type polymers is processed into SMILES sequences. If there are dummy atoms representing polymer sites, C atoms are used to replace the dummy atoms.

[0014] S2. Structural decoupling: Identify the SMILES sequence of N-type polymers and sever the bonds connecting the backbone and side chains. The backbone is defined as the conjugated system that dominates the transport of charge carriers in the molecule; the side chain is defined as an aliphatic chain segment that regulates the solubility or polarity of the molecule.

[0015] S3. Backbone Feature Extraction: Each atom in the backbone is encoded using node feature encoding to form encoded atom nodes; each chemical bond in the backbone is encoded using edge feature encoding to form encoded chemical bond edges; based on the encoded atom nodes and chemical bond edges, a molecular graph structure of the N-type polymer is constructed. The molecular graph structure of the N-type polymer is input into a GINE graph neural network, and the GINE graph neural network extracts the global graph feature vector V of the backbone. backbone ;

[0016] S4. Sidechain Feature Extraction: Extract physicochemical features from the sidechains to form sidechain features; aggregate multiple sidechain features to obtain aggregated sidechain features; perform logarithmic transformation on the aggregated sidechain features and input them into an ANN (Artificial Neural Network) to generate a sidechain feature vector V. sidechain ;

[0017] S5. Heterogeneous Fusion and Prediction: The feature vector V of the skeleton global graph is then used for...backbone With sidechain feature vector V sidechain Dimensional concatenation is performed to obtain a fused feature vector; the fused feature vector is input into a fusion perceptron, which then outputs the predicted thermoelectric properties of the N-type polymer. .

[0018] Preferably, the structural decoupling is achieved using the RDKit tool, which is used to identify and cleave the bonds connecting the backbone and side chains. Specifically, the SMILES sequence of the N-type polymer is first imported into the RDKit tool. The molecular topology corresponding to the SMILES sequence is analyzed, and the backbone belonging to the conjugated system and the side chains belonging to the aliphatic segments are identified. The bonds connecting the backbone and side chains are located and identified, and the bonds are cleaved using a chemical bond cleaving algorithm to separate the backbone and side chains. After cleaving, the backbone is represented as a topological diagram composed of atoms and chemical bonds, and the side chains are represented as a set of molecular chain segments.

[0019] Preferably, in the skeleton feature extraction, the node feature encoding includes element type One-hot encoding, geometric coordination environment One-hot encoding, physical property continuous variable normalization features, and local topological constraint features. The local topological constraint features include: hydrogenation number normalization value, aromaticity Boolean feature, intra-loop attribute Boolean feature, and side chain connection site identifier.

[0020] Preferably, in the skeleton feature extraction, the edge feature encoding includes bond type One-hot encoding and topological attribute Boolean features; the bond type is a single bond, double bond, triple bond, aromatic bond, or other bond type, and the topological attribute includes whether the chemical bond is inside a ring and whether it has conjugation properties.

[0021] Preferably, the GINE uses the GINEConv operator with residual connections;

[0022] ;

[0023] ;

[0024] in, Represents atoms In the Feature vectors after layer iteration Represents atoms The state at level k-1 Represents atoms Neighboring atoms The state at level k-1 Indicates with atoms The set of all directly connected atoms It is the first The multilayer perceptron of the GINE graph neural network corresponding to each layer. It is a learnable scalar parameter; This represents the atoms in the (k-1)th layer. With neighboring atoms The coding features corresponding to the chemical bonds between them.

[0025] Preferably, in the side chain feature extraction, the physicochemical features include the number of heavy atoms, LogP, TPSA, main chain length, side chain length, and the proportion of branch point positions;

[0026] In the side chain feature extraction, the aggregation of multiple side chain features adopts a weighted aggregation method based on the proportion of heavy atoms, and the aggregation formula is as follows: ;

[0027] in, For the first The number of heavy atoms in each side chain, This is the 6-dimensional feature vector of the sidechain; This is a side-chain aggregation characteristic.

[0028] Preferably, the ANN artificial neural network includes two linear perceptron layers. The first linear perceptron layer is used to map physical descriptors to a high-dimensional space, and the second linear perceptron layer is used to extract latent features. The two linear perceptron layers are each equipped with a Dropout layer.

[0029] Preferably, in the heterogeneous fusion and prediction, the skeleton global graph feature vector is... With sidechain feature vectors Dimensional concatenation yields the fused feature vector. :

[0030] ;

[0031] Fusion feature vectors The system enters a fusion perceptron, which includes a ReLU activation layer, a Dropout regularization layer, and a linear mapping layer.

[0032] The ReLU activation layer uses the ReLU activation function for nonlinear interaction, simulating the physical interaction between the backbone and sidechains; the probability of the Dropout regularization layer is 0.3-0.6, which is used to suppress overfitting during small sample training.

[0033] The linear mapping layer outputs predicted values ​​of the thermoelectric properties of N-type polymers. :

[0034] ;

[0035] in It is the final fused feature vector obtained by processing the ReLU activation layer and Dropout regularization layer of the fusion perceptron.

[0036] Preferably, the fusion perceptron includes a conductivity prediction weight branch, which comprises a first branch ReLU activation layer and a first branch linear mapping layer. The first branch ReLU activation layer is used to process the input fusion feature vector. Perform nonlinear transformation; the first branch linear mapping layer outputs conductivity. Predicted value ,

[0037] ,in, This is the weight matrix for conductivity prediction. This is the bias vector for predicting conductivity;

[0038] The fusion perceptron includes a power factor prediction weight branch, which comprises a second branch ReLU activation layer and a second branch linear mapping layer. The second branch ReLU activation layer is used to process the input fusion feature vector. Perform a nonlinear transformation; the second branch linear mapping layer outputs the predicted power factor PF. , ,in, The weight matrix for power factor prediction. This is the bias vector for power factor prediction.

[0039] Preferably, the heterogeneous fusion and prediction step further includes an iterative optimization step, comprising:

[0040] S11. Supervision label determination: The target value y is obtained by performing a logarithmic transformation on the experimentally measured thermoelectric properties of the N-type polymer.

[0041] S12, Loss Function Calculation: Calculate the predicted thermoelectric properties of N-type polymers. The predicted thermoelectric properties of the N-type polymer are calculated by comparing the predicted value with the target value y. The loss function between the target value y and the target value y, wherein the loss function uses the mean squared error;

[0042] S13. Gradient Calculation: Based on the backpropagation algorithm, starting from the loss function, backtrack and calculate the gradients of all learnable parameters in GINE, ANN and fused perceptron.

[0043] S14. Parameter Update: Following the gradient descent strategy, update all learnable parameters in the direction of gradient descent. The parameter update formula is: w represents the parameter to be updated, and η represents the learning rate. The gradient corresponding to the parameter w. These are the updated parameters;

[0044] Repeat steps S12, S13, and S14 until the loss function value converges to the preset threshold, thus completing the iterative optimization.

[0045] The beneficial effects of this invention are as follows: By employing a deep learning architecture that decouples the backbone and side chains, this invention effectively eliminates the feature dilution effect of side chain atomic information on the electronic characteristics of the conjugated backbone in existing technologies. The prediction model includes GINE, ANN, and a fusion perceptron, enabling it to detect minute atomic substitutions in the conjugated main chain and subtle shifts in side chain branching points. As a prediction tool for N-type polymer systems, this invention effectively predicts the thermoelectric properties of N-type polymers, significantly reducing the expensive experimental costs and risks associated with traditional R&D models such as "blind trial and error." Attached Figure Description

[0046] Figure 1 This is a flowchart of the method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks according to the present invention;

[0047] Figure 2 This is the molecular structure diagram of the N-type conjugated polymer P(NDI2OD-T2). Detailed Implementation

[0048] The present invention will now be described in further detail with reference to the accompanying drawings and preferred embodiments. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0049] This invention provides a method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks. By decoupling the polymer's microstructure, it captures the electronic transport and regulation characteristics that determine thermoelectric properties. A prediction model is then used to predict the thermoelectric properties of N-type polymers. This prediction model includes GINE, ANN, and a fusion perceptron, such as... Figure 1 As shown, the prediction method includes the following steps:

[0050] S1. Data preprocessing: The molecular formula of the N-type polymer is processed into the SMILES sequence. If there are dummy atoms representing polymer sites, C atoms are used to replace the dummy atoms.

[0051] like Figure 2The diagram shows the molecular structure of the N-type conjugated polymer P(NDI2OD-T2). The SMILES sequence of P(NDI2OD-T2) is: O=C(C(C1=C2C(C3=O)=C([*])C=C41)=C(C5=C=C(C6=C=C([*])S6)S5)C=C2C(N3CC(CCCCCCCC)CCCCCCCCCC)=O)N(CC(CCCCCCCC)CCCCCCCCCC)C4=O. Replacing the dummy atom * representing the polymer site with a C atom, the resulting SMILES sequence is: O =C(C(C1=C2C(C3=O)=C(C)C=C41)=C(C5=CC=C(C6=CC=C(C)S6)S5)C=C2C(N3CC(CCCCCCCC)CCCCCCCCCC)=O)N(CC(CCCCCCCC)CCCCCCCCCC)C4=O, the processing results indicate that its repeating unit consists of one naphthalene diimide (NDI) and two thiophene (T2) units, with a side chain of 2-octyldodecyl.

[0052] S2. Structural decoupling: Identify the SMILES sequence of N-type polymers and sever the bonds connecting the backbone and side chains. The backbone is defined as the conjugated system that dominates the transport of charge carriers in the molecule; the side chain is defined as an aliphatic chain segment that regulates the solubility or polarity of the molecule.

[0053] The structural decoupling is achieved using the RDKit tool, which is used to identify and cleave the bonds connecting the backbone and side chains. Specifically, the SMILES sequence of the N-type polymer is first imported into the RDKit tool. The molecular topology corresponding to the SMILES sequence is analyzed, and the backbone belonging to the conjugated system and the side chains belonging to the aliphatic segments are identified. The bonds connecting the backbone and side chains are located and identified, and the bonds are cleaved using a chemical bond cleaving algorithm to separate the backbone and side chains. After cleaving, the backbone is represented as a topological map composed of atoms and chemical bonds, and the side chains are represented as a set of molecular chain segments.

[0054] Using the Python cheminformatics RDKit tool, the SMILES sequence of a molecule is input, and the bonds connecting the backbone and side chains are identified and severed.

[0055] S3. Backbone Feature Extraction: Each atom in the backbone is encoded using node feature encoding to form encoded atom nodes; each chemical bond in the backbone is encoded using edge feature encoding to form encoded chemical bond edges; based on the encoded atom nodes and chemical bond edges, a molecular graph structure of the N-type polymer is constructed; the molecular graph structure of the N-type polymer is input into GINE, and the backbone global graph feature vector V is extracted by GINE.backbone .

[0056] (31) In the skeleton feature extraction, the node feature encoding is a 17-dimensional feature vector, including 7-dimensional element type One-hot encoding, 4-dimensional geometric coordination environment One-hot encoding, 2-dimensional physical property continuous variable normalization feature and 4-dimensional local topological constraint feature.

[0057] Each atom in the skeleton is encoded using node features, constructing a feature vector of length 17. The specific physical meaning is as follows:

[0058] Element type identification (7-dimensional): One-hot encoding is used to distinguish {C, N, O, Cl, F, S, Se}. The seven atoms that play a key role in electron transport are: C (carbon), N (nitrogen), O (oxygen), Cl (chlorine), F (fluorine), S (sulfur), and Se (selenium). The encoding method adopts the one-hot rule: only the position corresponding to the current atom is 1, and all others are 0, enabling the prediction model to clearly identify the atom type.

[0059] Geometric coordination environment (4-dimensional): It is encoded based on the atomic degree of the atom in the molecular topology, that is, the number of directly connected neighboring atoms, corresponding to the categories {1,2,3,4}.

[0060] Example:

[0061] Atomic degree=1→[1,0,0,0]

[0062] Atomic degree=2→[0,1,0,0]

[0063] Atomic degree=3→[0,0,1,0]

[0064] Atomic degree=4→[0,0,0,1].

[0065] Continuous variables of physical properties (2-dimensional): 1st dimension: Normalized electronegativity (based on the Pauling scale), using the Pauling scale electronegativity, normalized to the [0,1] interval, used to characterize the electron-withdrawing ability of atoms. 2nd dimension: Normalized atomic mass (Mass / 100), the relative atomic mass (g / mol) of the atom is linearly normalized to Mass / 100 so that it falls into the [0,1] interval.

[0066] Local topological constraints (4-dimensional): include:

[0067] The hydrogen addition number is normalized (first dimension), where num_Hs represents the total number of hydrogen atoms directly bonded to the atom (including implicit hydrogens). In organic conjugated skeletons, the hydrogen addition number of common atoms (C, N, O, etc.) ranges from 0 to 3. This value is compressed to the [0,1] interval by num_Hs / 3.0, aligning with other continuous features (electronegativity, atomic mass) and avoiding gradient imbalance. For example, the hydrogen addition number of an aromatic ring C atom is 0, with a normalized value of 0.0; the hydrogen addition number of a double bond C atom is 1, with a normalized value of 0.333; and the hydrogen addition number of a saturated C atom connected to a side chain is 3, with a normalized value of 1.0.

[0068] Aromaticity Boolean feature (second dimension): IsAromatic is the aromaticity determination result from RDKit, indicating whether an atom belongs to an aromatic ring system. Encoding logic: binary 0 / 1 encoding, no normalization required, directly used as feature input. Aromaticity is a core property of conjugated systems, directly determining electron delocalization capability and carrier transport efficiency.

[0069] Boolean features within the ring (3rd dimension): IsInRing indicates whether an atom is within any ring structure (including aromatic and alicyclic rings). Encoding logic: binary 0 / 1 encoding, direct input. This distinguishes between conjugated atoms within the ring and chain-like side-chain atoms, assisting the prediction model in identifying the topological structure of the conjugated backbone and avoiding interference from chain-like side-chain atoms in backbone feature extraction.

[0070] Sidechain Link Identifier (4th Dimension): The sidechain link identifier (IsSidechainLink) is used to mark whether the backbone atom is connected to the cleaved sidechain. Encoding logic: binary 0 / 1 encoding, direct input. This feature serves as a bridge to decouple the backbone and sidechain information interaction in the prediction model, ensuring awareness of the inductive effect or steric interference of the sidechain on the backbone site.

[0071] For example, in P(NDI2OD-T2), the aromatic bonds within the thiophene ring simultaneously satisfy "intra-ring + conjugated", encoded as [1,1]; the single bond connecting the NDI core and the side chain is a non-intra-ring, non-conjugated bond, encoded as [0,0].

[0072] For P(NDI2OD-T2), 17-dimensional features are extracted for each atom in the framework (such as N or C atoms in NDI), as follows:

[0073] One-hot encoding of element type: If the current atom is N, then the one-hot encoding is [0,1,0,0,0,0,0];

[0074] One-hot encoding of geometric coordination environment: Count the number of neighbors of the atom in the skeleton diagram. For example, the degree of the carbonyl carbon atom in the NDI core is 3, which is encoded as [0,0,1,0].

[0075] Normalization characteristics of continuous variables in physical properties: For example, the Pauli electronegativity of the nitrogen atom is 3.04, and its normalized value is 3.04 / 4.0 = 0.76. The atomic mass is 14.01, and its normalized value is 14.01 / 100 = 0.1401.

[0076] Local topological constraint characteristics: Topological characteristics of atoms in P(NDI2OD-T2):

[0077] Aromatic carbon atoms (C) on the NDI ring:

[0078] It belongs to the naphthalene ring aromatic conjugated system;

[0079] For aromatic carbon atoms, num_Hs=0 → normalization=0.0;

[0080] IsAromatic = 1 for aromatic carbon atoms;

[0081] IsInRing=1 for aromatic carbon atoms;

[0082] IsSidechainLink=0 for aromatic carbon atoms;

[0083] The topological constraint features of aromatic carbon atoms are [0.0,1,1,0].

[0084] The two imide nitrogen atoms N on the NDI ring:

[0085] The imide nitrogen atom's num_Hs=0 → normalized = 0.0;

[0086] The nitrogen atom of the imide has IsAromatic=1 (in an aromatic conjugated system);

[0087] IsInRing=1 for the nitrogen atom of the imide;

[0088] IsSidechainLink=1 for the nitrogen atom of the imide (connecting the 2-octyldodecyl side chain).

[0089] The topological constraint feature of the imide nitrogen atom is [0.0,1,1,1].

[0090] Carbon or sulfur atoms on the thiophene (T2) ring:

[0091] Thiophene is a typical aromatic ring;

[0092] For carbon or sulfur atoms, num_Hs = 0 → normalization = 0.0;

[0093] IsAromatic = 1 for carbon or sulfur atoms;

[0094] IsInRing=1 for carbon or sulfur atoms;

[0095] IsSidechainLink=0 for carbon or sulfur atoms;

[0096] The topological constraint features of carbon or sulfur atoms are [0.0,1,1,0].

[0097] (32) In the skeleton feature extraction, the edge feature encoding is a 7-dimensional feature vector, including a 5-dimensional bond type One-hot encoding and a 2-dimensional topological attribute Boolean feature; the bond type is a single bond, double bond, triple bond, aromatic bond or other bond type, and the topological attribute includes whether the chemical bond is in a ring and whether it has conjugation properties.

[0098] Bond type identification (5-dimensional): Encoding {single bond, double bond, triple bond, aromatic bond, others}.

[0099] Topological properties (2D): Whether the chemical bond is inside the ring and whether it has conjugation properties are also Boolean features.

[0100] For P(NDI2OD-T2), bond type identification: for the carbon-oxygen double bond in NDI, the code is [0,1,0,0,0]; for the aromatic bond in the thiophene ring, the code is [0,0,0,1,0].

[0101] For P(NDI2OD-T2), in the topological properties, all aromatic bonds and conjugated double bonds within the ring are both within the ring and conjugated, encoded as [1,1]; the skeleton conjugated bridge bonds are not within the ring but are conjugated, encoded as [0,1]; the side-chain connecting single bonds and the internal bonds of the side chain are both not within the ring and not conjugated, encoded as [0,0].

[0102] (33) GINE Graph Neural Network:

[0103] The GINE graph neural network employs the GINEConv operator with residual connections to prevent the gradient vanishing problem in deep networks and maintain the transmission of underlying chemical features.

[0104] node In the The update formula for the layer is:

[0105] ;

[0106] in, Represents atoms In the Feature vectors after layer iteration Represents atoms The state at level k-1 Represents atoms Neighboring atoms The state at level k-1 Indicates with atoms The set of all directly connected atoms It is the first The multilayer perceptron of the GINE graph neural network corresponding to each layer. It is a learnable scalar parameter; This represents the atoms in the (k-1)th layer. With neighboring atoms The coding features corresponding to the chemical bonds between them.

[0107] node In aggregated neighbors When receiving information, through The operation linearly superimposes the atomic states with the properties of the chemical bonds that connect them, which means that when the model learns atomic features, it is embedded with the physical constraint of how they are connected.

[0108] Global summation pooling is used to aggregate the hidden states of all nodes:

[0109] ;

[0110] Compared to average pooling, additive pooling can preserve molecular scale information and reflect the influence of molecular weight and conjugated system size on conductivity.

[0111] GINE consists of multiple cascaded GINEConv convolutional layers, activation layers, normalization layers, and global pooling layers. The functions of each layer are as follows:

[0112] (331) GINEConv convolutional layer: performs message passing and feature aggregation on atomic nodes and chemical bond information, and learns the local molecular structure and electronic conjugation features.

[0113] Input: Atomic features, edge (chemical bond) features

[0114] Output: Updated atomic depth features

[0115] (332) ReLU activation layer: Introduces nonlinear expressive power, enabling the model to learn complex structure-performance relationships.

[0116] (333) Batch normalization layer / graph normalization layer: stabilizes training, accelerates convergence, and reduces the differences in the characteristic distribution of different atoms and bonds.

[0117] (334) Residual Connection: solves the gradient vanishing problem in deep networks and ensures that the chemical information of the underlying layer is not lost in the transmission of multiple layers.

[0118] (335) Global AddPooling: Aggregates the features of all atoms into a global representation of the whole molecule for subsequent prediction.

[0119] S4. Sidechain Feature Extraction: Extract physicochemical features from the sidechains to form sidechain features; aggregate multiple sidechain features to obtain aggregated sidechain features; perform logarithmic transformation on the aggregated sidechain features and input them into an ANN to generate a sidechain feature vector V. sidechain .

[0120] In the side chain feature extraction, the physicochemical features are 6-dimensional features, including the number of heavy atoms, LogP, TPSA, main chain length, branch length, and the proportion of branch points. For P(NDI2OD-T2), the number of heavy atoms is 20; the calculated LogP is approximately 7.9 (high hydrophobicity), and the TPSA is 0 (alkyl chain nonpolarity). Branching topology: its main chain length is identified as 19, the branch length as 1, and the branch point is located at the 9th position of the main chain. Therefore, the proportion of branch points is calculated to be 9 / 19.

[0121] The number of heavy atoms reflects the volumetric size of the side chain. LogP and TPSA are used to characterize the hydrophobicity and polarity of the side chain. Parameters such as main chain length, branch length, and the proportion of branching points can effectively distinguish between straight-chain alkyl groups and branches at different branching sites, thereby capturing the interference of side chain steric hindrance on the crystallinity of the skeleton.

[0122] In the side chain feature extraction, the aggregation of multiple side chain features adopts a weighted aggregation method based on the proportion of heavy atoms, and the aggregation formula is as follows: ;

[0123] in, For the first The number of heavy atoms in each side chain, This is the 6-dimensional feature vector of the sidechain; This is a side-chain aggregation characteristic.

[0124] ANN consists of two linear perceptron layers. The first linear perceptron is used to map physical descriptors to a high-dimensional space, and the second linear perceptron is used to extract latent features. Each of the two linear perceptron layers is equipped with a Dropout layer.

[0125] Logarithmic transformation is performed on the sidechain aggregation features to enhance the smoothness of the feature distribution. The logarithmically transformed features are then input into an ANN for feature learning and representation, realizing the mapping of sidechain structure information to a low-dimensional dense vector, and finally generating a sidechain feature vector V. sidechain This is used for subsequent feature concatenation with skeleton features and model training.

[0126] The ANN (Artificial Neural Network) is a two-layer fully connected neural network with Dropout, specifically including:

[0127] The first-layer linear perceptron has an input dimension of 6 (corresponding to the 6-dimensional physical descriptor of the side chain) and an output dimension of 64. It is used to map the low-dimensional physical descriptor to the high-dimensional feature space and learn the complex interaction relationship between the physical features of the side chain.

[0128] The first Dropout layer: The Dropout probability is set to 0.3 to suppress overfitting during small sample training.

[0129] ReLU activation layer: Introduces nonlinear transformation to improve the model's ability to fit nonlinear relationships;

[0130] The second-layer linear perceptron has an input dimension of 64 and an output dimension of 64. It is used to refine and integrate high-dimensional features, remove redundant information, and enhance effective features.

[0131] The second Dropout layer further suppresses overfitting;

[0132] The second ReLU activation layer: The final output is a fixed-dimensional (64-dimensional) sidechain latent feature vector V. sidechain .

[0133] In a preferred embodiment, the following characteristics are evaluated: number of heavy atoms, lipid-water partition coefficient (LogP), topological polar surface area (TPSA), main chain length, and branch chain length. Logarithmic transformation is performed on the data used as input to the ANN.

[0134] Finally, the conjugate skeleton of P(NDI2OD-T2) is mapped to a graph object (DataObject), whose mathematical representation and physical meaning are as follows:

[0135] Node feature matrix This tensor represents the 32 atomic nodes contained in the skeleton. Each row is a 17-dimensional feature vector;

[0136] Topology connection matrix This matrix describes the interatomic connections within the framework. The value 74 indicates that there are 37 chemical bonds within the framework of this repeating unit.

[0137] Edge feature matrix There are 74 directed edges, each with 7-dimensional features.

[0138] For the two 2-octyldodecyl side chains carried by P(NDI2OD-T2), after weighted aggregation and logarithmic transformation, the final 6-dimensional input vector is: [1.3222,0.9496,0.0,1.3010,0.3010,0.4737].

[0139] S5. Heterogeneous Fusion and Prediction: The feature vector V of the skeleton global graph is then used for... backbone With sidechain feature vector V sidechain Dimensional concatenation is performed to obtain a fused feature vector; the fused feature vector is input into a fusion perceptron, which then outputs the predicted thermoelectric properties of the N-type polymer. .

[0140] In the heterogeneous fusion and prediction, the skeleton global graph feature vector is... With sidechain feature vectors Dimensional concatenation yields the fused feature vector. :

[0141] ;

[0142] Fusion feature vectors The system enters a fusion perceptron, which includes a ReLU activation layer, a Dropout regularization layer, and a linear mapping layer.

[0143] The ReLU activation layer uses the ReLU activation function for nonlinear interaction, simulating the physical interaction between the backbone and sidechains; the probability of the Dropout regularization layer is 0.3-0.6, which is used to suppress overfitting during small sample training.

[0144] The linear mapping layer outputs predicted values ​​of the thermoelectric properties of N-type polymers. :

[0145] ;

[0146] in It is the final fused feature vector obtained by processing the ReLU activation layer and Dropout regularization layer of the fusion perceptron.

[0147] The fusion perceptron comprises a three-layer structure: a ReLU activation layer, a Dropout regularization layer, and a linear mapping layer.

[0148] ReLU activation layer: Performs nonlinear transformation on the fused feature vector to simulate the physical interaction between the skeleton and sidechains, enhancing the model's ability to express structure-performance relationships;

[0149] Dropout regularization layer: suppresses overfitting with a random deactivation probability of 0.3-0.6, and improves the stability of the model under small sample training.

[0150] Linear mapping layer: The final fused features after nonlinear and regularization processing are mapped to predicted values ​​of thermoelectric properties of N-type polymers, so as to achieve accurate output of thermoelectric properties.

[0151] In a further preferred embodiment, the fusion perceptron includes a conductivity prediction weight branch, which comprises a first branch ReLU activation layer and a first branch linear mapping layer. The first branch ReLU activation layer is used to process the input fusion feature vector. Perform nonlinear transformation; the first branch linear mapping layer outputs conductivity. Predicted value ,

[0152] ,in, This is the weight matrix for conductivity prediction. This is the bias vector for predicting conductivity;

[0153] The fusion perceptron includes a power factor prediction weight branch, which comprises a second branch ReLU activation layer and a second branch linear mapping layer. The second branch ReLU activation layer is used to process the input fusion feature vector. Perform a nonlinear transformation; the second branch linear mapping layer outputs the predicted power factor PF. , ,in, The weight matrix for power factor prediction. This is the bias vector for power factor prediction.

[0154] This invention effectively eliminates the feature dilution effect of side-chain atomic information on the electronic features of the conjugated backbone in existing technologies through a deep learning architecture that decouples the backbone and side chains. The prediction model includes GINE, ANN, and a fused perceptron, enabling it to detect minute atomic substitutions in the conjugated main chain and subtle shifts in side chain branching points. As a prediction tool for N-type polymer systems, this invention effectively predicts the thermoelectric properties of N-type polymers, significantly reducing the expensive experimental costs and risks associated with traditional R&D models such as "blind trial and error."

[0155] In a further optimization scheme, the heterogeneous fusion and prediction step also includes an iterative optimization step, including:

[0156] S11. Supervision label determination: The target value y is obtained by performing a logarithmic transformation on the experimentally measured thermoelectric properties of the N-type polymer.

[0157] Data acquisition and correction:

[0158] Experimental data on N-type polymers in an N-DMBI doped system were extracted from existing academic literature in the field of organic thermoelectrics. The data covers electrical conductivity (…). The values ​​of the power factor (PF) and the corresponding repeating unit structure of the polymer are mapped using the following logarithmic transformation formula:

[0159] ;in, The experimentally measured thermoelectric properties of N-type polymers are logarithmically transformed to make the distribution of the target value y approach a normal distribution, thereby improving the training stability of the prediction model.

[0160] S12, Loss Function Calculation: Calculate the predicted thermoelectric properties of N-type polymers. The predicted thermoelectric properties of the N-type polymer are calculated by comparing the predicted value with the target value y. The loss function between the target value y and the target value y, wherein the loss function uses the mean squared error;

[0161] S13. Gradient Calculation: Based on the backpropagation algorithm, starting from the loss function, backtrack and calculate the gradients of all learnable parameters in GINE, ANN and fused perceptron.

[0162] S14. Parameter Update: Following the gradient descent strategy, update all learnable parameters in the direction of gradient descent. The parameter update formula is: w represents the parameter to be updated, and η represents the learning rate. The gradient corresponding to the parameter w. These are the updated parameters;

[0163] Repeat steps S12, S13, and S14 until the loss function value converges to the preset threshold, thus completing the iterative optimization.

[0164] Comparative experimental design and results analysis:

[0165] Group A input data:

[0166] In constructing the input data for the comparative model, the N-type polymer repeating units were first represented as the string "SMILES," and molecular descriptors were calculated using the Mordred tool to obtain initial features of over 1800 dimensions. Since the original descriptors were numerous and contained redundant and irrelevant features, a recursive feature elimination cross-validation method (RFECV) was further employed for feature selection.

[0167] Specifically, RFECV uses a random forest regressor as the base learner, sets the step size to 1, meaning that one feature with the lowest importance is deleted in each iteration; the minimum number of features to retain is set to 10; the cross-validation method uses five-fold cross-validation, and random shuffling is performed during the partitioning process, with a random seed of 42; this method continuously trains the model iteratively, calculates feature importance and removes the least important features, while recording the cross-validation scores under different feature counts, and finally selects the feature subset corresponding to the optimal cross-validation performance as the model input.

[0168] After the above screening, the power factor (PF) prediction task ultimately retained 17 molecular descriptors: SpDiam_A, AATTS0i, ATSC8Z, ATSC3v, AATSC4s, AATSC6p, MATS3Z, MATS6se, GATS8dv, GATS5d, GATS8s, BCUTZ-1l, Xch-5d, Xch-5dv, SMR_VSA7, AMID_N, and n6Ring. The conductivity (σ) prediction task ultimately retained 16 molecular descriptors: SpDiam_A, AATTS0Z, AATTS5i, AATSC5Z, AATSC6se, AATSC6p, MATS3Z, MATS6se, GATS7dv, GATS2d, GATS5d, GATS6s, GATS6i, BCUTZ-1l, PEOE_VSA7, and AMID_N.

[0169] Finally, the optimal combination of descriptors obtained through screening is used as input data for the traditional machine learning comparison model, and is used for subsequent regression prediction of power factor and conductivity.

[0170] Parameter settings for each model in Group A:

[0171] (41) Random Forest (Group A)

[0172] For the random forest comparison model, a Bagging-based ensemble learning approach was adopted, integrating multiple decision tree weak learners to reduce model variance and suppress overfitting. During parameter optimization, the search range for the number of decision trees was set to 50 to 500, the maximum tree depth to 5 to 30, the minimum number of samples required for internal node splits to 2 to 10, the minimum number of samples in a leaf node to 1 to 5, and the maximum number of features was selected between square root and base-2 logarithmic methods. The final optimization results show that for the power factor prediction task, the optimal parameter combination is: 135 decision trees, 16 maximum depth, 5 minimum number of split samples, 1 minimum number of leaf node samples, and logarithmic maximum number of features; for the conductivity prediction task, the optimal parameter combination is: 50 decision trees, 11 maximum depth, 2 minimum number of split samples, 2 minimum number of leaf node samples, and square root maximum number of features.

[0173] (42) Limiting gradient hints (Group A)

[0174] For the XGBoost comparison model, the parameter optimization range was set as follows: learning rate 0.01 to 0.3, number of trees 50 to 500, maximum depth 5 to 30, split threshold gamma 0 to 2, minimum child node weight 1 to 10, row sampling ratio and column sampling ratio 0.1 to 1.0, L1 regularization coefficient 1 to 5, and L2 regularization coefficient 1 to 10. The optimal parameters for the power factor prediction task are: learning rate 0.0287, number of trees 500, maximum depth 17, gamma 0, minimum child node weight 10, row sampling ratio 0.5751, column sampling ratio 0.5498, L1 regularization coefficient 1.0, and L2 regularization coefficient 10.0. The optimal parameters for the conductivity prediction task are: learning rate 0.1802, number of trees 128, maximum depth 14, gamma 0.0264, minimum child node weight 8, row sampling ratio 0.5287, column sampling ratio 0.5744, L1 regularization coefficient 1.2686, and L2 regularization coefficient 6.2692.

[0175] (43) Support Vector Regression (Group A)

[0176] For the support vector regression comparison model, the parameter optimization range was set as follows: penalty coefficient 0.1–1000, pipe width 0.001–1.0, kernel function parameter 0.0001–10.0, and radial basis function (RBF) kernel type. The optimal parameters for the power factor prediction task were: penalty coefficient 40.4593, pipe width 0.7692, and kernel function parameter 0.0001804; the optimal parameters for the conductivity prediction task were: penalty coefficient 71.9465, pipe width 1.0, and kernel function parameter 0.0008902, with the kernel function type being RBF in both cases.

[0177] (44) -Support Vector Regression (Group A)

[0178] for - Support Vector Regression comparison model, parameter optimization range set as follows: parameter nu = 0.1–0.8, penalty coefficient = 0.1–1000, kernel function parameter = 10. −4 ~10.0, the kernel function type uses a radial basis kernel. The final optimal parameters for the power factor prediction task are: With a value of 0.3171, a penalty coefficient of 48.7753, and a kernel function parameter of 0.0007780, the optimal parameters for the conductivity prediction task are: The value is set to 0.6619, the penalty coefficient is 177.6577, the kernel function parameter is 0.0007234, and the kernel function type is radial basis kernel.

[0179] (45) Linear Support Vector Regression (Group A)

[0180] For the linear support vector regression comparison model, the key hyperparameters include the penalty coefficient and the pipe width. The search range for the penalty coefficient is set to 0.01 to 1000, and the search range for the pipe width is set to 0.001 to 10.0. The optimal parameters for the power factor prediction task are: penalty coefficient 0.2411, pipe width 0.3761; and for the conductivity prediction task, the optimal parameters are: penalty coefficient 0.1186, pipe width 0.6922.

[0181] (46) Decision Tree (Group A)

[0182] For the decision tree comparison model, the parameter optimization range is set as follows: maximum depth 3–20, minimum number of split samples 2–20, minimum number of leaf nodes 1–10, maximum number of features selected from three options: square root, base-2 logarithm, and no restriction, loss function criterion selected between squared error and absolute error, and pruning parameter values ​​range as follows. The optimal parameters for the power factor prediction task are: maximum depth 6, minimum number of split samples 5, minimum number of leaf node samples 9, maximum number of features using a logarithmic approach, loss function of squared error, and pruning parameter of 0.00336. The optimal parameters for the conductivity prediction task are: maximum depth 3, minimum number of split samples 18, minimum number of leaf node samples 10, maximum number of features without restriction, loss function of squared error, and pruning parameter of 0.03427.

[0183] (47) Ordinary least squares regression (Group A)

[0184] For the multiple linear regression comparison model, since it is an analytical solution model, the regression coefficients can be uniquely determined directly by the least squares criterion after given the input feature matrix and the target output, so there is no need to optimize the hyperparameters.

[0185] (48) Partial Least Squares Regression (Group A)

[0186] For the partial least squares regression comparison model, the key hyperparameter is the number of principal components, with a search range of 1 to 15. Ultimately, the optimal number of principal components for the power factor prediction task is 2, and for the conductivity prediction task, it is 4.

[0187] (49) Ridge regression (Group A)

[0188] For the ridge regression comparison model, the key hyperparameters include the L2 penalty term coefficient and the solver type. The search range for the L2 penalty term coefficient is set from 0.01 to 100.0, and the solver is selected from auto, svd, cholesky, and lsqr. The optimal parameters for the power factor prediction task are an L2 penalty term coefficient of 29.4365 and an svd solver; the optimal parameters for the conductivity prediction task are an L2 penalty term coefficient of 21.5409 and an svd solver.

[0189] (50) Lasso Regression (Group A)

[0190] For the lasso regression comparison model, the key hyperparameter is the L1 penalty term coefficient, and the search range is set to... Up to 10.0. The optimal parameter for the power factor prediction task is 0.1058, and the optimal parameter for the conductivity prediction task is 0.05993.

[0191] (51) Elastic network regression (Group A)

[0192] For the ElasticNet comparison model, key hyperparameters include total regularization strength and the proportion of L1 regularization terms. The search range for total regularization strength is set to... Up to version 10.0, the search range for the L1 regularization ratio was set to 0.01 to 1.0. The optimal parameters for the power factor prediction task were: total regularization strength 0.1736, L1 regularization ratio 0.6707; and for the conductivity prediction task, the optimal parameters were: total regularization strength 0.2070, L1 regularization ratio 0.01.

[0193] (52) K-Nearest Neighbor Algorithm (Group A)

[0194] For the K-Nearest Neighbors algorithm comparison model, its key hyperparameters include the number of neighbors, weighting strategy, distance metric, and leaf node size. The search range for the number of neighbors is set to 2 to 30. The weighting strategy is chosen between uniform weighting and distance-weighted weighting. The distance metric is chosen between 1 and 2, corresponding to Manhattan distance and Euclidean distance, respectively. The search range for leaf node size is set to 20 to 60. The optimal parameters for the power factor prediction task are: 14 neighbors, distance-weighted weighting, a distance metric of 2, and a leaf node size of 52. The optimal parameters for the conductivity prediction task are: 2 neighbors, uniform weighting, a distance metric of 1, and a leaf node size of 20.

[0195] (53) Gaussian process regression (Group A)

[0196] For the Gaussian process regression comparison model, the key hyperparameter is the initial noise term, and the search range is set as follows: Up to 10. The optimal parameters for the power factor prediction task are ultimately 0.00563, and the optimal parameters for the conductivity prediction task are 1.0023 × 10⁻⁶. .

[0197] (54) Artificial Neural Networks (Group A)

[0198] For the artificial neural network comparison model, its key hyperparameters include hidden layer structure, activation function, solver, L2 regularization coefficient, learning rate update method, and initial learning rate. Specifically, the hidden layer structure is selected from a preset combination of single-layer, two-layer, and three-layer networks; the activation function is selected from ReLU, tanh, and logistic; the solver is selected from LBFGS and Adam; and the search range for the L2 regularization coefficient is set to [value missing]. Up to 1, the learning rate update method can be selected from constant and adaptive, and the initial learning rate search range is set to [value missing]. to The optimal parameters for the power factor prediction task are: a single-layer hidden layer with 32 neurons, a logistic activation function, a lbfgs solver, an L2 regularization coefficient of 0.5501, a constant learning rate, and an initial learning rate of 0.1. The optimal parameters for the conductivity prediction task are: a single-layer hidden layer with 16 neurons, a logistic activation function, a lbfgs solver, an L2 regularization coefficient of 1.0, a constant learning rate, and an initial learning rate of [missing value]. .

[0199] Group B input data:

[0200] Control group B uses a graph neural network, which does not decouple the conjugated backbone and side chains of the molecular structure, but instead uses the entire molecular graph as the input to the graph neural network to obtain the prediction results.

[0201] Group B parameter settings

[0202] For the designed graph neural network comparison model, its key hyperparameters include the hidden layer dimension of the graph neural network, the dropout ratio of the graph branch, the hidden dimension of the prediction layer, the dropout ratio of the prediction layer, the learning rate, and the weight decay coefficient. Specifically, the optimal parameters for the power factor prediction task are set as follows: hidden layer dimension of the graph neural network is 128, graph branch dropout is 0, hidden layer dimension of the prediction layer is 128, prediction layer dropout is 0.1, learning rate is 0.0001, and weight decay coefficient is... The optimal parameters for the conductivity prediction task are set as follows: the hidden layer dimension of the graph neural network is 48, the graph branch dropout is 0.0571, the hidden layer dimension of the prediction layer is 74, the prediction layer dropout is 0.1142, the learning rate is 0.00216, and the weight decay coefficient is 0.01.

[0203] Error index:

[0204] First, define the following symbols: They represent the first The true value of the nth sample, the nth Calculate the following metrics using the predicted value of each sample, the total number of samples, and the average of all true values.

[0205] Mean Absolute Error (MAE):

[0206] ;

[0207] Mean Squared Error (MSE):

[0208] ;

[0209] Root Mean Square Error (RMSE):

[0210] ;

[0211] Coefficient of Determination (R-Squared / Coefficient of Determination) 2 ):

[0212] .

[0213] The final experimental results are as follows:

[0214] Table 1: Comparative Experimental Results of Power Factor Prediction Task

[0215]

[0216] Table 2: Comparative Experiment Results of Conductivity Prediction Task

[0217]

[0218] Results analysis:

[0219] This embodiment focuses on two core indicators of N-type organic thermoelectric materials—power factor ( ) and conductivity ( A full model comparison was conducted. Based on the LOOCV validation results of 125 samples (Tables 1 and 2), the present invention demonstrates significant technical advantages across all evaluation dimensions. Experiments show that the present invention achieves significant advantages in power factor (R... 2 =0.6514) and conductivity (R 2 In the prediction of conductivity (=0.6540), compared with the best-performing Random Forest (RF) and Limiting Gradient Boosting (XGB) in Group A, the coefficient of determination of this invention is significantly improved. This proves that the "skeleton-sidechain decoupling" architecture can more fundamentally reveal the relationship between the microscopic topology and macroscopic transport properties of polymers, and has predictive universality across indicators. Data comparison shows that the unprocessed baseline graphical neural network (Group B GNN) is inferior in conductivity prediction. The MSE is only 0.4716, far lower than the 0.6540 of this invention. This result strongly supports the improvement logic of this invention. In terms of the MSE and MAE metrics, which measure error, this invention demonstrates stronger robustness. Taking conductivity prediction as an example, the MSE of this invention (0.7002) is about 35% lower than that of GNN (1.0694). This invention effectively avoids the overfitting trap commonly found in models on small sample tasks, providing performance predictions for the rational design of N-type thermoelectric materials.

[0220] The above description is merely a specific embodiment of the present invention. Various examples and illustrations do not constitute a limitation on the substantive content of the present invention. Those skilled in the art can modify or transform the specific embodiments described above after reading the specification without departing from the essence and scope of the invention.

Claims

1. A method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks, characterized in that, The method includes: S1. Data preprocessing: The molecular formula of N-type polymers is processed into SMILES sequences. If there are dummy atoms representing polymer sites, C atoms are used to replace the dummy atoms. S2. Structural decoupling: Identify the SMILES sequence of N-type polymers and sever the bonds connecting the backbone and side chains. The backbone is defined as the conjugated system that dominates the transport of charge carriers in the molecule; the side chain is defined as an aliphatic chain segment that regulates the solubility or polarity of the molecule. The structural decoupling is achieved using the RDKit tool, which is used to identify and cleave the bonds connecting the backbone and side chains. The specific method is as follows: First, the SMILES sequence of the N-type polymer is imported into the RDKit tool. The molecular topology corresponding to the SMILES sequence is analyzed, and the backbone belonging to the conjugated system and the side chains belonging to the aliphatic segments are identified. The bonds connecting the backbone and side chains are located and identified. The bonds are cleaved using a chemical bond cleaving algorithm to separate the backbone and side chains. After cleaving, the backbone is represented as a topological map composed of atoms and chemical bonds, and the side chains are represented as a set of molecular chain segments. S3. Backbone Feature Extraction: Each atom in the backbone is encoded with node features to form encoded atom nodes; each chemical bond in the backbone is encoded with edge features to form encoded chemical bond edges; based on the encoded atom nodes and chemical bond edges, a molecular graph structure of the N-type polymer is constructed, and the molecular graph structure of the N-type polymer is input into the GINE graph neural network, and the global graph feature vector Vbackbone of the backbone is extracted by the GINE graph neural network. In the skeleton feature extraction, the node feature encoding includes element type One-hot encoding, geometric coordination environment One-hot encoding, physical property continuous variable normalization feature and local topological constraint feature. The local topological constraint feature includes: hydrogenation number normalization value, aromaticity Boolean feature, intra-loop attribute Boolean feature, and side chain connection site identifier. In the skeleton feature extraction, the edge feature encoding includes bond type One-hot encoding and topological attribute Boolean features; the bond type is a single bond, double bond, triple bond, aromatic bond, or other bond type, and the topological attribute includes whether the chemical bond is inside a ring and whether it has conjugation properties; S4. Sidechain Feature Extraction: Extract physicochemical features from the sidechain to form sidechain features, aggregate multiple sidechain features to obtain sidechain aggregate features; perform logarithmic processing on the sidechain aggregate features and input them into an ANN artificial neural network to generate a sidechain feature vector Vsidechain. In the side chain feature extraction, the physicochemical features include the number of heavy atoms, lipid-water partition coefficient, topological polar surface area of ​​molecules, main chain length, branch chain length, and the proportion of branch point positions. S5. Heterogeneous Fusion and Prediction: The global graph feature vector Vbackbone and the sidechain feature vector Vsidechain are concatenated dimensionally to obtain a fused feature vector; the fused feature vector is input into a fusion perceptron, which outputs the predicted thermoelectric properties of the N-type polymer. .

2. The method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks according to claim 1, characterized in that, The GINE graph neural network employs the GINEConv operator with residual connections; ; ; in, Represents atoms In the Feature vectors after layer iteration Represents atoms The state at level k-1, Represents atoms Neighboring atoms The state at level k-1 Indicates with atoms The set of all directly connected atoms It is the first The multilayer perceptron of the GINE graph neural network corresponding to each layer. It is a learnable scalar parameter; This represents the atoms in the (k-1)th layer. With neighboring atoms The coding features corresponding to the chemical bonds between them.

3. The method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks according to claim 1, characterized in that, In the side chain feature extraction, the aggregation of multiple side chain features adopts a weighted aggregation method based on the proportion of heavy atoms, and the aggregation formula is as follows: ; in, For the first The number of heavy atoms in each side chain, This is the 6-dimensional feature vector of the sidechain; This is a side-chain aggregation characteristic.

4. The method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks according to claim 1, characterized in that, The ANN artificial neural network includes two linear perceptron layers. The first linear perceptron layer is used to map physical descriptors to a high-dimensional space, and the second linear perceptron layer is used to extract latent features. Each of the two linear perceptron layers is equipped with a Dropout layer.

5. The method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks according to claim 1, characterized in that, In the heterogeneous fusion and prediction, the skeleton global graph feature vector is... With sidechain feature vectors Dimensional concatenation yields the fused feature vector. : ; Fusion feature vectors The system enters a fusion perceptron, which includes a ReLU activation layer, a Dropout regularization layer, and a linear mapping layer. The ReLU activation layer uses the ReLU activation function for nonlinear interaction, simulating the physical interaction between the backbone and sidechains; the probability of the Dropout regularization layer is 0.3-0.6, which is used to suppress overfitting during small sample training. The linear mapping layer outputs predicted values ​​of the thermoelectric properties of N-type polymers. : ; in It is the final fused feature vector obtained by processing the ReLU activation layer and Dropout regularization layer of the fusion perceptron.

6. The method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks according to claim 1, characterized in that, The fusion perceptron includes a conductivity prediction weight branch, which comprises a first branch ReLU activation layer and a first branch linear mapping layer. The first branch ReLU activation layer is used to process the input fusion feature vector. Perform nonlinear transformation; the first branch linear mapping layer outputs conductivity. Predicted value , ,in, This is the weight matrix for conductivity prediction. This is the bias vector for predicting conductivity; The fusion perceptron includes a power factor prediction weight branch, which comprises a second branch ReLU activation layer and a second branch linear mapping layer. The second branch ReLU activation layer is used to process the input fusion feature vector. Perform a nonlinear transformation; the second branch linear mapping layer outputs the predicted power factor PF. , ,in, The weight matrix for power factor prediction. This is the bias vector for power factor prediction.

7. The method for predicting the thermoelectric properties of N-type polymers based on structural decoupling and neural networks according to claim 1, characterized in that, The heterogeneous fusion and prediction step also includes an iterative optimization step, including: S11. Supervision label determination: The target value y is obtained by performing a logarithmic transformation on the experimentally measured thermoelectric properties of the N-type polymer. S12, Loss Function Calculation: Calculate the predicted thermoelectric properties of N-type polymers. The predicted thermoelectric properties of the N-type polymer are calculated by comparing the predicted value with the target value y. The loss function between the target value y and the target value y, wherein the loss function uses the mean squared error; S13. Gradient Calculation: Based on the backpropagation algorithm, starting from the loss function, backtrack and calculate the gradients of all learnable parameters in GINE, ANN and fused perceptron. S14. Parameter Update: Following the gradient descent strategy, update all learnable parameters in the direction of gradient descent. The parameter update formula is: w represents the parameter to be updated, and η represents the learning rate. The gradient corresponding to the parameter w. These are the updated parameters; Repeat steps S12, S13, and S14 until the loss function value converges to the preset threshold, thus completing the iterative optimization.