A glass transition temperature prediction model determination method, application method and device

CN122551990APending Publication Date: 2026-08-11EAST CHINA UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

与此同时,传统基于单一结构式的表征方式难以有效刻画共聚体系在相同配比条件下由于不同序列排列引起的性能波动,导致模型对序列差异不敏感,难以实现对共聚聚酰胺玻璃化转变温度的准确预测

Benefits of technology

本申请提供了一种玻璃化转变温度预测模型确定方法、应用方法及装置,通过获取原始数据集,解决了聚酰胺、聚酰亚胺两类高分子材料样本与对应玻璃化转变温度基础数据缺失的问题,实现了建模所需原料样本数据与物性标签的完备归集;通过将样本重复单元扩展为预设长度长链SMILES表示,解决了短链结构式无法真实体现高分子长链聚合结构特征的问题,实现了聚合态分子链空间结构与键合特征的精准数字化表征;通过长链SMILES转换为高维第一描述符向量,解决了分子结构式无法被机器学习模型直接识别运算的问题,实现了高分子微观结构特征向量化提取;通过对高维描述符向量开展降维处理,解决了高维特征冗余、运算效率低、模型易过拟合的问题,实现了核心有效结构特征精简留存与特征维度优化;通过结合降维后描述符向量与温度数据训练高斯过程回归模型,解决了传统实验测试玻璃化转变温度周期长、成本高、效率低的问题,实现了玻璃化转变温度预测模型的构建。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551990A_ABST
    Figure CN122551990A_ABST
Patent Text Reader

Abstract

The application discloses a glass transition temperature prediction model determination method, an application method and a device, relates to the technical field of polymer material performance prediction and machine learning aided design, and comprises the following steps: acquiring an original data set, wherein the original data set comprises a plurality of single-component polyamide samples and corresponding glass transition temperature data, and a plurality of single-component polyimide samples and corresponding glass transition temperature data; expanding the repeating units of each single-component polyamide sample and each single-component polyimide sample into long-chain SMILES representation of a preset length; converting each long-chain SMILES into a corresponding descriptor vector higher than a preset dimension, and performing dimension reduction processing; training a Gaussian process regression model based on the descriptor vector after the dimension reduction processing and the corresponding glass transition temperature data, and obtaining a glass transition temperature prediction model. The glass transition temperature prediction model can be used to realize the prediction of the glass transition temperature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of polymer material performance prediction and machine learning-aided design technology, and in particular to a method, application method and apparatus for determining a glass transition temperature prediction model. Background Technology

[0002] Polyamide materials possess excellent thermal and mechanical properties, finding wide application in automotive, electronics, electrical, textile, and medical fields. For copolymer polyamide systems, monomer type, composition ratio, and chain segment sequence all influence the glass transition temperature (GVT). Due to the diverse monomer combinations, wide ratio ranges, and complex sequence arrangements in copolymer polyamides, experimental data acquisition typically relies on individual synthesis and testing, resulting in long cycles and high costs. This limits the amount of direct GVT data available for modeling, hindering the development of high-precision prediction models. In contrast, structural and performance data for single-component polyamides are easier to obtain, with relatively abundant data accumulation. Currently, there is a lack of methods to effectively predict the GVT of multi-component copolymer polyamides using single-component polyamide data combined with the composition ratio and sequence information of the copolymer system. Furthermore, traditional characterization methods based on a single structural formula struggle to effectively depict performance fluctuations caused by different sequence arrangements under the same ratio conditions, leading to model insensitivity to sequence differences and hindering accurate prediction of the GVT of copolymer polyamides.

[0003] Therefore, it is necessary to provide a method that can explicitly characterize sequence information and predict the glass transition temperature of multicomponent copolyamides, and also a method that can predict the glass transition temperature of multicomponent copolyimides. Summary of the Invention

[0004] The purpose of this application is to provide a method, application method, and apparatus for determining a glass transition temperature prediction model, which can be used to predict the glass transition temperature of multi-component copolyamides and multi-component copolyimides.

[0005] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a method for determining a glass transition temperature prediction model, including: Obtain the original dataset; the original dataset includes: several one-component polyamide samples and corresponding glass transition temperature data, and several one-component polyimide samples and corresponding glass transition temperature data; the one-component polyamide samples and the one-component polyimide samples are both made by copolymerization of one or more monomers.

[0006] The repeating units of each single-component polyamide sample and each single-component polyimide sample are extended into long-chain SMILES of a preset length.

[0007] Each long chain SMILES is transformed into a corresponding first descriptor vector with a dimension higher than the preset dimension.

[0008] The first descriptor vector is subjected to dimensionality reduction processing to obtain the dimensionality-reduced first descriptor vector.

[0009] The Gaussian process regression model is trained based on the first descriptor vector after dimensionality reduction and the corresponding glass transition temperature data to obtain a glass transition temperature prediction model.

[0010] Secondly, this application provides a method for applying a glass transition temperature prediction model, including: Obtain the monomer composition and target ratio of the sample to be predicted.

[0011] Based on the monomer composition and target ratio of the sample to be predicted, several long chains of SMILES with different sequence arrangements are randomly generated.

[0012] The long chains SMILES with different sequence arrangements are transformed into corresponding second descriptor vectors with a higher dimension than a preset number.

[0013] The second descriptor vector is then subjected to dimensionality reduction processing to obtain the dimensionality-reduced second descriptor vector.

[0014] The second descriptor vector after dimensionality reduction is input into the glass transition temperature prediction model to obtain the glass transition temperature prediction values ​​corresponding to different sequence samples; the glass transition temperature prediction model is a model obtained based on the glass transition temperature prediction model determination method described above.

[0015] The predicted glass transition temperatures of the different sequence samples are screened and statistically analyzed to obtain the glass transition temperature of the sample to be predicted under the target ratio.

[0016] Thirdly, this application provides a device for determining a glass transition temperature prediction model, comprising: The original dataset acquisition module is used to acquire the original dataset, which includes: several one-component polyamide samples and corresponding glass transition temperature data, and several one-component polyimide samples and corresponding glass transition temperature data; the one-component polyamide samples and the one-component polyimide samples are both made by copolymerization of one or more monomers.

[0017] An extension module is used to extend the repeating units of each single-component polyamide sample and each single-component polyimide sample into long-chain SMILES representations of a preset length.

[0018] The first transformation module is used to transform each long chain SMILES into a corresponding first descriptor vector with a higher dimension than the preset dimension.

[0019] The first dimensionality reduction processing module is used to perform dimensionality reduction processing on the first descriptor vector to obtain the dimensionality-reduced first descriptor vector.

[0020] The training module is used to train the Gaussian process regression model based on the first descriptor vector after dimensionality reduction and the corresponding glass transition temperature data, so as to obtain the glass transition temperature prediction model.

[0021] Fourthly, this application provides an application device for a glass transition temperature prediction model, comprising: The test data acquisition module is used to acquire the monomer composition and target ratio of the sample to be predicted.

[0022] The sequence generation module is used to randomly generate several long chain SMILES representations with different sequence arrangements based on the monomer composition and target ratio of the sample to be predicted.

[0023] The second conversion module is used to convert the long chains of SMILES with different sequence arrangements into corresponding second descriptor vectors with a higher dimension than a preset dimension.

[0024] The second dimensionality reduction processing module is used to perform dimensionality reduction processing on the second descriptor vector to obtain the dimensionality-reduced second descriptor vector.

[0025] The prediction module is used to input the second descriptor vector after dimensionality reduction into the glass transition temperature prediction model to obtain the glass transition temperature prediction values ​​corresponding to different sequence samples; the glass transition temperature prediction model is a model obtained based on the glass transition temperature prediction model determination method described above.

[0026] The screening and statistics module is used to screen and statistically analyze the predicted glass transition temperature values ​​corresponding to the different sequence samples, so as to obtain the glass transition temperature of the sample to be predicted under the target ratio.

[0027] Fifthly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the glass transition temperature prediction model determination method or the glass transition temperature prediction model application method described above.

[0028] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the glass transition temperature prediction model determination method or the glass transition temperature prediction model application method described above.

[0029] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the glass transition temperature prediction model determination method or the glass transition temperature prediction model application method described above.

[0030] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a method, application method, and apparatus for determining a glass transition temperature prediction model. By acquiring the original dataset, it solves the problem of missing basic data on the glass transition temperatures of two types of polymer materials, polyamide and polyimide, achieving complete collection of raw material sample data and property labels required for modeling. By expanding the sample repeating units into long-chain SMILES representations of a preset length, it solves the problem that short-chain structures cannot truly reflect the polymer long-chain polymerization structure characteristics, achieving accurate digital characterization of the spatial structure and bonding features of polymeric molecular chains. By converting long-chain SMILES into high-dimensional first descriptor vectors, it solves the problem that molecular structures cannot be directly recognized and calculated by machine learning models, achieving vectorized extraction of polymer microstructure features. By performing dimensionality reduction processing on the high-dimensional descriptor vectors, it solves the problems of high-dimensional feature redundancy, low computational efficiency, and easy model overfitting, achieving simplification and retention of core effective structural features and optimization of feature dimensions. By combining the dimensionality-reduced descriptor vectors with temperature data to train a Gaussian process regression model, it solves the problems of long cycle, high cost, and low efficiency of traditional experimental testing of glass transition temperature, achieving the construction of a glass transition temperature prediction model. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is an application environment diagram of a method for determining a glass transition temperature prediction model and its application method in one embodiment of this application. Figure 2 A flowchart illustrating a method for determining a glass transition temperature prediction model according to an embodiment of this application; Figure 3A flowchart illustrating an application method for a glass transition temperature prediction model provided in an embodiment of this application; Figure 4 A schematic diagram of the functional modules of a glass transition temperature prediction model determination device provided in an embodiment of this application; Figure 5 A schematic diagram of the functional modules of a glass transition temperature prediction model application device provided in an embodiment of this application; Figure 6 This is a schematic diagram showing the data distribution in a dataset corresponding to the glass transition temperatures of polyamide and polyimide provided in an embodiment of this application. Figure 7 This is a schematic diagram illustrating the expansion of corresponding repeating units of single-component polyamide and polyimide into long chains of a predetermined length, according to an embodiment of this application. Figure 8 A schematic diagram of multiple long-chain SMILES representing different sequence arrangements is provided for an embodiment of this application for a copolyamide to be predicted, based on the monomer composition and ratio. Figure 9 A schematic diagram illustrating the prediction accuracy of a glass transition temperature model provided in an embodiment of this application; Figure 10 A schematic diagram illustrating the sorting results of the predicted values ​​provided in an embodiment of this application; Figure 11 A schematic diagram of the chemical structure of PA10T provided in an embodiment of this application; Figure 12 A schematic diagram of the chemical structure of PA56 provided in an embodiment of this application; Figure 13 A comparison chart of experimental and predicted values ​​for verifying the accuracy of a model provided in an embodiment of this application; Figure 14 A schematic diagram of the chemical structure of BFA_TERE provided in an embodiment of this application; Figure 15 A schematic diagram of the chemical structure of BFA_ISO provided in an embodiment of this application; Figure 16 A comparison chart of experimental and predicted values ​​for verifying the accuracy of a model provided in another embodiment of this application; Figure 17 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0034] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] The glass transition temperature prediction model determination method provided in this application embodiment can be applied to, for example... Figure 1 The application environment shown is illustrated. Terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on another server. Terminal 102 can send the acquired raw dataset to server 104. The raw dataset includes: several one-component polyamide samples and their corresponding glass transition temperature data, and several one-component polyimide samples and their corresponding glass transition temperature data. Both the one-component polyamide and one-component polyimide samples are made from one or more monomers through a copolymerization reaction. Upon receiving the raw dataset, server 104 expands the repeating units of each one-component polyamide and one-component polyimide sample into long-chain SMILES of a preset length. Each long-chain SMILES is converted into a corresponding first descriptor vector with a dimension higher than a preset value. The first descriptor vector is then dimensionality-reduced to obtain a dimensionality-reduced first descriptor vector. Based on the dimensionality-reduced first descriptor vector and the corresponding glass transition temperature data, a Gaussian process regression model is trained to obtain a glass transition temperature prediction model. Server 104 can then feed back the obtained glass transition temperature prediction model to terminal 102. In addition, in some embodiments, the method for determining the glass transition temperature prediction model can also be implemented by the server 104 or the terminal 102 separately. For example, the terminal 102 can directly process the original data and obtain the glass transition temperature prediction model based on the processed original data. Alternatively, the server 104 can obtain the original data from the data storage system, process it, and obtain the glass transition temperature prediction model based on the processed original data.

[0036] The glass transition temperature prediction model and application method provided in this application embodiment can be applied to, for example... Figure 1The application environment shown is illustrated. Terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on another server. Terminal 102 can send the monomer composition and target ratio of the acquired sample to be predicted to server 104. Server 104 receives the monomer composition and target ratio of the sample to be predicted. For the monomer composition and target ratio of the sample to be predicted, server 104 randomly generates several long chains of SMILES with different sequence arrangements based on the monomer composition and target ratio of the sample to be predicted. Each long chain of SMILES with different sequence arrangements is converted into a corresponding second descriptor vector with a dimension higher than a preset value. The second descriptor vector is then subjected to dimensionality reduction processing to obtain a dimensionality-reduced second descriptor vector. The dimensionality-reduced second descriptor vector is input into a glass transition temperature prediction model to obtain predicted glass transition temperatures for different sequence samples. The glass transition temperature prediction model is a model obtained based on the glass transition temperature prediction model determination method. The predicted glass transition temperatures for different sequence samples are filtered and statistically analyzed to obtain the glass transition temperature of the sample to be predicted under the target ratio. Server 104 can then feed back the obtained glass transition temperature of the sample to be predicted under the target ratio to terminal 102. Furthermore, in some embodiments, the glass transition temperature prediction model application method can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly obtain the monomer composition and target ratio of the sample to be predicted, and obtain the glass transition temperature of the sample under the target ratio based on the monomer composition and target ratio of the sample to be predicted and the glass transition temperature prediction model. Alternatively, the server 104 can obtain the monomer composition and target ratio of the sample to be predicted from the data storage system, and obtain the glass transition temperature of the sample under the target ratio based on the monomer composition and target ratio of the sample to be predicted and the glass transition temperature prediction model.

[0037] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.

[0038] In one exemplary embodiment, such as Figure 2As shown, a method for determining a glass transition temperature prediction model is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 205.

[0039] Step 201: Obtain the original dataset; the original dataset includes: several one-component polyamide samples and corresponding glass transition temperature data, and several one-component polyimide samples and corresponding glass transition temperature data; the one-component polyamide samples and the one-component polyimide samples are both made by copolymerization of one or more monomers.

[0040] Step 202: Expand the repeating units of each single-component polyamide sample and each single-component polyimide sample into long-chain SMILES representations of a preset length.

[0041] Step 203: Convert each long chain SMILES into a corresponding first descriptor vector with a dimension higher than the preset dimension.

[0042] Step 204: Perform dimensionality reduction on the first descriptor vector to obtain the dimensionality-reduced first descriptor vector.

[0043] Step 205: Train the Gaussian process regression model based on the first descriptor vector after dimensionality reduction and the corresponding glass transition temperature data to obtain the glass transition temperature prediction model.

[0044] By implementing steps 201 to 205 above, an original dataset is constructed by integrating single-component polyamide and polyimide samples and corresponding glass transition temperature data. The polymer molecular structure is accurately characterized by long-chain SMILES. Effective features are simplified through high-dimensional descriptor transformation and dimensionality reduction optimization. Then, Gaussian process regression model training is completed by combining the corresponding glass transition temperature data. This not only fully restores the molecular configuration by using the complete long-chain structure characterization and accurately establishes the correspondence between structure and thermal properties, but also simplifies the calculation process and improves the model training efficiency by eliminating redundant features through dimensionality reduction, effectively improving the fitting accuracy and generalization prediction ability of the glass transition temperature prediction model.

[0045] As an optional implementation, the original dataset includes: 568 one-component polyamide samples and their corresponding glass transition temperature data, and 1367 one-component polyimide samples and their corresponding glass transition temperature data.

[0046] As an optional implementation, the long chain SMILES has a length of 10 repeating units.

[0047] In another exemplary embodiment of this application, in order to accurately convert each long chain SMILES into a corresponding first descriptor vector with a higher dimension than a preset dimension, the above step 203 is replaced by the following steps 2031 to 2033: Step 2031: Use Python's RDKit package to parse each long chain SMILES into a molecular Mol object. Construct the internal topological structure of the molecules' atoms and bonds through the molecular Mol object. Add hydrogen atoms or generate three-dimensional conformations to the molecular Mol object as needed to obtain the processed molecular Mol object.

[0048] Step 2032: Call the Mordred library to perform descriptor calculation on the processed molecular Mol object, and generate a numerical descriptor with a higher dimension than the preset dimension, which includes physicochemical properties, topological structure features, molecular shape and geometric features.

[0049] Step 2033: Based on the numerical descriptors with dimensions higher than the preset dimension, obtain the descriptor vector corresponding to each long chain SMILES with dimensions higher than the preset dimension.

[0050] Specifically, the SMILES string for each polymer chain is first parsed into a molecular object (Mol object) using Python's RDKit package. This object establishes the internal topological structure of the molecules' atoms and bonds, and hydrogen atoms can be added or a three-dimensional conformation can be generated for geometric optimization as needed. Then, the Mordred library is called to perform descriptor calculations on the molecular object, automatically generating high-dimensional numerical descriptors that include physicochemical properties (such as molecular weight, polarizability, hydrophobicity, etc.), topological features (such as Randic index, Wiener number), molecular shape, and geometric features. Each SMILES ultimately corresponds to a high-dimensional vector, and each dimension of the vector is the numerical value of the corresponding descriptor, which can be used for subsequent machine learning analysis or performance prediction.

[0051] As an optional implementation, the dimensionality reduction process employs principal component analysis (PCA).

[0052] As an optional implementation, the Gaussian process regression model employs the Matern kernel function; the expression for the Matern kernel function is: ; in, For the sample and samples The Matern kernel function values ​​between; This is a smoothness parameter used to adjust the smoothness of the prediction function; For Gamma function, it is a generalization of factorial function in the real number field and is used for normalization of kernel function; The variance of the signal; For length scale; Input spacing; To correct the Bessel function.

[0053] Specifically, the Gaussian process regression model is a probability-based nonparametric supervised learning method used to predict continuous variables. It assumes that the observed data are generated by a latent function and follow a multivariate Gaussian distribution.

[0054] For the training input vector: .

[0055] Corresponding output vector: 。

[0056] Satisfies a joint Gaussian distribution: .

[0057] in, This represents the first descriptor vector after dimensionality reduction corresponding to the training sample set; This represents the number of training samples; Indicates the first The first descriptor vector after dimensionality reduction corresponding to each training sample, and , This indicates the dimension of the descriptor after dimensionality reduction. y This represents the observation vector consisting of the glass transition temperature observations corresponding to each training sample; Indicates the first Observed glass transition temperature values ​​corresponding to each training sample; For kernel function The covariance matrix formed To observe the noise variance, It is an identity matrix.

[0058] For a new prediction point, the prediction output is... The conditional distribution is: .

[0059] in For the predicted mean: .

[0060] To predict uncertainty: ; ; in, This represents the first descriptor vector after dimensionality reduction of the sample to be predicted. This represents the observation vector consisting of the observed glass transition temperatures corresponding to the sample to be predicted; This represents the covariance between the first descriptor vector after dimensionality reduction of the sample to be predicted and itself, calculated using a kernel function. This represents the first descriptor vector after dimensionality reduction for the sample to be predicted and the first descriptor vector after dimensionality reduction for the training sample set. The covariance row vectors between them are calculated using a kernel function; This represents the first descriptor vector after dimensionality reduction corresponding to the training sample set. The covariance column vector between the first descriptor vectors after dimensionality reduction and corresponding to the sample to be predicted is obtained by calculating the kernel function. Indicates the first The covariance between the first descriptor vector after dimensionality reduction of each training sample and the first descriptor vector after dimensionality reduction of the sample to be predicted is calculated using a kernel function.

[0061] In this embodiment, the kernel function used is the Matern kernel function, and its expression is: .

[0062] when When the function is continuous but not differentiable, it has finite smoothness; when It degenerates into a squared exponential kernel (RBF kernel) and has infinite smoothness; the median value The controllable function balances local variations and global trends, thus capturing the nonlinear relationship between descriptors and performance in polymer performance prediction while ensuring the smoothness and generalization ability of the prediction results.

[0063] In one exemplary embodiment, such as Figure 3 As shown, a method for applying a glass transition temperature prediction model is provided, including the following steps 301 to 306: Step 301: Obtain the monomer composition and target ratio of the sample to be predicted.

[0064] Step 302: Based on the monomer composition and target ratio of the sample to be predicted, randomly generate several long chain SMILES representations with different sequence arrangements.

[0065] Step 303: Convert the long chains SMILES with different sequence arrangements into corresponding second descriptor vectors with a higher dimension than the preset dimension.

[0066] Step 304: Perform dimensionality reduction processing on the second descriptor vector to obtain the dimensionality-reduced second descriptor.

[0067] Step 305: Input the second descriptor into the glass transition temperature prediction model to obtain the predicted glass transition temperature values ​​corresponding to different sequence samples; the glass transition temperature prediction model is a model obtained based on the glass transition temperature prediction model determination method described above.

[0068] Step 306: Filter and statistically analyze the predicted glass transition temperature values ​​corresponding to the different sequence samples to obtain the glass transition temperature of the sample to be predicted under the target ratio.

[0069] As an optional implementation, 50 long chain SMILES representations with different sequence arrangements are randomly generated based on the monomer composition and target ratio of the sample to be predicted.

[0070] In another exemplary embodiment of this application, step 306 is replaced by steps 3061 to 3063: Step 3061: Based on the monomer composition of the copolyamide to be predicted, obtain the glass transition temperature of the homopolymer corresponding to each monomer in the copolyamide to be predicted, and select the maximum value among the glass transition temperatures of the homopolymers corresponding to each monomer in the copolyamide to be predicted as the screening basis.

[0071] Step 3062: Sort the predicted glass transition temperature values ​​corresponding to different sequence samples according to their numerical values, and remove the portion of the predicted glass transition temperature values ​​corresponding to different sequence samples that are higher than the maximum value to obtain the filtered predicted glass transition temperature values.

[0072] Step 3063: Average the glass transition temperature predictions of the previous screenings to obtain the glass transition temperature of the sample to be predicted under the target ratio.

[0073] Specifically, the predicted glass transition temperatures of different sequence samples are sorted according to their numerical values. Predicted values ​​that are higher than the glass transition temperatures of the corresponding homopolymers of each monomer are removed. The average of the top 20 predicted values ​​is then used as the glass transition temperature of the sample to be predicted under the target ratio.

[0074] In one exemplary embodiment, such as Figure 4 As shown, a device for determining a glass transition temperature prediction model is provided, comprising the following modules: The original dataset acquisition module is used to acquire the original dataset, which includes: several one-component polyamide samples and corresponding glass transition temperature data, and several one-component polyimide samples and corresponding glass transition temperature data; the one-component polyamide samples and the one-component polyimide samples are both made by copolymerization of one or more monomers.

[0075] An extension module is used to extend the repeating units of each single-component polyamide sample and each single-component polyimide sample into long-chain SMILES representations of a preset length.

[0076] The first transformation module is used to transform each long chain SMILES into a corresponding first descriptor vector with a higher dimension than the preset dimension.

[0077] The first dimensionality reduction processing module is used to perform dimensionality reduction processing on the first descriptor vector to obtain the dimensionality-reduced first descriptor vector.

[0078] The training module is used to train the Gaussian process regression model based on the first descriptor vector after dimensionality reduction and the corresponding glass transition temperature data, so as to obtain the glass transition temperature prediction model.

[0079] In one exemplary embodiment, such as Figure 5 As shown, a glass transition temperature prediction model application device is provided, including the following modules: The test data acquisition module is used to acquire the monomer composition and target ratio of the sample to be predicted.

[0080] The sequence generation module is used to randomly generate several long chain SMILES representations with different sequence arrangements based on the monomer composition and target ratio of the sample to be predicted.

[0081] The second conversion module is used to convert the long chains of SMILES with different sequence arrangements into corresponding second descriptor vectors with a higher dimension than a preset dimension.

[0082] The second dimensionality reduction processing module is used to perform dimensionality reduction processing on the second descriptor vector to obtain the dimensionality-reduced second descriptor vector.

[0083] The prediction module is used to input the second descriptor vector after dimensionality reduction into the glass transition temperature prediction model to obtain the glass transition temperature prediction values ​​corresponding to different sequence samples; the glass transition temperature prediction model is a model obtained based on the glass transition temperature prediction model determination method described above.

[0084] The screening and statistics module is used to screen and statistically analyze the predicted glass transition temperature values ​​corresponding to the different sequence samples, so as to obtain the glass transition temperature of the sample to be predicted under the target ratio.

[0085] In one example, this application uses long-chain SMILES to represent the sequence information of copolyamides and utilizes a combination of Mordred descriptor extraction, principal component analysis dimensionality reduction, and machine learning modeling to predict the glass transition temperature of the target copolyamide.

[0086] Acquire several single-component polyamide samples and their corresponding glass transition temperature data, and extend the repeating units of each single-component polyamide sample into long-chain SMILES of a preset length.

[0087] The Mordred descriptor of the long chain SMILES is calculated to generate a descriptor vector with a dimension higher than a preset value. Principal component analysis is then performed on the descriptor vector to reduce its dimension, resulting in a dimension-reduced descriptor vector.

[0088] The Gaussian process regression model is trained based on the dimensionality-reduced descriptor vector and the corresponding glass transition temperature data to obtain a glass transition temperature prediction model.

[0089] For the copolyamide to be predicted, multiple long-chain SMILES representing different sequence arrangements are randomly generated according to the monomer composition and ratio to characterize the sequence changes of the copolymerization system. At the same time, the Mordred descriptors of the generated long-chain SMILES are calculated to generate descriptor vectors with a higher dimension than the preset dimension. Principal component analysis is then performed to reduce the dimension of the descriptor vector of the copolyamide to be predicted.

[0090] The descriptor vector of the copolyamide to be predicted is input into the glass transition temperature prediction model to obtain the predicted glass transition temperature values ​​corresponding to different sequence samples.

[0091] The predicted values ​​are sorted, and abnormal predictions that are higher than the glass transition temperature of the corresponding single-component polyamide are removed. The high-confidence predictions after screening are statistically analyzed to obtain the glass transition temperature of the copolyamide under the target ratio.

[0092] This application can characterize the effect of sequence differences of different repeating units in copolyamide on glass transition temperature, and has high prediction accuracy, sequence sensitivity and generalization ability. It can be used for copolyamide formulation screening and performance design.

[0093] In another example, this application uses the prediction of the glass transition temperature of a copolyimide as an illustration.

[0094] Acquire several single-component polyimide samples and their corresponding glass transition temperature data, and expand the repeating units of each single-component polyimide sample into long-chain SMILES of a preset length.

[0095] The Mordred descriptor of the long chain SMILES is calculated to generate a descriptor vector with a dimension higher than a preset value. Principal component analysis is then performed on the descriptor vector to reduce its dimension, resulting in a dimension-reduced descriptor vector.

[0096] The Gaussian process regression model is trained based on the dimensionality-reduced descriptor vector and the corresponding glass transition temperature data to obtain a glass transition temperature prediction model.

[0097] For the copolyimide to be predicted, multiple long-chain SMILES representing different sequence arrangements are randomly generated based on the monomer composition and ratio to characterize the sequence variation of the copolymerization system. At the same time, the Mordred descriptors of the generated long-chain SMILES are calculated to generate descriptor vectors with a higher dimension than the preset dimension. Principal component analysis is then performed to reduce the dimension of the descriptor vector of the copolyimide to be predicted.

[0098] The descriptor vector of the copolyimide to be predicted is input into the glass transition temperature prediction model to obtain the predicted glass transition temperature values ​​corresponding to different sequence samples.

[0099] The predicted values ​​are sorted, and abnormal predictions that are higher than the glass transition temperature of the corresponding single-component polyimide are removed. The high-confidence predictions after screening are statistically analyzed to obtain the glass transition temperature of the copolyimide under the target ratio.

[0100] In some embodiments, to accurately predict the glass transition temperature of the copolyamide, the method further includes: acquiring a training dataset, which includes several one-component polyamide samples and their corresponding glass transition temperature data, preferably including several one-component polyimide samples with similar structures and their corresponding glass transition temperature data; expanding the repeating units of each one-component polyamide sample and each one-component polyimide sample into long chain SMILES representations of a preset length; converting each long chain SMILES into a corresponding descriptor vector with a dimension higher than a preset dimension, and performing dimensionality reduction using principal component analysis; training a Gaussian process regression model using the dimensionality-reduced descriptor vector as input and the corresponding glass transition temperature as a label to obtain a glass transition temperature prediction model.

[0101] For the copolyamide to be predicted, multiple long chains with different sequence arrangements are randomly generated under a given composition ratio. The Mordred descriptor corresponding to the long chain SMILES is calculated to generate a descriptor vector with a dimension higher than the preset dimension, and the dimension is reduced by principal component analysis. The dimension-reduced descriptor vector is input into the glass transition temperature prediction model to obtain the glass transition temperature of the copolyamide to be predicted under different random sequences. The prediction results are then filtered and statistically analyzed to obtain the final predicted value under the target ratio.

[0102] Specifically, glass transition temperature data for single-component polyamides are collected, and glass transition temperature data for polyimides with similar structures are preferably supplemented, such as... Figure 6 It is known that this includes 568 glass transition temperature data points for polyamides and 1367 glass transition temperature data points for polyimides. For example... Figure 7 It is known that the repeating units of a single-component polyamide are extended into long chains of 10 repeating units SMILES by connecting them through amide bonds, which are then used for subsequent feature extraction and model training.

[0103] For a two-repeating-unit copolymer polyamide system formed from three monomers, random copolymerization under given ratios generates multiple long-chain smiles. Each long-chain smile corresponds to a possible sequence arrangement, thus characterizing differences in sequence under the same composition ratio. For example, ... Figure 8 As shown, using diamine A, diacid B, and diacid C as monomers, and feeding them in a molar ratio of 10:6:4, a long chain of copolymer polyamide composed of two repeating units, diamine A-diacid B and diamine A-diacid C, is formed through random copolymerization. Under the same mixing ratio, multiple long chain structures with different sequence arrangements are randomly generated and converted into SMILES.

[0104] Mordred descriptors for long-chain SMILES in the training dataset are calculated and their features extracted to obtain descriptor vectors with dimensions higher than a preset limit. Principal component analysis is then used to reduce the dimensionality of these descriptor vectors. A Gaussian process regression model is established based on the dimensionality-reduced descriptor vectors to train the glass transition temperature prediction relationship between single-component polyamides and polyimides. The polyamide data is split into training and testing sets in a 9:1 ratio, and all polyimide data is added to the training set. Figure 9 It is known that the training set contains 1878 data points, and the test set contains 57 data points. For the polyamide glass transition temperature prediction model, the R-value of the model on the test set is... 2 The MAE values ​​are 0.98 and 10℃, respectively, and the R value of the training set is... 2 The MAE values ​​are 0.99 and 5℃.

[0105] For the copolyamide to be predicted, the Mordred descriptor vector of the randomly generated sequence is calculated and subjected to principal component analysis dimensionality reduction processing consistent with the training phase. The resulting vector is input into a Gaussian process regression model to obtain multiple predicted glass transition temperatures. The predicted values ​​are sorted, and outliers with glass transition temperatures higher than the corresponding single-component polyamide are removed. The top 20 predicted values ​​are then averaged to obtain the final predicted glass transition temperature.

[0106] Figure 10The results of 50 random copolymerization simulations of two different polyamide structures (blue squares represent segments containing terephthalic acid and decanediamine units, and green squares represent segments containing adipic acid and pentanediamine units) in a 5:5 ratio are presented. The ranking results are arranged from high temperature to low temperature, corresponding to the degree of order of the copolymer segments from high to low, with glass transition temperatures of 227℃, 117℃, 90℃, and 80℃, respectively.

[0107] The above method can be used to predict the glass transition temperature of copolyamide systems with different ratios. For sequence-sensitive systems, this application can characterize the glass transition temperature fluctuations caused by sequence arrangement and provide a basis for subsequent formulation design and experimental screening.

[0108] This application utilizes a repeating unit PA10T ( Figure 11 ) and PA56 ( Figure 12 The composition of the copolymer system validated the reliability of the glass transition temperature prediction model trained in this application. Specific results are as follows: Figure 13 As shown.

[0109] Figure 13 The experimental and model predicted values ​​for glass transition temperatures (GLTs) under different ratios of PA10T:PA56 are shown as follows: 123℃ and 117℃ for PA10T:PA56 at a ratio of 10:0; 103.5℃ and 109℃ for PA10T:PA56 at a ratio of 9:1; 94.7℃ and 95℃ for PA10T:PA56 at a ratio of 8:2; 84.1℃ and 99℃ for PA10T:PA56 at a ratio of 7:3; 97.1℃ and 88℃ for PA10T:PA56 at a ratio of 7:3; 90.8℃ and 84℃ for PA10T:PA56 at a ratio of 6:4; 82.2℃ and 73℃ for PA10T:PA56 at a ratio of 5:5; 72.9℃ and 72℃ for PA10T:PA56 at a ratio of 4:6; 63.4℃ and 55℃ for PA10T:PA56 at a ratio of 3:7; 60.5℃ and 58℃ for PA10T:PA56 at a ratio of 1:9; and 73℃ and 67℃ for PA10T:PA56 at a ratio of 0:10. It is evident that the predicted glass transition temperature values ​​obtained by the method of this application under different ratios are generally close to the experimental values, and can well reflect the overall trend of the glass transition temperature of the PA10T / PA56 copolymer system with the change of composition, indicating that the prediction method proposed in this application has good reliability and practicality.

[0110] This application also uses the repeating unit BFA_TERE ( Figure 14 ) and BFA_ISO ( Figure 15 The composition of the copolymer system validated the reliability of the glass transition temperature prediction model trained in this application. Specific results are as follows: Figure 16 As shown.

[0111] Figure 16The experimental and model-predicted glass transition temperatures (GTH) for different ratios are shown as follows: 323.2℃ and 325.5℃ for BFA_TERE:BFA_ISO of 10:0; 307.2℃ and 314.9℃ for 8:2; 307.2℃ and 305.8℃ for 6:4; 302.8℃ and 305.0℃ for 5:5; 298.7℃ and 298.0℃ for 4:6; 290.4℃ and 293.0℃ for 2:8; and 285.5℃ and 292.8℃ for 0:10. For ratios of 9:1, 7:3, 3:7, and 1:9, only model-predicted values ​​are currently available: 324.0℃, 314.7℃, 297.5℃, and 293.8℃, respectively. It is evident that, within the range of proportions with available experimental data, the predicted glass transition temperature obtained by the method of this application is generally close to the experimental value, and can well reflect the overall trend of glass transition temperature of the BFA_TERE / BFA_ISO copolymer system as a function of composition. This indicates that the prediction method proposed in this application has good reliability and application value.

[0112] Compared with the prior art, this application has at least the following advantages: (1) By generating multiple long-chain SMILES under the same composition ratio, the sequence changes of copolyamide and copolyimide can be explicitly characterized.

[0113] (2) By combining the Mordred descriptor with PCA dimensionality reduction, the feature dimension can be reduced while preserving the main structural information, thereby improving modeling efficiency.

[0114] (3) A high-precision prediction of the glass transition temperature of copolyamide and copolyimide is achieved by using a Gaussian process regression model based on dimensionality reduction features.

[0115] (4) By screening and statistically analyzing the prediction results of multiple sequence samples, the interference of abnormal sequences on the final prediction results can be reduced, and the robustness of the prediction can be improved.

[0116] (5) This application can be used for the formulation design, performance screening and material development efficiency improvement of copolyamide and copolyimide.

[0117] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 17As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data from the training and application processes of the glass transition temperature prediction model. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for determining and applying a glass transition temperature prediction model.

[0118] Figure 17 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0119] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0120] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0121] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with relevant regulations and be authorized by the owner of the corresponding device.

[0122] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0123] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0124] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0125] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A glass transition temperature prediction model determination method, characterized by, The method for determining the glass transition temperature prediction model includes: Obtain the original dataset; the original dataset includes: several one-component polyamide samples and corresponding glass transition temperature data, and several one-component polyimide samples and corresponding glass transition temperature data; the one-component polyamide samples and the one-component polyimide samples are both made by copolymerization of one or more monomers; The repeating units of each single-component polyamide sample and each single-component polyimide sample are expanded into long-chain SMILES of a preset length. Transform each long chain SMILES into a corresponding first descriptor vector with a dimension higher than the preset value; The first descriptor vector is reduced in dimension to obtain the first descriptor vector after dimension reduction. The Gaussian process regression model is trained based on the first descriptor vector after dimensionality reduction and the corresponding glass transition temperature data to obtain a glass transition temperature prediction model.

2. The glass transition temperature prediction model determination method according to claim 1, characterized in that, Each long chain SMILES is transformed into a corresponding first descriptor vector with a dimension higher than the preset limit, specifically including: The RDKit package in Python is used to parse each long chain SMILES into a molecular Mol object. The internal topological structure of the atoms and bonds of the molecule is constructed through the molecular Mol object. Hydrogen atoms are added to the molecular Mol object or a three-dimensional conformation is generated as needed to obtain the processed molecular Mol object. The Mordred library is called to perform descriptor calculation on the processed Mol object, generating a numerical descriptor with a higher dimension than the preset dimension, which includes physicochemical properties, topological structure features, molecular shape and geometric features. Based on the numerical descriptors with dimensions higher than the preset limit, a first descriptor vector with dimensions higher than the preset limit is obtained for each long chain SMILES.

3. The glass transition temperature prediction model determination method according to claim 1, wherein, The Gaussian process regression model uses the Matern kernel function; the expression for the Matern kernel function is: ; in, For the sample and samples The Matern kernel function values ​​between; This is a smoothness parameter used to adjust the smoothness of the prediction function; For Gamma function, it is a generalization of factorial function in the real number field and is used for normalization of kernel function; The variance of the signal; For length scale; Input spacing; To correct the Bessel function.

4. A glass transition temperature prediction model application method characterized by, The application methods of the glass transition temperature prediction model include: Obtain the monomer composition and target ratio of the sample to be predicted; Based on the monomer composition and target ratio of the sample to be predicted, several long chains SMILES with different sequence arrangements are randomly generated; The long chains SMILES with different sequence arrangements are transformed into corresponding second descriptor vectors with a higher dimension than a preset number. The second descriptor vector is reduced in dimension to obtain the second descriptor vector after dimension reduction. The second descriptor vector after dimensionality reduction is input into the glass transition temperature prediction model to obtain the glass transition temperature prediction values ​​corresponding to different sequence samples; the glass transition temperature prediction model is a model obtained based on the glass transition temperature prediction model determination method according to any one of claims 1-3. The predicted glass transition temperatures of the different sequence samples are screened and statistically analyzed to obtain the glass transition temperature of the sample to be predicted under the target ratio.

5. The glass transition temperature prediction model application method according to claim 4, characterized in that, The predicted glass transition temperatures of the different sequence samples are screened and statistically analyzed to obtain the glass transition temperature of the sample to be predicted under the target ratio, specifically including: Based on the monomer composition of the sample to be predicted, the glass transition temperature of the homopolymer corresponding to each monomer in the sample to be predicted is obtained, and the maximum value among the glass transition temperatures of the homopolymer corresponding to each monomer in the sample to be predicted is selected as the screening criterion. The predicted glass transition temperature values ​​corresponding to different sequence samples are sorted according to their numerical values. The portion of the predicted glass transition temperature values ​​corresponding to different sequence samples that is higher than the maximum value is removed to obtain the filtered predicted glass transition temperature values. The glass transition temperature predictions of the first few screened samples are averaged to obtain the glass transition temperature of the sample to be predicted under the target ratio.

6. A device for determining a glass transition temperature prediction model, characterized in that, The device for determining the glass transition temperature prediction model includes: The original dataset acquisition module is used to acquire the original dataset, which includes: several one-component polyamide samples and their corresponding glass transition temperature data, and several one-component polyimide samples and their corresponding glass transition temperature data; the one-component polyamide samples and the one-component polyimide samples are both made by copolymerization of one or more monomers. An extension module is used to extend the repeating units of each single-component polyamide sample and each single-component polyimide sample into long-chain SMILES representations of a preset length; The first transformation module is used to transform each long chain SMILES into a corresponding first descriptor vector with a higher dimension than the preset dimension; The first dimension reduction processing module is used to perform dimension reduction processing on the first descriptor vector to obtain the dimension-reduced first descriptor vector. The training module is used to train the Gaussian process regression model based on the first descriptor vector after dimensionality reduction and the corresponding glass transition temperature data, so as to obtain the glass transition temperature prediction model.

7. A glass transition temperature prediction model application apparatus characterized by comprising: a glass transition temperature prediction model application program; and a glass transition temperature prediction model application program execution unit. The device for applying the glass transition temperature prediction model includes: The test data acquisition module is used to acquire the monomer composition and target ratio of the sample to be predicted; The sequence generation module is used to randomly generate several long chains of SMILES with different sequence arrangements based on the monomer composition and target ratio of the sample to be predicted. The second conversion module is used to convert the long chains of SMILES with different sequence arrangements into corresponding second descriptor vectors with a higher than preset dimension. The second dimension reduction processing module is used to perform dimension reduction processing on the second descriptor vector to obtain the dimension-reduced second descriptor vector. The prediction module is used to input the second descriptor vector after dimensionality reduction into the glass transition temperature prediction model to obtain the glass transition temperature prediction values ​​corresponding to different sequence samples; the glass transition temperature prediction model is a model obtained based on the glass transition temperature prediction model determination method according to any one of claims 1-3. The screening and statistics module is used to screen and statistically analyze the predicted glass transition temperature values ​​corresponding to the different sequence samples, so as to obtain the glass transition temperature of the sample to be predicted under the target ratio.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that the processor executes the computer program to implement the method for determining the glass transition temperature prediction model according to any one of claims 1-3 or the method for applying the glass transition temperature prediction model according to any one of claims 4-5.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, When executed by a processor, the computer program implements the method for determining the glass transition temperature prediction model according to any one of claims 1-3 or the method for applying the glass transition temperature prediction model according to any one of claims 4-5.

10. A computer program product comprising a computer program, characterized in that, When executed by a processor, the computer program implements the method for determining the glass transition temperature prediction model according to any one of claims 1-3 or the method for applying the glass transition temperature prediction model according to any one of claims 4-5.