Training method, prediction method and system of model for predicting polyamide performance parameters, structural design method, equipment and storage medium

Through the machine learning model of GCN and KAN, the efficiency and cost problems of polyamide materials in high-performance design are solved, and high-precision prediction of glass transition temperature, melting point and tensile modulus are achieved, which promotes the application in the aerospace field.

CN120260741APending Publication Date: 2025-07-04EAST CHINA UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510114192.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When designing polyamide materials, it is difficult to simultaneously improve the performance of high glass transition temperature, high melting point and high tensile modulus, resulting in low efficiency and high cost of material research and development, especially in applications in the aerospace field.

Method used

The machine learning model of graph convolutional neural network (GCN) combined with Kolmogorov-Arnold Networks (KAN) is used to predict the performance parameters of polyamide, including glass transition temperature, melting point and tensile modulus by training the relationship between the polyamide structural formula and performance parameters in the dataset.

Benefits of technology

It improves the R&D efficiency of polyamide materials, reduces R&D costs, realizes high-precision prediction of polyamide performance parameters, enhances the generalization ability of the model, and can quickly screen out new polyamide structures that meet high performance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260741A_ABST
    Figure CN120260741A_ABST
Patent Text Reader

Abstract

The invention discloses a training method, a prediction method and a system of a model for predicting polyamide performance parameters, a structural design method, equipment and a storage medium. The model training method comprises the following steps that S1, data sets are obtained, the data sets comprise a first data set, and the first data set comprises different structural formulas of polyamide and performance parameters corresponding to the structural formulas; the performance parameters are one or more of glass transition temperature, melting point and tensile modulus; and S2, inputting the data set into a graph convolutional neural network, and training by taking a KAN as a middle layer to obtain a prediction model. The method provided by the invention has higher precision and generalization ability, and has higher interpretability compared with a traditional neural network; the method has relatively high prediction accuracy on performance parameters of polyamide; the research and development efficiency of the polyamide material is greatly improved, the research and development cost is reduced, and the reasonable design of an advanced polyamide structure can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a training method, a prediction method and system, a structural design method, a device, and a storage medium for a model for predicting performance parameters of polyamide. Background Art

[0002] Polyamide is a thermoplastic resin polymerized from monomers containing carboxyl and amino groups through amide bonds. Due to its unique molecular structure, it has excellent heat resistance and mechanical properties, and has broad application prospects in the fields of aerospace and national defense industries. In the actual applications in these fields, materials often need to have multiple excellent properties to meet the application requirements. For example, the aircraft protective cover needs to face extreme temperature changes, huge mechanical stresses, and complex aerodynamic environments. Therefore, the selection of its materials must meet a series of strict performance requirements, including high glass transition temperature, high melting point, and high tensile modulus.

[0003] Since there are often contradictory and restrictive relationships among multiple properties, it is difficult for materials to simultaneously have a high glass transition temperature, a high melting point, and a high tensile modulus. To obtain high-performance materials that meet actual needs, researchers need to make trade-offs among various properties. The traditional material research and development mode of "experience + experimental trial and error" for traditional materials not only has low efficiency and high design costs, but also is difficult to solve the problem of simultaneously improving multiple properties of polyamide.

[0004] Therefore, in the case where the mechanical property data of polyamide is scarce, especially the tensile modulus is difficult to determine, how to establish a prediction method for polyamide performance parameters, how to quickly and accurately predict polyamide performance parameters, and how to guide researchers to reasonably design the chemical structure of polyamide are technical problems that need to be solved urgently at present. Summary of the Invention

[0005] In order to overcome the technical problems such as low efficiency and high cost in the process of designing polyamide with multiple contradictory properties in the existing traditional material research and development mode of "experience + experimental trial and error", the present invention provides a training method, a prediction method and system, a structural design, a device, and a storage medium for a model for predicting performance parameters of polyamide. The method of the present application has higher accuracy and generalization ability, and has higher interpretability compared with traditional neural networks; this method has high prediction accuracy for the performance parameters of polyamide; it greatly improves the research and development efficiency of polyamide materials, reduces the research and development cost, and can realize the reasonable design of advanced polyamide structures; the designed polyamide meets the requirements of high glass transition temperature, high melting point, and high tensile modulus, and has broad application prospects in high-tech fields such as aerospace.

[0006] The present invention solves the above technical problems through the following technical solutions:

[0007] The present invention provides a method for training a model for predicting performance parameters of polyamides, which comprises the following steps:

[0008] S1. Obtain a data set, where the data set includes a first data set, and the first data set includes different structural formulas of polyamides and the corresponding performance parameters of each structural formula; the performance parameters are one or more of glass transition temperature, melting point, and tensile modulus;

[0009] S2. Input the data set into a graph convolutional neural network, and use KAN as the intermediate layer to train the model.

[0010] In the present invention, the graph convolutional neural network is generally referred to as GCN, and the full name of KAN is Kolmogorov - Arnold Networks; constructing a machine learning model combining GCN and KAN can be simply referred to as the GC - KAN model, which can be used to automatically learn the relationship between polymer structural formulas and performance parameters.

[0011] In the present invention, those skilled in the art know that the structural formula of the polyamide is generally represented by SMILES, and the full name is Simplified molecular input line entry specification, which is a specification for clearly describing the molecular structural formula with an ASCII string.

[0012] In the present invention, those skilled in the art know that the performance parameters of the polymers in the data set can generally be obtained by collecting in public literature and / or obtaining from public databases.

[0013] In the present invention, the polyamide refers to a repeating unit containing an amide group.

[0014] In some embodiments, when the performance parameter is the glass transition temperature, the data volume of the first data set is 100 - 1000, preferably 501.

[0015] In some embodiments, when the performance parameter is the melting point, the data volume of the first data set is 100 - 1000, preferably 900.

[0016] In some embodiments, when the performance parameter is the tensile modulus, the data volume of the first data set is 100 - 1000, preferably 131.

[0017] In some embodiments, the first data set includes a training set and a test set, and the ratio of the training set to the test set is 9:1 or 8:2.

[0018] In a specific embodiment, when the performance parameter is the glass transition temperature, the data volume ratio of the training set to the test set is 9:1.

[0019] In a specific embodiment, when the performance parameter is the melting point, the data volume ratio of the training set to the test set is 9:1.

[0020] In a specific embodiment, when the performance parameter is the tensile modulus, the data volume ratio of the training set to the test set is 8:2.

[0021] In some embodiments, the data set further includes a second data set for supplementing the first data set; the second data set includes different structural formulas of polyimide and the corresponding performance parameters of each structural formula, and the performance parameter is one or more of the glass transition temperature, melting point, and tensile modulus; preferably, the data volume of the second data set is 386.

[0022] In the present invention, the experimental data of the tensile modulus of polyamide are scarce. Directly using this part of the data to establish a model results in a low-precision model with weak generalization ability; introducing the experimental data of polyimide as a supplement improves the accuracy of the machine learning model in predicting the tensile modulus of polyamide and solves the problem of small data of materials.

[0023] Among them, preferably, supplementing the structural formula and performance parameter data of polyimide to the training set of the first data set can automatically learn the potential relationship between the structural formula and performance parameters of polyamide similar to polyimide; the supplementary data introduced through the second data set can expand the chemical space learned by the model and enhance the generalization ability of the model.

[0024] In some embodiments, in step S1, after obtaining the data set, a format normalization processing step is further included; preferably, the format normalization processing step includes using the repeating unit as the basic format and adding identifiers at both ends of the molecule of the repeating unit, such as adding "*" as the identifier.

[0025] In some embodiments, in step S1, after obtaining the data set, a cleaning step is further included; preferably, the cleaning step includes removing different performance parameters of polyamide with the same structural formula and only retaining the highest value of the performance parameter.

[0026] In some embodiments, in step S2, the graph convolutional neural network includes 3 convolutional layers; preferably, the graph convolutional neural network further includes 3 dropout layers, and each dropout layer is respectively arranged after each convolutional layer; more preferably, the dropout probability of each dropout layer is 0-0.2.

[0027] Among them, each of the dropout layers is used to randomly discard the outputs of some neurons to prevent overfitting.

[0028] In the present invention, each convolutional layer of the graph convolutional neural network uses a linear layer as a convolutional kernel to map edge features into a weight matrix matching the node feature dimension, and weights the features of adjacent nodes.

[0029] In a specific embodiment, each of the convolutional layers is a dynamic edge-conditioned filter, and the dynamic edge-conditioned filter includes an edge feature dimension, a node feature dimension, an input dimension, and an output dimension of the convolutional layer; preferably, the input dimensions of the convolutional layer are 40-1000 respectively; preferably, the output dimensions of the convolutional layer are 40-1000 respectively; preferably, the edge feature dimension is 15, and the node feature dimension is 42-1000.

[0030] In one embodiment, the input dimension of the first convolutional layer is 42, the output dimension is 1000, the node feature dimension is 42, and the edge feature dimension is 15; the input dimension of the second convolutional layer is 1000, the output dimension is 250; the input dimension of the third convolutional layer is 250, the output dimension is 40, the node feature dimension is 250, and the edge feature dimension is 15.

[0031] Among them, the "dynamic edge-conditioned filter" is also called edge-conditioned convolution, which refers to the linear layer as the convolutional kernel and is used to map edge features into a weight matrix matching the node feature dimension, and its parameters are jointly determined by the edge feature dimension, the node feature dimension, and the output dimension of the convolutional layer.

[0032] In some embodiments, in step S2, the graph convolutional neural network is used to obtain the atomic information, bond information, and adjacent atomic positions of the polymer; preferably, the atomic information is one or more of atomic number, aromaticity, total number of connected chemical bonds, number of hydrogen atoms, and hybridization type to represent the atomic information; preferably, the bond information is bond type, stereochemical property, direction, and aromaticity; preferably, the adjacent atomic positions are the indices of the two atoms connected by the bond.

[0033] Among them, the atomic information and the bond information include but are not limited to the above types.

[0034] In some embodiments, in step S2, the KAN includes 3 hidden layers; preferably, the input dimensions of each of the hidden layers are 5-50; preferably, the output dimensions of each of the hidden layers are 5-50; preferably, each of the hidden layers includes a dynamic activation function; more preferably, the grid value of the dynamic activation function is 5; preferably, the k value of the dynamic activation function is 3.

[0035] Among them, "grid" represents a grid, which defines the node positions of the activation function and determines the partitioning of the input space; "k" represents the order of the dynamic activation function. The higher the order, the stronger the expression ability of the activation function. By setting the two parameters of "grid" and "k", the shape of the dynamic activation function is controlled.

[0036] In one embodiment, the input dimension of the first hidden layer is 40 and the output dimension is 20; the input dimension of the second hidden layer is 20 and the output dimension is 5; the input dimension of the third hidden layer is 5 and the output dimension is 1.

[0037] In step S2 of the present invention, the specific implementation manner of the graph convolutional neural network may be: edge features generate a weight matrix through a multi-layer perceptron. After the adjacent node features are transformed by the weight matrix of the connected edges, they are added to the features of the central node to form updated node features. This process enables the edge features to dynamically adjust the propagation weight of information between nodes; finally, all the updated node features are added together to obtain the feature vector of the molecular graph.

[0038] In step S2 of the present invention, the specific implementation manner of the KAN may be: mapping the feature vector of the molecular graph to a high-dimensional space to capture the complex structural information of the graph. In each layer, neurons receive the input from the previous layer and perform a non-linear transformation through an activation function. The activation function is dynamically adjusted during the training process to approximate the complex functional relationship between the input data and the high-dimensional space, so as to better fit the molecular graph features. Through layer-by-layer non-linear transformation and the representation in the high-dimensional space, KAN can effectively capture the local and global structural features in the graph and enhance the expression ability of the model in tasks such as molecular property prediction.

[0039] The present invention also provides a method for predicting polyamide performance parameters, which includes the following steps: inputting the structural formula of the polyamide to be measured into the prediction model to obtain the performance parameters of the polyamide to be measured; the prediction model is obtained by using the training method of the above-mentioned model; the types of the performance parameters of the polyamide to be measured are the same as those in the first dataset.

[0040] In the present invention, a machine learning model constructed by combining a graph convolutional neural network with KAN dynamically adjusts the information transmission between nodes through the weight matrix generated by edge features, enabling the model to flexibly adjust the influence of adjacent nodes on the central node according to different molecular structures and chemical environments, improving the modeling ability for complex relationships; the dynamic activation function of KAN in polyamide performance prediction enhances the modeling ability of the model for complex molecular graph features, enabling the model to flexibly adapt to the characteristics of different data, and ultimately improving the prediction accuracy and generalization ability.

[0041] In some embodiments, the polyamide to be measured includes a polyamide with a known chemical structure or a virtual polyamide.

[0042] Among them, the polyamide with a known chemical structure may be a polyamide with both a known chemical structure and performance parameters, or a polyamide with a known chemical structure but unknown performance parameters.

[0043] In a specific embodiment, the virtual polyamide is obtained by combining diamine groups and diacid groups through a chemical reaction template; preferably, the number of both the diamine groups and the diacid groups is 200 - 400. For example, the number of diamine groups is 361, and the number of diacid groups is 309.

[0044] Among them, preferably, the combination of the virtual polyamides includes the following steps:

[0045] (1) Define the monomer types in the polyamide chemical structure; among them, the monomer types preferably include: ① diamine monomers with amino groups at both ends of the main chain; ② diacid monomers with carboxyl groups at both ends of the main chain.

[0046] (2) Search for diamine monomers and diacid monomers in a public database for combination to obtain the chemical structures of several virtual polyamides, forming a gene pool.

[0047] Among them, the combination method is preferably implemented through Python programming.

[0048] In a specific embodiment, the number of the virtual polyamides is more than 100,000, preferably 100,000 - 120,000, for example, 111,549.

[0049] In the present invention, those skilled in the art know that the chemical structure information of the polyamide to be measured is generally the input obtained by converting the chemical structure of the repeating unit of the polyamide to be measured into the corresponding SMILES.

[0050] The present invention provides a prediction system for polyamide performance parameters, which includes a prediction module. The prediction module is used to input the chemical structure of the polyamide to be measured into a prediction model to obtain the performance parameters of the polyamide to be measured; the prediction model is obtained by using the training method of the model as described above.

[0051] The present invention provides a method for polyamide structure design, which includes the following steps:

[0052] S1. Use a prediction model to predict the candidate chemical structures of virtual polyamides to obtain the performance parameters of the candidate chemical structures; the prediction model is obtained by using the training method of any one of the above models; the performance parameters of the candidate chemical structures are glass transition temperature, melting point, and tensile modulus.

[0053] s2. Calculate the comprehensive scores of the candidate structural formulas of each of the virtual polyamides, and sort the comprehensive scores to obtain the structural formulas of the polyamides.

[0054] In some embodiments, the candidate structural formulas are obtained by combining diamine monomers and diacid monomers through a chemical reaction template.

[0055] In a specific embodiment, the number of both the diamine monomers and the diacid monomers is 200 - 400. For example, the number of the diamine monomers is 361, and the number of the diacid monomers is 309.

[0056] In some embodiments, the number of the candidate structural formulas is more than 100,000, more preferably 100,000 - 120,000, such as 111,549.

[0057] In some embodiments, the method for calculating the comprehensive score includes calculating using an evaluation function in a weighted form.

[0058] Among them, the calculation method of the comprehensive score is as follows:

[0059]

[0060] Among them, Score is the comprehensive score, T g , T m , E respectively represent the glass transition temperature, melting point, and tensile modulus; T g min and T g max are respectively the minimum value and the maximum value of the glass transition temperature among all virtual polyamides, T m min and T m max are respectively the minimum value and the maximum value of the melting point among all virtual polyamides, E min and E max are respectively the minimum value and the maximum value of the tensile modulus among all virtual polyamides.

[0061] In the present invention, preferably, sort in the order of the magnitude of the comprehensive score, and select the top 50 polyamides with the highest comprehensive scores from the structural formulas ranked in the top 5%.

[0062] In the present invention, compared with the machine learning model constructed by combining a graph convolutional neural network (GCN) and a Gaussian process regression (GPR), the machine learning model constructed by combining the graph convolutional neural network (GCN) and a multi-layer perceptron (MLP) in the present application has higher prediction accuracy and stronger generalization ability.

[0063] The present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the training method of the model as described above, or the prediction method of the polyamide performance parameters as described above is implemented.

[0064] The present invention also provides a computer-readable storage medium, on which a computer program is stored. It is characterized in that when the computer program is executed by a processor, the training method of the model as described above, or the prediction method of the polyamide performance parameters as described above is implemented.

[0065] The present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the training method of the model as described above, or the prediction method of the polyamide performance parameters as described above is implemented.

[0066] On the basis of conforming to the common knowledge in the art, the above preferred conditions can be arbitrarily combined to obtain various preferred examples of the present invention.

[0067] The reagents and raw materials used in the present invention are all commercially available.

[0068] The positive and progressive effects of the present invention are as follows:

[0069] (1) By introducing KAN into the GCN machine learning model, the present invention realizes the efficient approximation and flexible modeling of high-dimensional complex functions through a dynamic activation function; the dynamic activation function adaptively adjusts according to the input features, avoiding the model expression limitations that may be caused by using a fixed activation function in a traditional multi-layer perceptron (MLP). This dynamic characteristic enables KAN to have higher accuracy and generalization ability when processing data with strong non-linear relationships, and the visualization of the KAN framework and the activation function has higher interpretability compared to traditional neural networks.

[0070] (2) The performance prediction method of the present invention has high prediction accuracy for the performance parameters of polyamide; specifically, in terms of predicting the glass transition temperature, the R of the test set 2 is 0.98 and the MAE is 12 °C; in terms of predicting the melting point, the R of the test set 2 is 0.90 and the MAE is 17 °C; in terms of predicting the tensile modulus, the R of the test set 2 is 0.93 and the MAE is 0.14 GPa.

[0071] (3) The present invention uses a machine learning model to perform high-throughput performance prediction on a large number of combinatorially generated virtual structural formulas of polyamides, and screens out a series of preferred new structures with high glass transition temperature, high melting point, and high tensile modulus, thereby providing a material design method based on the combination of virtual screening and machine learning for the high-performance optimization of polyamide materials and the R & D experiments of new polyamides. Description of the Drawings

[0072] Figure 1 It is the data distribution in the dataset corresponding to the glass transition temperature of the polyamide in Example 1;

[0073] Figure 2 It is the data distribution in the dataset corresponding to the melting point of the polyamide in Example 1;

[0074] Figure 3 It is the data distribution in the dataset of the tensile modulus of the polyamide and polyimide in Example 1;

[0075] Figure 4 It is the flow chart for constructing a polyamide performance parameter prediction model using the method of combining GCN and KAN in Example 1;

[0076] Figure 5 It is the prediction accuracy of the model for the glass transition temperature of the polyamide in Example 1;

[0077] Figure 6 It is the prediction accuracy of the model for the melting point of the polyamide in Example 1;

[0078] Figure 7 It is the prediction accuracy of the model for the tensile modulus of the polyamide in Example 1;

[0079] Figure 8 It is the KAN visualization architecture and some dynamic activation functions corresponding to the model for the glass transition temperature of the polyamide in Example 1;

[0080] Figure 9 It is the KAN visualization architecture and some dynamic activation functions corresponding to the model for the melting point of the polyamide in Example 1;

[0081] Figure 10 It is the KAN visualization architecture and some dynamic activation functions corresponding to the model for the tensile modulus of the polyamide in Example 1;

[0082] Figure 11 It is the schematic diagram of the obtaining process of the virtual polyamide structural formula in Example 1;

[0083] Figure 12 It is some diamine monomers used to combine into the virtual polyamide in Example 1;

[0084] Figure 13 It is a partial diacid monomer used to form a virtual polyamide in Example 1;

[0085] Figure 14 It is a comparison chart of experimental values and predicted values for verifying the model accuracy in Effect Example 1;

[0086] Figure 15 It is the prediction performance of different modeling methods in Example 1 and Comparative Examples 1-2 for the glass transition temperature of polyamide; Figure 15 Part a of it is the prediction effect R of different modeling methods on the glass transition temperature of polyamide 2 , Figure 15 Part b of it is the prediction effect MAE of different modeling methods on the glass transition temperature of polyamide;

[0087] Figure 16 It is the prediction performance of different modeling methods in Example 1 and Comparative Examples 1-2 for the melting point of polyamide; Figure 16 Part a of it is the prediction effect R of different modeling methods on the melting point of polyamide 2 , Figure 16 Part b of it is the prediction effect MAE of different modeling methods on the melting point of polyamide;

[0088] Figure 17 It is the prediction performance of different modeling methods in Example 1 and Comparative Examples 1-2 for the tensile modulus of polyamide; Figure 17 Part a of it is the prediction effect R of different modeling methods on the tensile modulus of polyamide 2 , Figure 17 Part b of it is the prediction effect MAE of different modeling methods on the tensile modulus of polyamide;

[0089] Figure 18 It is a three-dimensional distribution diagram of the glass transition temperature, melting point and tensile modulus of the candidate polyamide structure predicted in Example 2.

[0090] Figures 19 - 23 It is the chemical structure of the polyamide repeating unit that meets the top 5% of the comprehensive scores of high glass transition temperature, high melting point and high tensile modulus screened in Example 2. Detailed implementation mode

[0091] The present invention will be further described below by way of examples, but the present invention is not limited to the scope of the described examples. For the experimental methods without specific conditions noted in the following examples, they are carried out according to conventional methods and conditions, or selected according to the product instructions.

[0092] Example 1

[0093] This example discloses a training method for a model for predicting polyamide performance parameters, which includes the following steps:

[0094] S1. Collect the performance parameters of different structural formulas of polyamides and their corresponding glass transition temperatures, melting points, and tensile moduli from the literature as the first dataset; collect the different structural formulas of polyimides and their corresponding tensile moduli from the literature as the second dataset to supplement the first dataset. The data distributions in the datasets of glass transition temperature, melting point, and tensile modulus are as Figures 1 - 3 shown. Figure 1 This is the data distribution in the dataset corresponding to the glass transition temperature of the polyamide in this example; Figure 2 This is the data distribution in the dataset corresponding to the melting point of the polyamide in this example; Figure 3 This is the data distribution in the datasets of the tensile moduli of the polyamide and polyimide in this example.

[0095] Among them, the dataset of glass transition temperature contains 501 experimental data of polyamides, the dataset of melting point contains 900 experimental data of polyamides, and the dataset of tensile modulus contains 131 experimental data of polyamides and 386 experimental data of polyimides.

[0096] S2. Perform format normalization and data cleaning on the collected datasets. Specifically:

[0097] Perform format normalization on the SMILES representations of the collected datasets, unify them into a format based on the repeating unit, and add "*" identifiers at both ends of the molecule of the repeating unit;

[0098] For the glass transition temperature, melting point, and tensile modulus data of polyamides with the same structural formula collected from different literatures, only retain the highest value among them and remove the others.

[0099] S3. Use the cleaned dataset as the input unit to construct a machine learning model combining GCN and KAN, as Figure 4 shown; Figure 4 This is the flowchart for constructing a polyamide performance parameter prediction model using the method of combining GCN and KAN in Example 1. Specifically, this process includes the following steps:

[0100] a. For the two types of datasets of glass transition temperature and melting point, divide them according to the ratio of training set:test set = 9:1;

[0101] For the tensile modulus dataset, divide it according to the ratio of training set:test set = 8:2, and supplement all the polyimide datasets to the training set;

[0102] b. Represent the molecular structural formula in the form of a molecular graph. Specifically, it is obtained by representing the atomic information through the atomic number, aromaticity, total number of connected chemical bonds, number of hydrogen atoms, hybridization type, formal charge, number of radical electrons, chirality label, implicit valence, explicit valence, and size of the smallest ring of each atom in the molecule; representing the bond information through the bond type, whether it belongs to a ring, whether it is conjugated, stereochemical properties, direction, and aromaticity of each bond in the molecule; and representing the adjacent atoms through the indices of the two atoms connected by the bond.

[0103] c. Use the feature vector encoded by the molecular graph as the model input and the corresponding performance as the model output. Add KAN as an intermediate layer in the model and obtain the performance prediction model of polyamide through forward deduction and backpropagation training. Figure 5 This is the prediction accuracy of the model for the glass transition temperature of polyamide in this example; from Figure 5 the results, for the model of the glass transition temperature of polyamide, the R 2 of the training set is 0.98, the MAE is 8 °C, and the R 2 of the test set is 0.98, the MAE is 12 °C. Figure 6 This is the prediction accuracy of the model for the melting point of polyamide in this example; from Figure 6 the results, for the model of the melting point of polyamide, the R 2 of the training set is 0.96, the MAE is 8 °C, and the R 2 of the test set is 0.90, the MAE is 17 °C. Figure 7 This is the prediction accuracy of the model for the tensile modulus of polyamide in this example; from Figure 7 the results, for the model of the tensile modulus of polyamide, the R 2 of the training set is 0.98, the MAE is 0.06 GPa, and the R 2 of the test set is 0.93, the MAE is 0.14 GPa.

[0104] Figure 8 This is the KAN visualization architecture and partial dynamic activation functions corresponding to the model of the glass transition temperature of polyamide in this example; Figure 9 This is the KAN visualization architecture and partial dynamic activation functions corresponding to the model of the melting point of polyamide in this example; Figure 10 This is the KAN visualization architecture and partial dynamic activation functions corresponding to the model of the tensile modulus of polyamide in this example; in Figures 8 - 10 , the number of nodes represents the input and output dimensions of each layer, each edge represents an activation function, and the transparency of the edge represents the importance of the activation function. The lower the transparency, the more important the activation function. The upper part of the figure is the activation function graph of the last layer.

[0105] Among them, the GCN has 3 convolutional layers and 3 dropout layers. Each dropout layer is respectively set after each convolutional layer, and the dropout probability of the dropout layer is 0 - 0.2; each convolutional layer is a dynamic edge-conditioned filter, and the dynamic edge-conditioned filter includes an edge feature dimension, a node feature dimension, the input dimension and the output dimension of the convolutional layer; the input dimension of the first convolutional layer is 42, the output dimension is 1000, the node feature dimension is 42, and the edge feature dimension is 15; the input dimension of the second convolutional layer is 1000, the output dimension is 250, the edge feature dimension is 15, and the node feature dimension is a matrix of 1000 * 250 weighted to 1000 node features; the input dimension of the third convolutional layer is 250, the output dimension is 40, the node feature dimension is 250, and the edge feature dimension is 15.

[0106] KAN includes 3 hidden layers; the input dimension of the first hidden layer is 40, and the output dimension is 20; the input dimension of the second hidden layer is 20, and the output dimension is 5; the input dimension of the third hidden layer is 5, and the output dimension is 1; each hidden layer includes a dynamic activation function; the grid value of each dynamic activation function is 5; the k value of each dynamic activation function is 3.

[0107] This embodiment also discloses a method for predicting polyamide performance parameters, which includes the following steps: input the structural formula of the polyamide to be tested into the above prediction model to obtain the performance data of the glass transition temperature, melting point and tensile modulus of the polyamide to be tested.

[0108] The polyamide to be tested includes polyamides with known structural formulas and virtual polyamides;

[0109] Among them, the polyamide with a known structural formula can be a polyamide with both a known structural formula and performance parameters, or a polyamide with a known structural formula but unknown performance parameters. When analyzing polyamides with both known structural formulas and performance parameters, the prediction results can illustrate the accuracy of the model.

[0110] Among them, the method for obtaining virtual polyamides includes the following steps:

[0111] ss1. Define the monomer types in the polyamide structural formula; among them, the monomer types include diamine monomers and diacid monomers; the chemical structures of diamine monomers and diacid monomers can be defined as the gene structure of polyamides, and can also be called diamine genes and diacid genes.

[0112] ss2. Search for diamine monomers and diacid monomers on the public databases SciFinder, PubChem, and ChemSpider to obtain 361 diamine monomers and 309 diacid monomers;

[0113] ss3. By using Python programming to combine diamine monomers and diacid monomers, 111,549 structural formulas of virtual polyamides are obtained, as shown in Figure 11 shown Figure 11 This is a schematic diagram of the process for obtaining the structural formula of the virtual polyamide in this embodiment. Figure 12 These are some of the diamine monomers used to form the virtual polyamide in this embodiment; Figure 13 These are some of the diacid monomers used to form the virtual polyamide in this embodiment.

[0114] Effect Example 1

[0115] The test object of this effect example is the prediction model established by the polyamide performance parameter prediction method in Example 1 to verify the reliability of the prediction model:

[0116] Search for a polyamide called PA10T, which is a polyamide not in the model but with known structural formula and performance parameters. Its structural formula is as follows:

[0117]

[0118] Figure 14 This is a comparison chart of experimental values and predicted values for verifying the model accuracy in this effect example. Among them, the experimental value and model predicted value of the glass transition temperature of PA10T are 110°C and 106°C respectively, the experimental value and model predicted value of the melting point of PA10T are 310°C and 307°C respectively, and the experimental value and model predicted value of the tensile modulus of PA10T are 1.30 GPa and 1.19 GPa respectively. It can be seen that the predicted values of the three models are all relatively close to the experimental values, verifying the reliability of the prediction method in Example 1.

[0119] Comparative Example 1

[0120] This comparative example discloses a method for predicting polyamide performance parameters, which uses the same dataset as in step S3 of Example 1 for modeling;

[0121] The difference between this comparative example and Example 1 is that a machine learning model (abbreviated as GC-MLP model) is constructed by combining a graph convolutional neural network (GCN) and a multi-layer perceptron (MLP), and the graph convolutional operation is the same as in Example 1. Other conditions are the same as in Example 1.

[0122] The MLP performs a non-linear transformation on the molecular graph features extracted by the GCN through a fixed ReLU activation function. Since the MLP depends on a fixed non-linear activation function, it is difficult to adapt to more complex non-linear relationships in the data; in addition, each layer of the MLP contains a large number of parameters, resulting in excessive consumption of computing resources.

[0123] Comparative Example 2

[0124] This comparative example discloses a method for predicting the performance parameters of polyamide, which uses the same data set as in step S3 of Example 1 for modeling;

[0125] The difference between this comparative example and Example 1 is that a machine learning model (abbreviated as GC-GPR model) is constructed by combining a graph convolutional neural network (GCN) with Gaussian process regression (GPR), and the graph convolution operation therein is the same as that in Example 1. Other conditions are the same as those in Example 1.

[0126] GPR models the molecular graph features as a Gaussian process, uses a kernel function to calculate the covariance matrix to describe the similarity between input data points, and performs non-linear fitting through Bayesian inference. As the amount of data increases, calculating the covariance matrix requires a large amount of time, resulting in a very slow training process and high memory consumption. In addition, GPR is very sensitive to the selection of the kernel function and the adjustment of hyperparameters, and an inappropriate kernel function will lead to a poor fitting effect.

[0127] Effect Example 2

[0128] The test object of this effect example is the prediction model recommended by the polyamide performance parameter prediction methods in Example 1 and Comparative Examples 1-2, and the reliability of this prediction model is verified;

[0129] Specifically, the prediction models obtained in Example 1 and Comparative Examples 1-2 are randomly run 100 times, and the R 2 (coefficient of determination) and MAE (mean absolute error) are as Figures 15 - 17 shown.

[0130] Figure 15 shows the prediction performance of different modeling methods in Example 1 and Comparative Examples 1-2 for the glass transition temperature of polyamide; Figure 15 Part a of shows the prediction effect R 2 of different modeling methods on the glass transition temperature of polyamide, Figure 15 and part b of shows the prediction effect MAE of different modeling methods on the glass transition temperature of polyamide. It can be seen that for the glass transition temperature prediction model, the average R 2 and MAE of the test set of the GC-KAN model of this application are 0.96 and 14 °C, respectively. R 2 is higher than the average value of 0.93 of the GC-MLP model in Comparative Example 1, and MAE is lower than the average value of 16 °C of the GC-MLP model. R 2 is higher than the average value of 0.95 of the GC-GPR model in Comparative Example 2, and MAE is equal to the average value of the GC-GPR model.

[0131] Figure 16The prediction performance of different modeling methods in Example 1 and Comparative Examples 1-2 for the melting point of polyamide; Figure 16 Part a of Figure 16 shows the prediction effect R of different modeling methods on the melting point of polyamide 2 , Figure 16 Part b of Figure 16 shows the prediction effect MAE of different modeling methods on the melting point of polyamide. It can be seen that for the melting point prediction model, the average R of the test set of the GC-KAN model of this application 2 and MAE are 0.83 and 21 °C. The R 2 is higher than the average value of 0.73 of the GC-MLP model in Comparative Example 1, and the MAE is lower than the average value of 27 °C of the GC-MLP model. The R 2 is equal to the average value of the GC-GPR model in Comparative Example 2, and the MAE is equal to the average value of the GC-GPR model.

[0132] Figure 17 The prediction performance of different modeling methods in Example 1 and Comparative Examples 1-2 for the tensile modulus of polyamide; Figure 17 Part a of Figure 17 shows the prediction effect R of different modeling methods on the tensile modulus of polyamide 2 , Figure 17 Part b of Figure 17 shows the prediction effect MAE of different modeling methods on the tensile modulus of polyamide. It can be seen that for the tensile modulus prediction model, the average R of the test set of the GC-KAN model of this application 2 and MAE are 0.71 and 0.25 GPa. The R 2 is higher than the average value of 0.65 of the GC-MLP model in Comparative Example 1, and the MAE is lower than the average value of 0.28 GPa of the GC-MLP model. The R 2 is higher than the average value of 0.68 of the GC-GPR model in Comparative Example 2, and the MAE is equal to the average value of the GC-GPR model.

[0133] From the above analysis, it can be obtained that in the prediction comparison of the above three polyamide performance parameters, the GC-KAN model has higher accuracy and stronger generalization ability, which proves the advantages of the method proposed in the present invention. Among them, the data samples of the tensile modulus of polyamide are less, and adding the data samples of the tensile modulus of polyimide can improve the prediction ability of the model.

[0134] Example 2

[0135] This example provides a structural design method for polyamide; in this example, based on the virtual polyamide structural formula formed in the prediction method of Example 1 and the model trained in Example 1 for predicting the performance parameters of polyamide, a new polyamide structure with both high glass transition temperature, high melting point and high tensile modulus was quickly screened; the specific research process includes the following steps:

[0136] sss1. Use the model trained in Example 1 to predict the 111,549 virtual polyamide candidate structural formulas in Example 1, and obtain the performance data of the glass transition temperature, melting point, and tensile modulus of each candidate structural formula.

[0137] sss2. Use a weighted evaluation function (also known as a scalarization function, which is the calculation formula for the following Score) to represent the above three performances with a comprehensive score, and use the comprehensive score to color the material performance space, as Figure 8 shown. Figure 18 This is the three-dimensional distribution diagram of the glass transition temperature, melting point, and tensile modulus of the candidate polyamide structure predicted in this example; the colors in the figure from blue to red represent the comprehensive performance from low to high.

[0138]

[0139] Among them, Score is the comprehensive score, T g , T m , and E represent the glass transition temperature, melting point, and tensile modulus respectively; T g min and T g max are the minimum and maximum values of the glass transition temperature among all virtual polyamides respectively, T m min and T m max are the minimum and maximum values of the melting point among all virtual polyamides respectively, and E min and E max are the minimum and maximum values of the tensile modulus among all virtual polyamides respectively.

[0140] Sort them in ascending order according to the magnitude of the comprehensive score, and select the top 50 new polyamides from the top 5% of the preferred structures. Their chemical structures are as Figures 19 - 23 shown; Figures 19 - 23 This is the chemical structure of the repeating unit of the polyamide with the top 5% of the comprehensive score that meets the high glass transition temperature, high melting point, and high tensile modulus screened in this example. In the figure, the top 50 polyamides are sorted from high to low according to the comprehensive score and are marked as No.1 - No.50; Tg is the glass transition temperature, Tm is the melting point, and E is the tensile modulus.

[0141] The new polyamide has excellent high-temperature resistance and mechanical properties and has broad application prospects in the aerospace field.

[0142] Example 3

[0143] This embodiment discloses an electronic device, which includes a memory, a processor, and a computer program stored in the memory and configured to run on the processor. When the processor executes the computer program, it implements the training method for the model for predicting polyamide performance parameters and the prediction method for polyamide performance parameters provided in the above-mentioned Embodiment 1. The electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0144] The electronic device may be presented in the form of a general-purpose computing device. For example, it may be a server device. The components of the electronic device may include, but are not limited to: the above-mentioned at least one processor, the above-mentioned at least one memory, and a bus connecting different system components (including the memory and the processor).

[0145] The bus includes a data bus, an address bus, and a control bus.

[0146] The memory may include volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read-only memory (ROM).

[0147] The memory may also include program tools (or utilities) having a set (at least one) of program modules. Such program modules include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.

[0148] The processor executes various functional applications and data processing by running the computer program stored in the memory, such as the training method for the model for predicting polyamide performance parameters and the prediction method for polyamide performance parameters provided in the above-mentioned Embodiment 1.

[0149] The electronic device may also communicate with one or more external devices (such as a keyboard, a pointing device, etc.). Such communication may be carried out through an input / output (I / O) interface. Moreover, the electronic device may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter. As shown in the figure, the network adapter communicates with other modules of the electronic device through the bus. It should be understood that other hardware and / or software modules may be used in combination with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems, etc.

[0150] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple units / modules.

[0151] Embodiment 4

[0152] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the training method of the model for predicting polyamide performance parameters and the prediction method of polyamide performance parameters provided in Embodiment 1 above.

[0153] Among them, the more specific forms that the readable storage medium can adopt may include but are not limited to: portable disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0154] Embodiment 5

[0155] This embodiment provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the training method of the model for predicting polyamide performance parameters and the prediction method of polyamide performance parameters provided in Embodiment 1 above.

[0156] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages. The program code can be executed completely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or completely on a remote device.

[0157] Although the specific embodiments of the present application are described above, those skilled in the art should understand that this is only an example. The protection scope of the present disclosure is defined by the appended claims. Without departing from the principles and essence of the present disclosure, those skilled in the art can make various changes or modifications to these embodiments, but these changes and modifications all fall within the protection scope of the present disclosure.

Claims

1. A training method for a model for predicting polyamide performance parameters, characterized in that, It includes the following steps: S1. Obtain a data set, where the data set includes a first data set, and the first data set includes different structural formulas of polyamide and performance parameters corresponding to each structural formula; the performance parameters are one or more of glass transition temperature, melting point, and tensile modulus; S2. Input the data set into a graph convolutional neural network, and use KAN as the intermediate layer to train the model.

2. The training method of the model according to claim 1, wherein The training method of the model satisfies one or more of the following conditions: ① When the performance parameter is the glass transition temperature, the data volume of the first data set is 100 - 1000, preferably 501; ② When the performance parameter is the melting point, the data volume of the first data set is 100 - 1000, preferably 900; ③ When the performance parameter is the tensile modulus, the data volume of the first data set is 100 - 1000, preferably 131; ④ The first data set includes a training set and a test set, and the data volume ratio of the training set to the test set is 9:1 or 8:2; Preferably, when the performance parameter is the glass transition temperature, the data volume ratio of the training set to the test set is 9:1; Preferably, when the performance parameter is the melting point, the data volume ratio of the training set to the test set is 9:1; Preferably, when the performance parameter is the tensile modulus, the data volume ratio of the training set to the test set is 8:2; ⑤ The data set further includes a second data set for supplementing the first data set; the second data set includes different structural formulas of polyimide and performance parameters corresponding to each structural formula, and the performance parameters are one or more of glass transition temperature, melting point, and tensile modulus; preferably, the data volume of the second data set is 386.

3. The training method of the model according to claim 1, wherein The training method of the model satisfies one or more of the following conditions: ① In step S1, after obtaining the data set, it further includes a format normalization processing step; preferably, the format normalization processing step includes using the repeating unit as the basic format and adding identifiers at both ends of the molecule of the repeating unit; ② In step S1, after obtaining the data set, it further includes a cleaning step; preferably, the cleaning step includes removing different performance parameters of polyamide with the same structural formula and only retaining the highest value of the performance parameter; ③ In step S2, the graph convolutional neural network includes 3 convolutional layers; Preferably, the graph convolutional neural network further includes 3 dropout layers, and each dropout layer is respectively arranged after each convolutional layer; more preferably, the dropout probability of each dropout layer is 0 - 0.2; Preferably, each convolutional layer is a dynamic edge-conditioned filter, and the dynamic edge-conditioned filter includes an edge feature dimension, a node feature dimension, an input dimension, and an output dimension of the convolutional layer; more preferably, the input dimensions of the convolutional layer are respectively 40 - 1000; more preferably, the output dimensions of the convolutional layer are respectively 40 - 1000; ④In step S2, the graph convolutional neural network is used to obtain the atomic information, bond information, and adjacent atomic positions of the polyamide; preferably, the atomic information is one or more of atomic number, aromaticity, total number of connected chemical bonds, number of hydrogen atoms, and hybridization type; preferably, the bond information is bond type, stereochemical property, direction, and aromaticity; preferably, the adjacent atomic positions are the indices of the two atoms connected by the bond. ⑤In step S2, the KAN includes 3 hidden layers; preferably, the input dimension of each hidden layer is 5 - 50; preferably, the output dimension of each hidden layer is 5 - 50; preferably, each hidden layer includes a dynamic activation function; more preferably, the grid value of the dynamic activation function is 5; preferably, the k value of the dynamic activation function is 3.

4. A method for predicting the performance parameters of polyamide, characterized in that, It includes the following steps: input the structural formula of the polyamide to be tested into the prediction model to obtain the performance parameters of the polyamide to be tested; the prediction model is obtained by using the training method of the model according to any one of claims 1 - 3. The types of the performance parameters of the polyamide to be tested are the same as those in the first dataset.

5. The prediction method of polyamide performance parameters according to claim 4, characterized in that The polyamide to be tested includes a polyamide with a known structural formula or a virtual polyamide. Preferably, the virtual polyamide is obtained by combining diamine monomers and diacid monomers through a chemical reaction template; more preferably, the number of both the diamine monomers and the diacid monomers is 200 - 400, for example, the number of diamine monomers is 361, and the number of diacid monomers is 309. Preferably, the number of the virtual polyamides is more than 100,000, more preferably 100,000 - 120,000, for example, 111,549.

6. A prediction system for polyamide performance parameters, characterized in that, It includes a prediction module, and the prediction module is used to input the structural formula of the polyamide to be tested into the prediction model to obtain the performance parameters of the polyamide to be tested; the prediction model is obtained by using the training method of the model according to any one of claims 1 - 3.

7. A method for the structural design of a polyamide, characterized in that, It includes the following steps: s1. Use the prediction model to predict the candidate structural formulas of the virtual polyamides to obtain the performance parameters of the candidate structural formulas; the prediction model is obtained by using the training method of the model according to any one of claims 1 - 3; the performance parameters of the candidate structural formulas are glass transition temperature, melting point, and tensile modulus. s2. Calculate the comprehensive scores of the candidate structural formulas of each virtual polyamide and sort the comprehensive scores to obtain the structural formulas of the polyamides. Preferably, the candidate structural formulas are obtained by combining diamine monomers and diacid monomers through a chemical reaction template; more preferably, the number of both the diamine monomers and the diacid monomers is 200 - 400, for example, the number of diamine monomers is 361, and the number of diacid monomers is 309. Preferably, the number of the candidate structural formulas is more than 100,000, more preferably 100,000 - 120,000, for example, 111,549. Preferably, the method for calculating the comprehensive score includes calculating by using a weighted evaluation function.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the training method of the model described in any one of claims 1-3, or the prediction method of the polyamide performance parameters described in claim 4 or 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the training method of the model described in any one of claims 1-3, or the prediction method of the polyamide performance parameters described in claim 4 or 5.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the training method of the model described in any one of claims 1-3, or the prediction method of the polyamide performance parameters described in claim 4 or 5.