A method for predicting performance of polycarboxylate superplasticizer based on heterogeneous graph neural network

CN122314148BActive Publication Date: 2026-09-11FUJIAN AGRI & FORESTRY UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610787499.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-11
Estimated Expiration
2046-06-03

AI Technical Summary

Technical Problem

该类方法通常将各影响因素视为相互独立或弱相关变量,难以系统刻画单体组成、工艺条件和测试条件之间的多类型关联关系,也难以充分挖掘不同样本之间由于共享单体、共享工艺或共享测试环境而形成的潜在结构信息

Benefits of technology

1、本发明提出了一种基于异构图神经网络的聚羧酸减水剂性能预测方法,能够将聚羧酸减水剂样本、单体组成、合成工艺参数以及测试条件统一表示为异构图结构,克服了传统表格型机器学习方法难以显式表达多类型实体及其关联关系的不足,提高了聚羧酸减水剂数据的结构化表征能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122314148B_ABST
    Figure CN122314148B_ABST
Patent Text Reader

Abstract

The application discloses a polycarboxylate superplasticizer performance prediction method based on a heterogeneous graph neural network, and belongs to the technical field of artificial intelligence and civil engineering materials. The method collects polycarboxylate superplasticizer monomer composition, synthesis process, test conditions and corresponding net paste fluidity data and pre-processes, constructs a heterogeneous graph structure with a sample as a core node and containing multiple types of nodes and relationship edges, builds a neural network model containing multiple layers of heterogeneous graph convolution layers, establishes a nonlinear mapping relationship between multiple factors and performance through supervised training, and realizes net paste fluidity prediction. The application can effectively depict the coupling and correlation of multiple factors, improve the prediction accuracy and generalization ability, provide technical support for intelligent design and optimization of superplasticizers, and reduce the research and development cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and civil engineering materials technology, and in particular to a method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks. Background Technology

[0002] Polycarboxylate superplasticizers are among the most widely used high-performance superplasticizers in modern concrete. Due to their high water reduction rate, good slump retention, and strong adaptability, they are widely applied in engineering material systems such as high-performance concrete, self-compacting concrete, mass concrete, and ultra-high-performance concrete. The performance of polycarboxylate superplasticizers is closely related to their molecular composition, synthesis process, and application conditions. The types and proportions of different acid monomers, macromonomers, and functional monomers, as well as process parameters such as the initiation system, reaction temperature, reaction time, and degree of neutralization, all significantly affect the dispersion performance, slump retention, air-containing properties, and compatibility with cement systems of the superplasticizer.

[0003] However, the design and performance regulation of polycarboxylate superplasticizers involve multi-factor, multi-level coupling effects. On the one hand, structural factors such as monomer type, monomer ratio, macromonomer molecular weight, and functional group composition jointly determine the molecular skeleton and side chain characteristics of the superplasticizer. On the other hand, polymerization process conditions further affect its molecular weight distribution, functional group retention, and final dispersion behavior. Furthermore, the test results of superplasticizer performance are also affected by external conditions such as cement type, water-cement ratio, dosage, and test system. Therefore, the performance formation mechanism of polycarboxylate superplasticizers is essentially a complex nonlinear coupling process between "sample composition - synthesis process - test conditions - performance response," posing a significant challenge to the efficient design and accurate prediction of superplasticizer performance.

[0004] Currently, the molecular design, formulation optimization, and performance evaluation of polycarboxylate superplasticizers still mainly rely on experimental mixing, empirical judgment, or conventional statistical analysis. These methods typically treat influencing factors as independent or weakly correlated variables, making it difficult to systematically characterize the various relationships between monomer composition, process conditions, and testing conditions. They also struggle to fully exploit the potential structural information formed between different samples due to shared monomers, processes, or testing environments. Therefore, when dealing with superplasticizer systems with numerous data dimensions, complex variable types, implicit sample relationships, and significant nonlinear performance responses, traditional methods often suffer from limited predictive accuracy, insufficient generalization ability, and weak scalability.

[0005] Therefore, this invention proposes a performance prediction method for polycarboxylate superplasticizers based on heterogeneous graph neural networks, which can achieve unified modeling of complex multi-source information of polycarboxylate superplasticizers, improve the accuracy and stability of performance prediction, and provide reliable support for the intelligent design and optimization of superplasticizers. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks. This method can effectively model various relationships between polycarboxylate superplasticizer samples, monomer composition, synthesis process, and testing conditions, achieving high-precision prediction of superplasticizer performance and providing technical support for the intelligent design and optimization of polycarboxylate superplasticizers.

[0007] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention proposes a method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks, comprising the following steps: S1. Collect sample data of polycarboxylate superplasticizer, including superplasticizer monomer composition information, synthesis process parameters, performance test conditions and corresponding performance indicators. Preprocess the raw data of different types separately. After preprocessing, use superplasticizer monomer composition information, synthesis process parameters and test conditions as input data for the heterogeneous graph model, and use the corresponding performance indicators as model output. S2. Based on the polycarboxylate superplasticizer sample data obtained in step S1, construct a heterogeneous graph structure for predicting the fluidity of the paste. The heterogeneous graph uses the polycarboxylate superplasticizer sample node as the core node and introduces monomer nodes, process nodes, and test condition nodes. Different types of relational edges describe the correlation between the sample and monomer composition, synthesis process, and test conditions, thereby forming a graph structure data suitable for heterogeneous graph neural network modeling. S3. Based on the constructed heterogeneous graph data structure, a heterogeneous graph neural network model containing multiple heterogeneous graph convolutional layers is constructed. The heterogeneous graph neural network model learns the nonlinear mapping relationship between the composition, synthesis process and test conditions of polycarboxylate superplasticizer and the fluidity of the paste by performing information propagation and feature aggregation on the multi-type relationships between sample nodes, monomer nodes, process nodes and test condition nodes, thereby realizing the predictive output of the performance indicators of polycarboxylate superplasticizer. S4. Based on the constructed heterogeneous graph neural network model, the heterogeneous graph neural network model is trained under supervision using the polycarboxylate superplasticizer sample data obtained in step S1 and the heterogeneous graph input data constructed in step S2. By calculating the error between the model's predicted output and the measured paste fluidity label of the sample, the model parameters are updated using gradient backpropagation and parameter iterative optimization, thereby establishing a mapping relationship between the polycarboxylate superplasticizer sample composition, synthesis process, test conditions and paste fluidity. After training, the model can be used to predict the paste fluidity of any input polycarboxylate superplasticizer material sample.

[0008] S5. After completing the training of the heterogeneous graph neural network model, input the polycarboxylate superplasticizer sample to be predicted into the trained heterogeneous graph neural network model to predict the flowability of the paste of the sample to be predicted, and obtain the corresponding performance prediction results.

[0009] Furthermore, in step S1, the monomer composition information includes the types of unsaturated monomers, the types and molecular weights of polyether macromonomers, the types of functional monomers and their proportions; the synthesis process parameters include the dosage of initiator, chain transfer agent, reducing agent and oxidant, reaction temperature and reaction time; the performance test conditions include cement type, water-cement ratio and water-reducing agent dosage; and the performance index is the fluidity of the cement paste.

[0010] Furthermore, in step S1, different types of raw data are preprocessed separately, specifically: different types of raw data are cleaned, normalized, discretely encoded or numerically standardized, and all types of data are uniformly numbered and structured to form sample node information, individual node information, process node information and test condition node information.

[0011] Furthermore, step S1 specifically includes: Step S11: Obtain the polycarboxylate superplasticizer sample dataset, assuming a total of N One water-reducing agent sample, denoted as: ,in, Indicates the first i One polycarboxylate superplasticizer sample, for each sample Collect information on the corresponding monomer composition, synthesis process parameters, test conditions, and corresponding paste fluidity values. Step S12: For each polycarboxylate superplasticizer sample, collect its monomer composition information, including the types of unsaturated monomers, the types and molecular weights of polyether macromonomers, the types of functional monomers and their ratio parameters. All monomers constitute a monomer node set. ,in, Indicates the first j Individual node, This represents the total number of all individual nodes; For a single node Its node feature vector This is used to characterize the intrinsic properties of the monomer itself, including monomer type, molecular weight, and functional group type; wherein, d M Indicates the feature dimension of a single node; The feature matrix of a single node can be constructed from the feature vectors of all individual nodes. :

[0012] For any sample node and single node If the monomer Used for samples The synthesis is performed at the sample node. With a single node Establish a compositional relationship edge between them The edge features corresponding to the relation edges are denoted as: ,in, Indicates monomer In the sample The proportioning parameters in the formula include, but are not limited to, molar ratio, feed ratio, and mass ratio; The compositional relationship between all sample nodes and individual nodes is represented as a sample-individual weighted correlation matrix. Defined as: , where matrix elements Defined as:

[0013] The Used to characterize the compositional relationship and connection strength between different polycarboxylate superplasticizer samples and each monomer node, and as input to the relationship matrix between sample nodes and monomer nodes in heterogeneous graph neural network; Step S13: For each polycarboxylate superplasticizer sample, collect the synthesis process parameters, including the dosage of initiator, chain transfer agent, reducing agent and oxidant, reaction temperature and time; wherein, the dosage of initiator, chain transfer agent, reducing agent and oxidant are all converted into percentages of the total mass of all monomers; All process parameters constitute the process node set. ,in, Indicates the first j Each process node Indicates the total number of process nodes; For process nodes Its node feature vector This is used to characterize the inherent properties of the process node itself, including process parameter category, reagent category, whether it belongs to a redox system, and parameter value type; wherein, Represents the feature dimension of the process node; The feature matrix of all process nodes can be constructed from their feature vectors. :

[0014] For any polycarboxylate superplasticizer sample node and process nodes If process node Used for samples The synthesis process occurs at the sample node. With process nodes Establish a process relationship between them The edge features corresponding to the relation edge ,in, express In the sample The process parameters in the data, when the process node is an initiator, chain transfer agent, reducing agent, or oxidizing agent node, Indicates the corresponding reagent in the sample Its percentage of the total monomer mass; when the process node is a temperature node or a time node. This indicates the corresponding reaction temperature or reaction time value; The technological relationships between all sample nodes and process nodes are represented as a sample-process weighted correlation matrix. Defined as: ; where matrix elements Defined as:

[0015] The Used to characterize the process relationship and connection strength between different polycarboxylate superplasticizer samples and each process node, and as input to the relationship matrix between sample nodes and process nodes in heterogeneous graph neural network; Step S14: Collect performance test parameters for each polycarboxylate superplasticizer sample, including cement type, water-cement ratio, and superplasticizer dosage. All test condition parameters constitute the test condition node set. ;in, Indicates the first j Test condition nodes Indicates the total number of test condition nodes; For test condition nodes Its node feature vector This is used to characterize the inherent properties of the test condition node itself, including the condition parameter category, parameter value type, material category, and test system category; wherein, Indicates the feature dimension of the test condition node; The feature matrix of all test condition nodes can be constructed from their feature vectors. :

[0016] For any polycarboxylate superplasticizer sample node and test condition nodes If the test conditions Used for samples The synthesis process occurs at the sample node. With test condition nodes Establish a test condition relationship edge between them The edge features corresponding to the relation edge ,in, Indicates test condition node In the sample The corresponding test condition parameter values; when the test condition node is a cement type node. This indicates the corresponding category code value; when the test condition node is the water-cement ratio node or the water-reducing agent dosage node... This represents the corresponding numerical parameter; The test condition relationships between all sample nodes and test condition nodes are represented as a sample-test condition weighted correlation matrix. Defined as: , where matrix elements Defined as:

[0017] The Used to characterize the test condition relationship and connection strength between different polycarboxylate superplasticizer samples and each test condition node, and as the input of the relationship between sample nodes and test condition nodes in heterogeneous graph neural network; Step S15: For each polycarboxylate superplasticizer sample, collect its corresponding measured value of paste fluidity as the performance output label for that sample; wherein, the measured values ​​of paste fluidity of all samples are combined to form the model output label vector. Y Defined as:

[0018] in, For the first i One sample of polycarboxylate superplasticizer The measured value of the fluidity of the paste.

[0019] Furthermore, step S2 specifically includes: Step S21: Define the isomer diagram structure for predicting the performance of polycarboxylate superplasticizers. ,in, This represents the set of nodes in a heterogeneous graph. Represents the set of edges in a heterogeneous graph. Represents the node type mapping function, Represents the edge type mapping function; The set of nodes The node type mapping function is composed of a sample node set S, a single node set M, a process node set P, and a test condition node set T; the edge set E is composed of various relationship edges between sample nodes and other types of nodes; This is used to map each node in a heterogeneous graph to its corresponding node type; the edge type mapping function Used to map each edge in a heterogeneous graph to its corresponding edge type; Based on the above conditions, a heterogeneous graph structure containing multiple types of nodes and multiple types of relational edges is constructed to uniformly characterize the correlation between polycarboxylate superplasticizer samples, monomer composition, synthesis process and testing conditions; and to provide a graph structure foundation for subsequent heterogeneous graph input matrix construction and heterogeneous graph neural network model training.

[0020] Step S22: Based on the single-node feature matrix, process node feature matrix, test condition node feature matrix, and weighted correlation matrix between different types of nodes obtained in steps S12 to S14, construct the input representation of the polycarboxylate superplasticizer isomer diagram. The heterogeneous graph input representation includes two parts: node feature input and relation structure input; wherein, the node feature input includes the individual node feature matrix. Process node feature matrix and test condition node feature matrix The relational structure input includes a sample-unit weighted association matrix. Sample-process weighted correlation matrix Sample-Test Condition Weighted Association Matrix ; The node feature input and relation structure input are jointly represented as heterogeneous graph input data. Defined as:

[0021] This represents the input data set composed of node feature information and relational structure information, rather than a single matrix or a tensor of the same dimension. The feature dimensions of different types of nodes and the size of different relation matrices can be different, and feature mapping, relation propagation, and information aggregation can be performed separately in the subsequent heterogeneous graph neural network.

[0022] The heterogeneous graph input data Together with the output label vector Y defined in step S15, they constitute the training data for the heterogeneous graph neural network model; Step S23: Input data based on heterogeneous graph A heterogeneous graph neural network is used to propagate multi-relational information and aggregate features from sample nodes to obtain the representation vector of each polycarboxylate superplasticizer sample. The predicted value of the paste fluidity of the corresponding sample is then output through a fully connected layer. The predicted outputs of all samples are combined to form the model's predicted output vector. :

[0023] in, Represents the i-th sample node The predicted output, .

[0024] Furthermore, in step S3, a heterogeneous graph neural network is used as a learning framework to model the multi-type correlations between polycarboxylate superplasticizer samples, monomers, process parameters, and test conditions. The constructed heterogeneous graph neural network model mainly consists of the following structure: The model consists of a heterogeneous graph convolutional layer, a batch normalization layer, activation function layers following each convolutional layer, and a fully connected output layer. The activation function layer uses the ReLU function to introduce nonlinear expression and improve the model's ability to fit the complex relationship between the monomer composition, synthesis process, and test conditions of polycarboxylate superplasticizer and the fluidity of the paste. The fully connected output layer is used to output the predicted fluidity value of the paste corresponding to the polycarboxylate superplasticizer sample.

[0025] Furthermore, step S3 specifically includes: Step S31: Input the heterogeneous graph constructed in step S22 into the data. Input a heterogeneous graph neural network model, where the feature matrix of a single node is... Process node feature matrix and test condition node feature matrix These are used as the initial input features for the corresponding type of node; In this system, different types of nodes are set up with independent feature mapping layers to uniformly map the original node features of different dimensions to the same hidden feature space. Step S32: Employ a heterogeneous graph convolution module based on a multi-relationship propagation mechanism. By explicitly defining message passing functions between sample nodes and individual nodes, process nodes, and test condition nodes, efficient feature propagation and representation updates under multiple types of relationships are achieved; specifically as follows: In the l-th layer of heterogeneous graph convolution, the sample node obtains an updated sample node representation by aggregating the features of its connected individual nodes, process nodes, and test condition nodes, defined as:

[0026] in, , , and These represent the learnable parameter matrices of the correspondence type in the l-th layer, respectively. Represents a nonlinear activation function; Individual nodes, process nodes, and test condition nodes also update their features through the inverse relationship with sample nodes to achieve bidirectional information exchange between multiple types of nodes in the heterogeneous graph, as shown in the following expression:

[0027] in, , and This is the learnable parameter matrix corresponding to the reverse relation type. , and This is the learnable parameter matrix for the corresponding node type. This represents a node of type 'a' in the l-th layer; Step S33: After propagation through a multi-layer heterogeneous graph neural network, the high-order final representation vector of the sample nodes is obtained. ;in, This indicates that after L layers of heterogeneous graph convolution, the i-th sample node... The final representation vector, ; The final hidden layer representation dimension of the sample node; the final representation vector of the sample node. The monomer composition, synthesis process, and testing conditions of the sample were comprehensively characterized. Step S34: Convert the sample node representation vector The input prediction layer outputs the predicted fluidity value of the corresponding sample, expressed as: ,in, Represents the i-th sample node The predicted output, Represents the prediction function; The prediction function uses a fully connected layer to combine the prediction outputs of all samples to form the model's prediction output vector. The model predicts the output vector. This is used to characterize the prediction results of the heterogeneous graph neural network model on the flowability of the paste for each polycarboxylate superplasticizer sample, and is used for error calculation and parameter optimization in subsequent model training.

[0028] Furthermore, step S4 specifically includes: Step S41: Divide all polycarboxylate superplasticizer samples into training set, validation set and test set according to a preset ratio; wherein, the training set is used for model parameter learning, the validation set is used for hyperparameter adjustment and convergence monitoring during model training, and the test set is used to evaluate the model's predictive performance on unknown samples.

[0029] Step S42: Input the heterogeneous graph corresponding to the training set into the data. The heterogeneous graph neural network model constructed in step S3 is input, and multi-layer information propagation and feature aggregation are performed on sample nodes, individual nodes, process nodes, and test condition nodes through heterogeneous graph convolutional layers to obtain the final representation vector of the sample nodes. The final representation vector of the sample nodes is then output through a fully connected output layer to produce the predicted value of the paste fluidity of the corresponding sample, forming the model prediction output vector. The predicted output for the i-th sample in the training set. Used to correlate with the corresponding measured paste fluidity Perform error calculation; Step S43: The mean squared error is used as the loss function for the regression task to measure the difference between the prediction results and the actual results of the heterogeneous graph neural network model. The loss function is defined as follows:

[0030] in, This represents the model training loss. This represents the set of sample nodes in the training set; Step S44: Based on the loss function The gradient backpropagation algorithm is used to calculate the gradient of the loss function with respect to the model parameters, and the Adam optimizer is used to iteratively update the model parameters; wherein, the model parameters include node feature mapping parameters, relation type propagation parameters, self-representation parameters, and prediction layer parameters; Step S45: Model Convergence Determination and Performance Evaluation During iterative training, the loss on the validation set is monitored in real time. When the loss on the validation set reaches the minimum value, or when the loss on the validation set no longer decreases after several consecutive training iterations, the model training is determined to be complete and training is stopped, resulting in a trained heterogeneous graph neural network model. After training, the heterogeneous graph input data corresponding to the test set samples are input into the trained heterogeneous graph neural network model to predict the flowability of the paste of the test set samples. The prediction results are then compared with the measured values ​​in the test set to evaluate the predictive performance and generalization ability of the model. Among them, the coefficient of determination is used. Root mean square error and mean absolute error As a performance evaluation metric for the model.

[0031] Furthermore, step S5 specifically includes: Step S51: Obtain sample data of the polycarboxylate superplasticizer to be predicted. The sample data includes monomer composition information, synthesis process parameters and performance test conditions. The monomer composition information includes the types of unsaturated monomers, the types and molecular weights of polyether macromonomers, the types of functional monomers and their proportions; the synthesis process parameters include the dosages of initiators, chain transfer agents, reducing agents and oxidants, reaction temperature and reaction time; the performance test conditions include cement type, water-cement ratio and water-reducing agent dosage.

[0032] Step S52: Represent the sample to be predicted as a sample node in the heterogeneity graph, and combine it with the existing unit node, process node, and test condition node to construct the heterogeneity graph input data corresponding to the sample to be predicted. ; The heterogeneous graph input data includes a sample-to-unit weighted correlation matrix between the sample to be predicted and the unit node, a sample-to-process weighted correlation matrix between the sample to be predicted and the process node, and a sample-to-test condition weighted correlation matrix between the sample to be predicted and the test condition node. Among them, the heterogeneous graph input data of the sample to be predicted The expression is as follows:

[0033] in, This represents the weighted correlation matrix between the sample to be predicted and the individual node. This represents the weighted correlation matrix between the sample to be predicted and the process node. This represents the weighted correlation matrix between the sample to be predicted and the test condition nodes; , represents the set of samples to be predicted, and q represents the number of samples to be predicted; Step S53: Input the heterogeneity map of the samples to be predicted constructed in step S52 into the data. The heterogeneous graph neural network model trained in step S4 is input, and multi-layer information propagation and feature aggregation are performed through the heterogeneous graph convolutional layer to obtain the final representation vector of the node to be predicted. The predicted value of the paste fluidity of the sample to be predicted is then output through the fully connected output layer.

[0034] In a second aspect, the present invention also provides a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the polycarboxylate superplasticizer performance prediction method based on heterogeneous graph neural networks as described in the first aspect.

[0035] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: 1. This invention proposes a method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks. It can uniformly represent polycarboxylate superplasticizer samples, monomer composition, synthesis process parameters and test conditions as a heterogeneous graph structure, which overcomes the shortcomings of traditional tabular machine learning methods that are difficult to explicitly express multiple types of entities and their relationships, and improves the structured characterization ability of polycarboxylate superplasticizer data.

[0036] 2. By constructing sample nodes, monomer nodes, process nodes, and test condition nodes, and establishing multiple types of relationship edges such as sample-monomer, sample-process, and sample-test condition, this invention achieves joint modeling of the complex coupling relationship between "composition-process-test condition-performance" of polycarboxylate superplasticizer, which can more effectively explore the nonlinear correlation law between various factors.

[0037] 3. This invention uses a heterogeneous graph neural network to propagate and aggregate information between different types of nodes. Compared with traditional linear regression models, ordinary neural network models, or prediction methods based on single feature splicing, it can make fuller use of the interactive information in different node types and multi-relationship structures, thereby improving the accuracy and generalization ability of polycarboxylate superplasticizer pulp fluidity prediction.

[0038] 4. By unifying and organizing the monomer composition information, synthesis process parameters and test conditions of the water-reducing agent into a graphical structure, this invention can adapt to the performance prediction task of polycarboxylate water-reducing agent samples under different monomer systems, different process combinations and different test conditions, and has good method versatility and scalability.

[0039] 5. The method constructed in this invention can be directly used to predict the fluidity of the paste of polycarboxylate superplasticizer samples to be tested, providing data-driven technical support for the formulation screening, process optimization and performance evaluation of polycarboxylate superplasticizers, thereby reducing the dependence of traditional test methods on a large number of formulation trials and repeated tests, reducing R&D costs and improving R&D efficiency.

[0040] 6. The method of this invention has good engineering application prospects and can provide new technical means for intelligent design, rapid performance evaluation of polycarboxylate superplasticizers and digital R&D of concrete admixtures. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is an overall framework diagram of the polycarboxylate superplasticizer performance prediction method based on heterogeneous graph neural network provided by the present invention; Figure 2 A flowchart illustrating the steps of the method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks provided by this invention; Figure 3 This is a schematic diagram of the heterogeneous graph structure provided by the present invention; Figure 4 This is a schematic diagram of the structure of a computer-readable storage medium provided by the present invention. Detailed Implementation

[0043] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0044] See attached document Figure 1-2 As shown, this embodiment provides a method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks, including the following steps: Step 1: Data Acquisition and Preprocessing Data on polycarboxylate superplasticizer samples were collected, including monomer composition information, synthesis process parameters, performance testing conditions, and corresponding performance indicators. The sample data were derived from experimental results. The monomer composition information included the types of unsaturated monomers, polyether macromonomers, and functional monomers and their proportions. The synthesis process parameters included the dosages of initiators, chain transfer agents, reducing agents, and oxidants, as well as the reaction temperature and reaction time. The performance testing conditions included cement type, water-cement ratio, and superplasticizer dosage. The performance indicator was the fluidity of the cement paste.

[0045] Different types of raw data can be cleaned, normalized, discretely encoded, or numerically standardized. All data types are then uniformly numbered and structured to form sample node information, monomer node information, process node information, and test condition node information. After preprocessing, the water-reducing agent monomer composition information, synthesis process parameters, and test conditions are used as input data for the heterogeneous graph model, and the corresponding measured paste fluidity is used as the model output.

[0046] This embodiment uses three samples as example samples, denoted as Sample 1, Sample 2, and Sample 3. These are used to illustrate the data organization method, heterogeneity graph construction method, and performance prediction process of this scheme. In practical applications, a larger sample set is required to train and validate the model. The monomer composition information, synthesis process parameters, performance test conditions, and paste fluidity of the samples are as follows: Sample 1: The water-reducing agent monomers include unsaturated monomer acrylic acid (AA) and polyether macromonomer methyl allyl polyoxyethylene ether (HPEG), with a molecular weight of 2400 and an AA:HPEG molar ratio of 4:1. The synthesis process parameters used ammonium persulfate (APS) as the initiator, accounting for 3% of the total monomer mass, and mercaptoacetic acid (TGA) as the chain transfer agent, accounting for 1% of the total monomer mass. The reaction temperature was 65℃, and the reaction time was 5 hours. The performance test conditions used P.O42.5 cement, a water-cement ratio of 0.29, and a water-reducing agent dosage of 0.3%. The fluidity of the cement paste was 285 mm.

[0047] Sample 2: The water-reducing agent monomers include unsaturated monomer acrylic acid (AA), polyether macromonomer ethylene glycol monovinyl polyethylene glycol ether (EPEG) with a molecular weight of 3000, and functional monomer 2-hydroxyethyl methacrylate phosphate (HEMAP), with a molar ratio of EPEG:AA:HEMAP = 1:4:0.375; the synthesis process parameters used were: oxidant hydrogen peroxide (H2O2) accounting for 1.5% of the total monomer mass, reducing agent ascorbic acid (Vc) accounting for 0.3% of the total monomer mass, chain transfer agent TGA accounting for 0.5% of the total monomer mass, reaction temperature 20℃, and reaction time 3 hours; the performance test conditions used P.O42.5 cement, water-cement ratio 0.29, water-reducing agent dosage of 0.2%; and the fluidity of the cement paste was 290 mm.

[0048] Sample 3: The water-reducing agent monomers include unsaturated monomer acrylic acid (AA), polyether macromonomer isopentenyl alcohol polyoxyethylene ether (TPEG) with a molecular weight of 2400, and functional monomer sodium methyl allyl sulfonate (SMAS) with a molar ratio of TPEG:AA:SMAS = 1:4:0.3. The synthesis process parameters used ammonium persulfate (APS) as the initiator, accounting for 1% of the total monomer mass; ascorbic acid (Vc) as the reducing agent, accounting for 0.07% of the total monomer mass; mercaptoacetic acid (TGA) as the chain transfer agent, accounting for 0.6% of the total monomer mass; reaction temperature of 15℃; and reaction time of 4 hours. The performance test conditions used SH42.5 cement, water-cement ratio of 0.29, and water-reducing agent dosage of 0.4%. The fluidity of the cement paste was 251 mm.

[0049] In this embodiment, step 1 specifically includes the following steps: Step 11) Obtain the polycarboxylate superplasticizer sample dataset, which contains a total of N One water-reducing agent sample, denoted as: ,in, Indicates the first i One polycarboxylate superplasticizer sample. For each sample Collect information on the corresponding monomer composition, synthesis process parameters, test conditions, and corresponding paste fluidity values.

[0050] In this embodiment, take N =3, then the sample set is: .

[0051] Step 12) For each polycarboxylate superplasticizer sample, collect its monomer composition information, including the types of unsaturated monomers, the types and molecular weights of polyether macromonomers, the types of functional monomers and their proportions. All monomers constitute the monomer node set: ,in, Indicates the first j Individual node, This represents the total number of all individual nodes.

[0052] In this embodiment, the monomers present in all samples include AA, HPEG-2400, EPEG-3000, TPEG-2400, HEMAP, and SMAS. Therefore, a set of single nodes can be represented as:

[0053] For a single node Its node feature vector This is used to characterize the intrinsic properties of the monomer itself, including but not limited to monomer type, molecular weight, and functional group type (whether it contains a carboxyl group, whether it contains a polyether side chain, whether it contains a phosphate ester group, whether it contains a sulfonic acid group, etc.). d M This represents the feature dimension of a single node.

[0054] The feature matrix of a single node can be constructed from the feature vectors of all individual nodes. :

[0055] In this embodiment,

[0056] Wherein, the feature dimension of a single node is set to d M =8, where each dimension represents, in order, whether it is an unsaturated monomer, whether it is a polyether macromonomer, whether it is a functional monomer, monomer molecular weight, whether it contains a carboxyl group, whether it contains a polyether side chain, whether it contains a phosphate ester group, and whether it contains a sulfonic acid group. A monomer node feature matrix can be constructed from all the monomer node feature vectors. .

[0057] For any sample node and single node If the monomer Used for samples The synthesis is performed at the sample node. With a single node Establish a compositional relationship edge between them The edge features corresponding to the relation edges are denoted as: ,in, Indicates monomer In the sample The proportioning parameters in the formula include, but are not limited to, molar ratio, feed ratio, or mass ratio.

[0058] The compositional relationship between all sample nodes and individual nodes is represented as a sample-individual weighted correlation matrix. Defined as: , where matrix elements Defined as:

[0059] in, It is used to characterize the compositional relationship and connection strength between different polycarboxylate superplasticizer samples and each monomer node, and serves as the input for the relationship between sample nodes and monomer nodes in a heterogeneous graph neural network.

[0060] In this embodiment, the sample-unit weighted correlation matrix for: .

[0061] Step 13) For each polycarboxylate superplasticizer sample, collect the synthesis process parameters, including the dosage of initiator, chain transfer agent, reducing agent, and oxidant, as well as the reaction temperature and time. The dosages of initiator, chain transfer agent, reducing agent, and oxidant are all converted to percentages of the total mass of all monomers.

[0062] All process parameters constitute the process node set. ,in, Indicates the first j Each process node This indicates the total number of process nodes. Each process node can be further assigned a node type label to distinguish between initiator nodes, chain transfer agent nodes, reducing agent nodes, oxidizing agent nodes, temperature nodes, and time nodes.

[0063] In this embodiment, the process parameters appearing in all samples include APS, TGA, H2O2, Vc, 65 ℃, 20 ℃, 15 ℃, 5 h, 3 h, and 4 h. Therefore, the set of process nodes can be represented as: .

[0064] For process nodes Its node feature vector This is used to characterize the inherent properties of the process node itself. These inherent properties include, but are not limited to, process parameter type, reagent type, whether it belongs to a redox system, and parameter value type. This represents the feature dimension of the process node.

[0065] The feature matrix of all process nodes can be constructed from their feature vectors. :

[0066] In this embodiment, Each dimension represents, in order: whether it is an initiator node, whether it is a chain transfer agent node, whether it is a reducing agent node, whether it is an oxidizing agent node, whether it is a temperature node, and whether it is a time node. The feature vectors of all process nodes can form a process node feature matrix. :

[0067] For any polycarboxylate superplasticizer sample node and process nodes If process node Used for samples The synthesis process occurs at the sample node. With process nodes Establish a process relationship between them The edge features corresponding to the relation edge ,in, express In the sample The process parameters in the data, when the process node is an initiator, chain transfer agent, reducing agent, or oxidizing agent node, Indicates the corresponding reagent in the sample Its percentage of the total monomer mass; when the process node is a temperature node or a time node. This indicates the corresponding reaction temperature or reaction time value.

[0068] The technological relationships between all sample nodes and process nodes are represented as a sample-process weighted correlation matrix. Defined as: , where matrix elements Defined as:

[0069] The It is used to characterize the process relationship and connection strength between different polycarboxylate superplasticizer samples and each process node, and serves as the input for the relationship between sample nodes and process nodes in the heterogeneous graph neural network.

[0070] In this embodiment, the sample-process weighted correlation matrix for:

[0071] Step 14) Collect performance test parameters for each polycarboxylate superplasticizer sample, including cement type, water-cement ratio, and superplasticizer dosage.

[0072] All test condition parameters constitute the test condition node set. .in, Indicates the first j Test condition nodes This indicates the total number of test condition nodes. Each test condition node can be further assigned a node type label to distinguish between cement type nodes, water-cement ratio nodes, and water-reducing agent dosage nodes.

[0073] In this embodiment, the test conditions appearing in all samples include P.O42.5 cement, SH42.5 cement, water-cement ratio of 0.29, water-reducing agent dosage of 0.2%, water-reducing agent dosage of 0.3%, and water-reducing agent dosage of 0.4%. Therefore, the set of test condition nodes can be represented as follows:

[0074] For test condition nodes Its node feature vector This is used to characterize the inherent properties of the test condition node itself. These inherent properties include, but are not limited to, the type of condition parameter, the type of parameter value, the material type, and the test system type. This indicates the feature dimension of the test condition node.

[0075] The feature matrix of all test condition nodes can be constructed from their feature vectors. :

[0076] In this embodiment, Each feature dimension represents, in turn, whether it is a cement type node, a water-cement ratio node, and a content node. The feature matrix of all test condition nodes can be constructed from the feature vectors of all test condition nodes.

[0077] For any polycarboxylate superplasticizer sample node and test condition nodes If the test conditions Used for samples The synthesis process occurs at the sample node. With test condition nodes Establish a test condition relationship edge between them The edge features corresponding to the relation edge ,in, Indicates test condition node In the sample The corresponding test condition parameter values. When the test condition node is a cement type node. This indicates the corresponding category code value; when the test condition node is the water-cement ratio node or the water-reducing agent dosage node... This represents the corresponding numerical parameter.

[0078] The test condition relationships between all sample nodes and test condition nodes are represented as a sample-test condition weighted correlation matrix. Defined as: , where matrix elements Defined as:

[0079] The It is used to characterize the test condition relationship and connection strength between different polycarboxylate superplasticizer samples and each test condition node, and serves as the input for the relationship between sample nodes and test condition nodes in a heterogeneous graph neural network.

[0080] In this embodiment, the sample-test condition weighted correlation matrix for:

[0081] Step 15) For each polycarboxylate superplasticizer sample, collect the corresponding measured value of the paste fluidity as the performance output label for that sample. Let the first... i One sample of polycarboxylate superplasticizer The measured value of the fluidity of the paste is recorded as The measured values ​​of the fluidity of the paste from all samples are combined to form the model output label vector. Y Defined as:

[0082] In this embodiment, the flowability of the paste corresponding to the three samples is 285 mm, 290 mm, and 251 mm, respectively. Therefore, the output label vector is:

[0083] Step 2, Heterogeneous graph structure construction: Based on the polycarboxylate superplasticizer sample data obtained in step 1, a heterogeneous graph structure for predicting the fluidity of paste is constructed. The heterogeneous graph uses polycarboxylate superplasticizer sample nodes as core nodes and introduces monomer nodes, process nodes, and test condition nodes. Different types of relational edges describe the relationships between the sample and monomer composition, synthesis process, and test conditions, thus forming a graph structure data suitable for heterogeneous graph neural network modeling. (See attached...) Figure 3 As shown, the constructed heterogeneous graph includes four types of nodes: sample nodes, unit nodes, process nodes, and test condition nodes, as well as three types of relationship edges: sample-unit relationship edges, sample-process relationship edges, and sample-test condition relationship edges.

[0084] In this embodiment, sample-to-unit relationship edge sets are constructed based on the compositional relationship between sample nodes and unit nodes, the process relationship between sample nodes and process nodes, and the test condition relationship between sample nodes and test condition nodes. Sample-process relationship edge set And sample-test condition relationship edge set The three together constitute the edge set in the heterogeneous graph. Furthermore, different types of relation edges are assigned corresponding edge features and connection weights to construct a sample-individual weighted association matrix. Sample-process weighted correlation matrix and the sample-test condition weighted correlation matrix This is used to characterize the connection relationships and connection strength between different types of nodes; simultaneously, through the node type mapping function... and edge type mapping function Type identification is performed on nodes and edges to achieve a unified graph representation of multiple node types and relationships, thereby forming a heterogeneous graph data structure capable of expressing multi-dimensional relationships of "sample-unit-process-test conditions"; specifically, the following steps are included: Step 21) Definition of Heterogeneous Graph Define isomer diagram structure for predicting the performance of polycarboxylate superplasticizers ,in, This represents the set of nodes in a heterogeneous graph. Represents the set of edges in a heterogeneous graph. Represents the node type mapping function, This represents the edge type mapping function.

[0085] Node set From the set of sample nodes S Single node set M Process node set P and test condition node set T Together they constitute, that is: It is represented by the node feature matrix.

[0086] Edge set E It is composed of various relational edges between sample nodes and other types of nodes, namely: It is represented by a weighted correlation matrix.

[0087] Node type mapping function Used to map nodes in a heterogeneous graph to their corresponding node types, defined as: Among them, the set of node types Defined as: .

[0088] Edge type mapping function Used to map edges in a heterogeneous graph to their corresponding edge types, defined as: Among them, the set of edge types Defined as:

[0089] Based on the above, a heterogeneous graph structure containing multiple types of nodes and multiple types of relational edges is constructed to uniformly characterize the correlation between polycarboxylate superplasticizer samples, monomer composition, synthesis process and test conditions, providing a graph structure foundation for subsequent heterogeneous graph input matrix construction and heterogeneous graph neural network model training.

[0090] Step 22) Model Input Features Based on the monomer node feature matrix, process node feature matrix, test condition node feature matrix, and weighted correlation matrix between different types of nodes obtained in steps 12) to 14), an input representation of the polycarboxylate superplasticizer isomer diagram is constructed. The isomer diagram input representation includes two parts: node feature input and relational structure input.

[0091] Node feature input includes individual node feature matrix Process node feature matrix and test condition node feature matrix .

[0092] The relational structure input includes a sample-individual weighted association matrix. Sample-process weighted correlation matrix Sample-Test Condition Weighted Association Matrix .

[0093] The node feature input and relation structure input are jointly represented as heterogeneous graph input data. Defined as:

[0094] This represents the input data set composed of node feature information and relational structure information, rather than a single matrix or a tensor of the same dimension. The feature dimensions of different types of nodes and the size of different relation matrices can be different, and feature mapping, relation propagation, and information aggregation can be performed separately in the subsequent heterogeneous graph neural network.

[0095] Heterogeneous graph input data Compared with the output label vector defined in step 15) Y The training data together constitute the heterogeneous graph neural network model.

[0096] Step 23), Model Output (Performance Prediction Target) In this scheme, the model output is the predicted flowability of the paste for polycarboxylate superplasticizer samples. This is based on heterogeneous graph input data. A heterogeneous graph neural network is used to propagate multi-relational information and aggregate features from sample nodes to obtain the representation vector of each polycarboxylate superplasticizer sample. Furthermore, a prediction layer outputs the predicted value of the paste fluidity for the corresponding sample. Let the first... i Sample nodes The predicted output is denoted as ,but The predicted outputs of all samples are combined to form the model's predicted output vector. :

[0097] Step 3: Construction of Heterogeneous Graph Neural Network Model Based on the heterogeneous graph structure constructed in step 2, a heterogeneous graph neural network model containing multiple heterogeneous graph convolutional layers is constructed. This model learns the nonlinear mapping relationship between the composition, synthesis process, and test conditions of polycarboxylate superplasticizer and the fluidity of the paste by performing information propagation and feature aggregation on various types of relationships between sample nodes, monomer nodes, process nodes, and test condition nodes. This enables the prediction output of the fluidity of polycarboxylate superplasticizer paste.

[0098] In this embodiment, after completing node feature extraction and heterogeneous graph structure construction, a heterogeneous graph neural network is used as the learning framework to model the multi-type correlations between polycarboxylate superplasticizer samples, monomers, process parameters, and test conditions. The constructed heterogeneous graph neural network model mainly consists of the following structure: three heterogeneous graph convolutional layers, two batch normalization layers, an activation function layer following each convolutional layer, and a fully connected output layer. The activation function layer uses the ReLU function to introduce nonlinear expression and improve the model's ability to fit the complex relationships between polycarboxylate superplasticizer monomer composition, synthesis process, test conditions, and paste flowability; the fully connected output layer outputs the predicted paste flowability value corresponding to the polycarboxylate superplasticizer sample. Specifically, the following steps are included: Step 31) Node feature initialization and feature mapping Input the heterogeneous graph constructed in step 22) into the data. Input a heterogeneous graph neural network model, where the feature matrix of a single node is... Process node feature matrix and test condition node feature matrix These serve as the initial input features for their respective node types. Since individual nodes, process nodes, and test condition nodes have different data attributes and original feature dimensions, this scheme first sets up independent feature mapping layers for different node types, uniformly mapping the original node features of different dimensions to the same hidden feature space. l The type in the layer is a The node is represented as The initial node of layer 0 is denoted as:

[0099] in, This represents the initial hidden representation matrix for a single node; This represents the initial hidden representation matrix of the process node; This represents the initial hidden representation matrix of the test condition nodes; Learnable mapping parameter matrix; d For a unified hidden feature dimension.

[0100] In this scheme, the sample nodes serve as target prediction nodes in the heterogeneous graph, and their initial representations can be initialized using zero vector initialization, unit initialization, or learnable embedding initialization. The initial representation matrix of the sample nodes is: .

[0101] Step 32) Heterogeneous graph convolutional propagation and feature update To achieve information propagation and feature interaction among different types of nodes, a heterogeneous graph convolution module based on a multi-relationship propagation mechanism is adopted. This is achieved by explicitly defining message passing functions between sample nodes and individual nodes, process nodes, and test condition nodes, thus realizing efficient feature propagation and representation updates under multiple relationship types. In the... l In layer heterogeneous graph convolution, sample nodes obtain an updated sample node representation by aggregating the features of their connected individual nodes, process nodes, and test condition nodes, defined as:

[0102] in, , , and They represent the first l Learnable parameter matrices of correspondence type in layers, This represents a non-linear activation function.

[0103] To achieve bidirectional information exchange between multiple types of nodes in the heterogeneous graph, individual nodes, process nodes, and test condition nodes also update their features through inverse relationships with sample nodes, and are defined as follows:

[0104] in, , and This is the learnable parameter matrix corresponding to the reverse relation type. , and This is the learnable parameter matrix for the corresponding node type.

[0105] Step 33) Sample Node Representation Extraction After propagation through a multi-layer heterogeneous graph neural network, the high-order representation vectors of the sample nodes are obtained. Let's assume that after three layers of heterogeneous graph convolution, the... i Sample nodes The final representation vector is denoted as ,in, .in, This represents the dimension of the final hidden layer representation of the sample node. The final representation vector of the sample node. The monomer composition, synthesis process, and testing conditions of the sample were comprehensively characterized.

[0106] Step 34) Prediction layer construction The final representation vector of the sample node Input the prediction layer, output the predicted fluidity value of the corresponding sample. i Sample nodes The predicted output is denoted as , ;in, This represents the prediction function, which can be achieved by using fully connected layers to combine the prediction outputs of all samples to form the model's prediction output vector. Model predicted output vector This is used to characterize the prediction results of the heterogeneous graph neural network model on the flowability of the paste for each polycarboxylate superplasticizer sample, and is used for error calculation and parameter optimization in subsequent model training.

[0107] Step 4: Model Training After constructing the heterogeneous graph neural network model in step 3, supervised training of the model is performed based on the polycarboxylate superplasticizer sample data obtained in step 1 and the heterogeneous graph input data constructed in step 2. By calculating the error between the model's predicted output and the measured paste fluidity label of the sample, gradient backpropagation and parameter iterative optimization are used to update the model parameters, thereby establishing a mapping relationship between the polycarboxylate superplasticizer sample composition, synthesis process, test conditions, and paste fluidity. After training, the model can be used to predict the paste fluidity of any input polycarboxylate superplasticizer material sample; specifically, the following steps are included: Step 41) Dataset partitioning All polycarboxylate superplasticizer samples were divided into training, validation, and test sets according to a preset ratio. The training set was used for model parameter learning, the validation set was used for hyperparameter tuning and convergence monitoring during model training, and the test set was used to evaluate the model's predictive performance on unknown samples.

[0108] In this embodiment, the training set, validation set, and test set can be divided in a ratio of 8:1:1.

[0109] Step 42) Input the heterogeneous graph corresponding to the training set into the data. The heterogeneous graph neural network model constructed in step 3 is input, and multi-layer information propagation and feature aggregation are performed on sample nodes, individual nodes, process nodes, and test condition nodes through heterogeneous graph convolutional layers to obtain the final representation vector of the sample nodes. The final representation vector of the sample nodes is then output through a fully connected output layer to produce the predicted value of the paste fluidity of the corresponding sample, forming the model prediction output vector. .

[0110] This represents the model's predicted output vector for all samples, specifically the vector representing the nth sample in the training set. i Predicted output for each sample Used to correlate with the corresponding measured paste fluidity Error calculation is performed. During training, the training loss is calculated only for the predicted output and the true label corresponding to the training set samples; the validation set and test set samples do not participate in parameter updates and are only used for performance evaluation.

[0111] Step 43) Loss Function Construction To measure the difference between the predictions and experimental results of the heterogeneous graph neural network model, mean squared error is used as the loss function for the regression task. The loss function L is defined as:

[0112] in, This represents the model training loss. This represents the set of sample nodes in the training set.

[0113] Step 44) Based on the loss function The gradient backpropagation algorithm is used to calculate the gradient of the loss function with respect to the model parameters, and the Adam optimizer is used to iteratively update the model parameters; the parameter update can be expressed as:

[0114] in, For the first t Model parameters at the next iteration The learning rate is used; model parameters include node feature mapping parameters, relation type propagation parameters, self-representation parameters, and prediction layer parameters.

[0115] Step 45) Model convergence determination and performance evaluation During iterative training, the loss on the validation set is monitored in real time. When the validation set loss reaches its minimum value, or when the validation set loss no longer decreases after several consecutive training iterations, the model training is considered complete and training is stopped, resulting in a trained heterogeneous graph neural network model.

[0116] After training, the heterogeneous graph input data corresponding to the test set samples is input into the trained heterogeneous graph neural network model to predict the flowability of the paste sample in the test set. The prediction results are then compared with the measured values ​​in the test set to evaluate the model's prediction performance and generalization ability.

[0117] A trained heterogeneous graph neural network model was used to predict the fluidity of paste on a test set. The predicted results were compared with measured values ​​to evaluate the model's predictive performance. The coefficient of determination was used as the criterion. Root mean square error and mean absolute error As a performance evaluation metric for the model.

[0118]

[0119] in, This represents the average measured fluidity of the sample paste. n This indicates the number of samples participating in the evaluation.

[0120] Step 5, Performance Prediction: After training the heterogeneous graph neural network model, the polycarboxylate superplasticizer sample to be predicted is input into the trained heterogeneous graph neural network model to predict the flowability of the paste of the sample and obtain the corresponding performance prediction results; specifically, the following steps are included: Step 51) Obtain sample data of the polycarboxylate superplasticizer to be predicted. The sample data includes monomer composition information, synthesis process parameters, and test conditions. The monomer composition information includes the types of unsaturated monomers, the types and molecular weights of polyether macromonomers, the types of functional monomers, and their proportions. The synthesis process parameters include the dosages of initiator, chain transfer agent, reducing agent, and oxidant, as well as the reaction temperature and reaction time. The test conditions include cement type, water-cement ratio, and superplasticizer dosage.

[0121] Step 52) Represent the sample to be predicted as a sample node in the heterogeneous graph, and combine it with the existing unit node, process node, and test condition node to construct the heterogeneous graph input data corresponding to the sample to be predicted. The heterogeneous graph input data includes the sample-to-unit relationship matrix between the sample to be predicted and the unit node, the sample-to-process relationship matrix between the sample to be predicted and the process node, and the sample-to-test condition relationship matrix between the sample to be predicted and the test condition node.

[0122] The set of samples to be predicted is denoted as Where q represents the number of samples to be predicted, the heterogeneous graph input data of the samples to be predicted is denoted as:

[0123] in, This represents the relationship matrix between the sample to be predicted and the individual node. This represents the matrix showing the relationship between the sample to be predicted and the process nodes. This represents the relationship matrix between the sample to be predicted and the test condition nodes.

[0124] Step 53) Model Inference and Performance Output Input the heterogeneity map of the samples to be predicted constructed in step 52) into the data. The heterogeneous graph neural network model trained in step 4 is input, and multi-layer information propagation and feature aggregation are performed through the heterogeneous graph convolutional layer to obtain the final representation vector of the node to be predicted. The predicted value of the paste fluidity of the sample to be predicted is then output through the fully connected output layer.

[0125] The predicted value of the fluidity of the paste for the sample to be predicted can be expressed as: , Indicates the first j Predicted values ​​of the fluidity of the paste for each sample to be predicted. Indicates that the sample node to be predicted has passed through the first... L The final representation vector after convolution of the heterogeneous graph. This represents a fully connected prediction function.

[0126] Output the predicted value of the paste fluidity corresponding to the polycarboxylate superplasticizer sample to be predicted. This value serves as the performance prediction result of the sample under given monomer composition, synthesis process parameters, and test conditions, providing a basis for the formulation screening, process optimization, and performance evaluation of polycarboxylate superplasticizers.

[0127] As attached Figure 4 As shown, the present invention also provides a computer-readable storage medium having stored thereon computer program instructions for performing the above-described method steps.

[0128] The computer-readable storage medium can be a non-volatile storage device, such as a hard disk, solid-state drive, USB flash drive, read-only memory (ROM), flash memory, CD-ROM, or other media capable of storing information electronically. The computer program stored on this medium can be read and executed by a processor to implement the aforementioned method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks.

[0129] The program instructions are used to implement the various method steps described in this specification, including but not limited to: data acquisition and preprocessing, heterogeneous graph structure construction, heterogeneous graph neural network model construction and training, and performance prediction. The computing device executing this program can be a server, cloud platform, personal computer, or embedded system, etc.

[0130] By deploying the program in the computer-readable storage medium, the present invention can automate, intelligentize, and optimize the material design process, improve the efficiency and accuracy of the development of ultra-high performance concrete materials, and has good practical value and promotion prospects.

[0131] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks, characterized in that, Includes the following steps: S1. Collect sample data of polycarboxylate superplasticizer, including superplasticizer monomer composition information, synthesis process parameters, performance test conditions and corresponding performance indicators, and preprocess different types of raw data respectively. After preprocessing, the monomer composition information of the water-reducing agent, the synthesis process parameters and test conditions are used as input data for the isomer diagram model, and the corresponding performance indicators are used as the model output. S2. Based on the polycarboxylate superplasticizer sample data obtained in step S1, a heterogeneous graph structure for predicting the fluidity of the paste is constructed. The heterogeneous graph uses polycarboxylate superplasticizer sample nodes as core nodes and introduces monomer nodes, process nodes, and test condition nodes. Different types of relational edges describe the correlation between the sample and monomer composition, synthesis process, and test conditions, thereby forming a heterogeneous graph structure data suitable for heterogeneous graph neural network modeling. The constructed heterogeneous graph includes four types of nodes: sample nodes, monomer nodes, process nodes, and test condition nodes, as well as three types of relational edges: sample-monomer relational edges, sample-process relational edges, and sample-test condition relational edges. S3. Based on the constructed heterogeneous graph data structure, a heterogeneous graph neural network model containing multiple heterogeneous graph convolutional layers is constructed. The heterogeneous graph neural network model learns the nonlinear mapping relationship between the composition, synthesis process and test conditions of polycarboxylate superplasticizer and the fluidity of the paste by performing information propagation and feature aggregation on the multi-type relationships between sample nodes, monomer nodes, process nodes and test condition nodes, thereby realizing the predictive output of the performance indicators of polycarboxylate superplasticizer. S4. Based on the constructed heterogeneous graph neural network model, the heterogeneous graph neural network model is trained under supervision using the polycarboxylate superplasticizer sample data obtained in step S1 and the heterogeneous graph structure data constructed in step S2. By calculating the error between the model prediction output and the measured flowability label of the sample paste, the model parameters are updated using gradient backpropagation and parameter iterative optimization, thereby establishing the mapping relationship between the polycarboxylate superplasticizer sample composition, synthesis process, test conditions and paste flowability. S5. After completing the training of the heterogeneous graph neural network model, input the polycarboxylate superplasticizer sample to be predicted into the trained heterogeneous graph neural network model to predict the flowability of the paste of the sample to be predicted, and obtain the corresponding performance prediction results.

2. The method for predicting the performance of polycarboxylate superplasticizer based on heterogeneous graph neural networks according to claim 1, characterized in that, In step S1, the monomer composition information includes the types of unsaturated monomers, the types and molecular weights of polyether macromonomers, the types of functional monomers and their proportions; the synthesis process parameters include the dosage of initiator, chain transfer agent, reducing agent and oxidant, reaction temperature and reaction time; the performance test conditions include cement type, water-cement ratio and water-reducing agent dosage; and the performance index is the fluidity of the cement paste.

3. The method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks according to claim 2, characterized in that, In step S1, different types of raw data are preprocessed separately. Specifically, different types of raw data are cleaned, normalized, discretely encoded, or numerically standardized. All types of data are uniformly numbered and structured to form sample node information, individual node information, process node information, and test condition node information.

4. The method for predicting the performance of polycarboxylate superplasticizer based on heterogeneous graph neural networks according to claim 3, characterized in that, Step S1 specifically includes: Step S11: Obtain the polycarboxylate superplasticizer sample dataset, assuming a total of N One water-reducing agent sample, denoted as: ,in, Indicates the first i One polycarboxylate superplasticizer sample, for each sample Collect information on the corresponding monomer composition, synthesis process parameters, test conditions, and corresponding paste fluidity values. Step S12: For each polycarboxylate superplasticizer sample, collect its monomer composition information, including the types of unsaturated monomers, the types and molecular weights of polyether macromonomers, the types of functional monomers and their ratio parameters. All monomers constitute a monomer node set. ,in, Indicates the first j Individual node, This represents the total number of all individual nodes; For a single node Its node feature vector This is used to characterize the intrinsic properties of the monomer itself, including monomer type, molecular weight, and functional group type; wherein, d M Indicates the feature dimension of a single node; The feature matrix of a single node can be constructed from the feature vectors of all individual nodes. : For any sample node and single node If the monomer Used for samples The synthesis is performed at the sample node. With a single node Establish a compositional relationship edge between them The edge features corresponding to the relation edges are denoted as: ,in, Indicates monomer In the sample The proportioning parameters in the formula include, but are not limited to, molar ratio, feed ratio, and mass ratio; The compositional relationship between all sample nodes and individual nodes is represented as a sample-individual weighted correlation matrix. Defined as: , where matrix elements Defined as: The Used to characterize the compositional relationship and connection strength between different polycarboxylate superplasticizer samples and each monomer node, and as input to the relationship matrix between sample nodes and monomer nodes in heterogeneous graph neural network; Step S13: For each polycarboxylate superplasticizer sample, collect the synthesis process parameters, including the dosage of initiator, chain transfer agent, reducing agent and oxidant, reaction temperature and time; wherein, the dosage of initiator, chain transfer agent, reducing agent and oxidant are all converted into percentages of the total mass of all monomers; All process parameters constitute the process node set. ,in, Indicates the first j Each process node Indicates the total number of process nodes; For process nodes Its node feature vector This is used to characterize the inherent properties of the process node itself, including process parameter category, reagent category, whether it belongs to a redox system, and parameter value type; wherein, Represents the feature dimension of the process node; The feature matrix of all process nodes can be constructed from their feature vectors. : For any polycarboxylate superplasticizer sample node and process nodes If process node Used for samples The synthesis process occurs at the sample node. With process nodes Establish a process relationship between them The edge features corresponding to the relation edge ,in, express In the sample The process parameters in the data, when the process node is an initiator, chain transfer agent, reducing agent, or oxidizing agent node, Indicates the corresponding reagent in the sample Its percentage of the total monomer mass; when the process node is a temperature node or a time node. This indicates the corresponding reaction temperature or reaction time value; The technological relationships between all sample nodes and process nodes are represented as a sample-process weighted correlation matrix. Defined as: ; where matrix elements Defined as: The Used to characterize the process relationship and connection strength between different polycarboxylate superplasticizer samples and each process node, and as input to the relationship matrix between sample nodes and process nodes in heterogeneous graph neural network; Step S14: Collect performance test parameters for each polycarboxylate superplasticizer sample, including cement type, water-cement ratio, and superplasticizer dosage. All test condition parameters constitute the test condition node set. ;in, Indicates the first j Test condition nodes Indicates the total number of test condition nodes; For test condition nodes Its node feature vector This is used to characterize the inherent properties of the test condition node itself, including the condition parameter category, parameter value type, material category, and test system category; wherein, Indicates the feature dimension of the test condition node; The feature matrix of all test condition nodes can be constructed from their feature vectors. : For any polycarboxylate superplasticizer sample node and test condition nodes If the test conditions Used for samples The synthesis process occurs at the sample node. With test condition nodes Establish a test condition relationship edge between them The edge features corresponding to the relation edge ,in, Indicates test condition node In the sample The corresponding test condition parameter values; when the test condition node is a cement type node. This indicates the corresponding category code value; when the test condition node is the water-cement ratio node or the water-reducing agent dosage node... This represents the corresponding numerical parameter; The test condition relationships between all sample nodes and test condition nodes are represented as a sample-test condition weighted correlation matrix. Defined as: , where matrix elements Defined as: The Used to characterize the test condition relationship and connection strength between different polycarboxylate superplasticizer samples and each test condition node, and as the input of the relationship between sample nodes and test condition nodes in heterogeneous graph neural network; Step S15: For each polycarboxylate superplasticizer sample, collect its corresponding measured value of paste fluidity as the performance output label for that sample; wherein, the measured values ​​of paste fluidity of all samples are combined to form the model output label vector. Y Defined as: in, For the first i One sample of polycarboxylate superplasticizer The measured value of the fluidity of the paste.

5. The method for predicting the performance of polycarboxylate superplasticizer based on heterogeneous graph neural networks according to claim 4, characterized in that, Step S2 specifically includes: Step S21: Define the isomer diagram structure for predicting the performance of polycarboxylate superplasticizers. ,in, This represents the set of nodes in a heterogeneous graph. Represents the set of edges in a heterogeneous graph. Represents the node type mapping function, Represents the edge type mapping function; The set of nodes The node type mapping function is composed of a sample node set S, a single node set M, a process node set P, and a test condition node set T; the edge set E is composed of various relationship edges between sample nodes and other types of nodes; This is used to map each node in a heterogeneous graph to its corresponding node type; the edge type mapping function Used to map each edge in a heterogeneous graph to its corresponding edge type; Based on the above conditions, a heterogeneous graph structure containing multiple types of nodes and multiple types of relation edges is constructed to uniformly characterize the relationship between polycarboxylate superplasticizer samples, monomer composition, synthesis process and test conditions. Step S22: Based on the single-node feature matrix, process node feature matrix, test condition node feature matrix, and weighted correlation matrix between different types of nodes obtained in steps S12 to S14, construct the input representation of the polycarboxylate superplasticizer isomer diagram. The heterogeneous graph input representation includes two parts: node feature input and relation structure input; wherein, the node feature input includes the individual node feature matrix. Process node feature matrix and test condition node feature matrix The relational structure input includes a sample-unit weighted association matrix. Sample-process weighted correlation matrix Sample-Test Condition Weighted Association Matrix ; The node feature input and relation structure input are jointly represented as heterogeneous graph input data. Defined as: This represents the input data set composed of node feature information and relational structure information. The heterogeneous graph input data Together with the output label vector Y defined in step S15, they constitute the training data for the heterogeneous graph neural network model; Step S23: Input data based on heterogeneous graph A heterogeneous graph neural network is used to propagate multi-relational information and aggregate features from sample nodes to obtain the representation vector of each polycarboxylate superplasticizer sample. The predicted value of the paste fluidity of the corresponding sample is then output through a fully connected layer. The predicted outputs of all samples are combined to form the model's predicted output vector. : in, Represents the i-th sample node The predicted output, .

6. The method for predicting the performance of polycarboxylate superplasticizer based on heterogeneous graph neural networks according to claim 1, characterized in that, In step S3, a heterogeneous graph neural network is used as a learning framework to model the multi-type correlations between polycarboxylate superplasticizer samples, monomers, process parameters, and test conditions. The constructed heterogeneous graph neural network model mainly consists of the following structure: The model consists of a heterogeneous graph convolutional layer, a batch normalization layer, activation function layers following each convolutional layer, and a fully connected output layer. The activation function layer uses the ReLU function to introduce nonlinear expression and improve the model's ability to fit the complex relationship between the monomer composition, synthesis process, and test conditions of polycarboxylate superplasticizer and the fluidity of the paste. The fully connected output layer is used to output the predicted fluidity value of the paste corresponding to the polycarboxylate superplasticizer sample.

7. The method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks according to claim 5, characterized in that, Step S3 specifically includes: Step S31: Input the heterogeneous graph constructed in step S22 into the data. Input a heterogeneous graph neural network model, where the feature matrix of a single node is... Process node feature matrix and test condition node feature matrix These are used as the initial input features for the corresponding type of node; In this system, different types of nodes are set up with independent feature mapping layers to uniformly map the original node features of different dimensions to the same hidden feature space. Step S32: Employ a heterogeneous graph convolution module based on a multi-relationship propagation mechanism. By explicitly defining message passing functions between sample nodes and individual nodes, process nodes, and test condition nodes, efficient feature propagation and representation updates under multiple types of relationships are achieved; specifically as follows: In the l-th layer of heterogeneous graph convolution, the sample node obtains an updated sample node representation by aggregating the features of its connected individual nodes, process nodes, and test condition nodes, defined as: in, , , and These represent the learnable parameter matrices of the correspondence type in the l-th layer, respectively. Represents a nonlinear activation function; Individual nodes, process nodes, and test condition nodes also update their features through the inverse relationship with sample nodes to achieve bidirectional information exchange between multiple types of nodes in the heterogeneous graph, as shown in the following expression: in, , and This is the learnable parameter matrix corresponding to the reverse relation type. , and This is the learnable parameter matrix for the corresponding node type. This represents a node of type 'a' in the l-th layer; Step S33: After propagation through a multi-layer heterogeneous graph neural network, the high-order final representation vector of the sample nodes is obtained. ;in, This indicates that after L layers of heterogeneous graph convolution, the i-th sample node... The final representation vector, ; The final hidden layer representation dimension of the sample node; the final representation vector of the sample node. The monomer composition, synthesis process, and testing conditions of the sample were comprehensively characterized. Step S34: Convert the sample node representation vector The input prediction layer outputs the predicted fluidity value of the corresponding sample, expressed as: ,in, Represents the i-th sample node The predicted output, Represents the prediction function; The prediction function uses a fully connected layer to combine the prediction outputs of all samples to form the model's prediction output vector. The model predicts the output vector. This is used to characterize the prediction results of the heterogeneous graph neural network model on the flowability of the paste for each polycarboxylate superplasticizer sample, and is used for error calculation and parameter optimization in subsequent model training.

8. The method for predicting the performance of polycarboxylate superplasticizer based on heterogeneous graph neural networks according to claim 7, characterized in that, Step S4 specifically includes: Step S41: Divide all polycarboxylate superplasticizer samples into training set, validation set and test set according to a preset ratio; Step S42: Input the heterogeneous graph corresponding to the training set into the data. The heterogeneous graph neural network model constructed in step S3 is input, and multi-layer information propagation and feature aggregation are performed on sample nodes, individual nodes, process nodes, and test condition nodes through heterogeneous graph convolutional layers to obtain the final representation vector of the sample nodes. The final representation vector of the sample nodes is then output through a fully connected output layer to produce the predicted value of the paste fluidity of the corresponding sample, forming the model prediction output vector. The predicted output for the i-th sample in the training set. Used to correlate with the corresponding measured paste fluidity Perform error calculation; Step S43: The mean squared error is used as the loss function for the regression task to measure the difference between the prediction results and the actual results of the heterogeneous graph neural network model. The loss function is defined as follows: in, This represents the model training loss. This represents the set of sample nodes in the training set; Step S44: Based on the loss function The gradient backpropagation algorithm is used to calculate the gradient of the loss function with respect to the model parameters, and the Adam optimizer is used to iteratively update the model parameters; wherein, the model parameters include node feature mapping parameters, relation type propagation parameters, self-representation parameters, and prediction layer parameters; Step S45: Model Convergence Determination and Performance Evaluation During iterative training, the loss on the validation set is monitored in real time. When the loss on the validation set reaches the minimum value, or when the loss on the validation set no longer decreases after several consecutive training iterations, the model training is determined to be complete and training is stopped, resulting in a trained heterogeneous graph neural network model. After training, the heterogeneous graph input data corresponding to the test set samples are input into the trained heterogeneous graph neural network model to predict the flowability of the paste of the test set samples. The prediction results are then compared with the measured values ​​in the test set to evaluate the predictive performance and generalization ability of the model. Among them, the coefficient of determination is used. Root mean square error and mean absolute error As a performance evaluation metric for the model.

9. The method for predicting the performance of polycarboxylate superplasticizers based on heterogeneous graph neural networks according to claim 2, characterized in that, Step S5 specifically includes: Step S51: Obtain sample data of the polycarboxylate superplasticizer to be predicted. The sample data includes monomer composition information, synthesis process parameters and performance test conditions. The monomer composition information includes the types of unsaturated monomers, the types and molecular weights of polyether macromonomers, the types of functional monomers and their proportions; the synthesis process parameters include the dosages of initiators, chain transfer agents, reducing agents and oxidants, reaction temperature and reaction time; the performance testing conditions include cement type, water-cement ratio and water-reducing agent dosage. Step S52: Represent the sample to be predicted as a sample node in the heterogeneous graph, and combine it with the existing unit node, process node, and test condition node to construct the heterogeneous graph input data corresponding to the sample to be predicted. ; The heterogeneous graph input data includes a sample-to-unit weighted correlation matrix between the sample to be predicted and the unit node, a sample-to-process weighted correlation matrix between the sample to be predicted and the process node, and a sample-to-test condition weighted correlation matrix between the sample to be predicted and the test condition node. Among them, the heterogeneous graph input data of the sample to be predicted The expression is as follows: in, This represents the weighted correlation matrix between the sample to be predicted and the individual node. This represents the weighted correlation matrix between the sample to be predicted and the process node. This represents the weighted correlation matrix between the sample to be predicted and the test condition nodes; , represents the set of samples to be predicted, and q represents the number of samples to be predicted; Step S53: Input the heterogeneity map of the samples to be predicted constructed in step S52 into the data. The heterogeneous graph neural network model trained in step S4 is input, and multi-layer information propagation and feature aggregation are performed through the heterogeneous graph convolutional layer to obtain the final representation vector of the node to be predicted. The predicted value of the paste fluidity of the sample to be predicted is then output through the fully connected output layer.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the polycarboxylate superplasticizer performance prediction method based on heterogeneous graph neural networks as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Wireless communication network knowledge graph representation learning method based on heterogeneous graph neural network

    CN117196033A

  • Binder composition-performance coupling prediction method based on graph neural network

    CN120766810A