Method for predicting performance of repair materials and optimizing formulations based on machine learning algorithms

By constructing a structured multimodal feature library and a set of heteroproton models, and combining Bayesian optimization and a differentiable allocation formula simulator, the problems of data and physical processes being disconnected and low optimization efficiency in the prediction of repair material properties and formulation optimization in existing technologies have been solved, realizing an efficient and intelligent material research and development process.

CN121922286BActive Publication Date: 2026-05-22FUZHOU UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FUZHOU UNIV
Filing Date
2026-03-18
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing technologies for predicting the performance of repair materials and optimizing formulations suffer from crude data processing methods. They fail to integrate microscopic interface features, process dynamic response time-series data, and dynamic service performance that have a decisive impact on performance. This results in a disconnect between data and real physical processes, weak prediction foundations, and the inability of a single model to incorporate the physicochemical mechanisms of materials, leading to poor adaptability, low optimization efficiency, and a lack of closed-loop learning mechanisms.

Method used

A structured multimodal feature library is constructed, and a heterogeneous model set of graph neural networks, temporal convolutional networks and gradient boosting trees is adopted. Data fusion is performed in combination with a meta-learner, a closed-loop optimization process is established, gradient iterative optimization is performed using Bayesian optimization and a differentiable allocation square simulator, and the model is updated through local incremental learning.

Benefits of technology

It achieves more accurate performance prediction, has good physical interpretability, shortens the R&D cycle, reduces costs, realizes the transformation from blind trial and error to targeted intelligent design, and adapts to new material systems through self-improvement, thus realizing the sustainable intelligentization of the R&D process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121922286B_ABST
    Figure CN121922286B_ABST
Patent Text Reader

Abstract

The application discloses a method for predicting performance of repair materials and optimizing formula based on a machine learning algorithm, and particularly relates to the field of cross between artificial intelligence and material science, comprising the following steps: constructing a structured multi-modal feature library, adopting double indexing and differential noise filtering for data processing, constructing a heterogeneous model set guided by physical mechanism, including three targeted sub-models of a graph neural network, a time series convolution network and a gradient boosting tree, dynamically fusing outputs of the sub-models through a meta-learner to realize joint prediction of material performance, actively searching for potential formula in a high-dimensional solution space through dimension reduction mapping and Bayesian optimization, and generating an optimal formula scheme by using a differentiable physical-data fusion simulator for gradient-guided constraint optimization, and finally realizing automatic feedback of new data and self-evolution of the model through triggering rules and local incremental learning mechanisms based on uncertainty and deviation, so as to form a complete self-evolution closed loop system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and materials science, and more specifically, to a method for predicting the performance of repair materials and optimizing their formulations based on machine learning algorithms. Background Technology

[0002] The performance and formulation design of repair materials are key factors that limit their engineering application effectiveness. Traditional research and development models mainly rely on experimental trial and error and expert experience, which have problems such as long research and development cycles, high costs and difficulty in systematically exploring complex formulation spaces. With the development of the material genome concept and data science, the use of computational models to assist in material research and development has become an important trend.

[0003] Existing solutions attempt to apply a single machine learning model, such as random forest or traditional neural network, to predict individual performance indicators of repair materials based on a limited dataset of static components and process parameters. The implementation method is usually as follows: collect historical experimental data, extract the proportion of formulation components and basic process parameters as features, train a general prediction model, and filter within a preset formulation range through traversal or simple optimization algorithms to obtain formulation suggestions that meet specific performance requirements.

[0004] However, in practical use, it still has some shortcomings. For example, the data processing method is crude and fails to integrate the microscopic interface features, process dynamic response time series data and dynamic service performance that have a decisive impact on performance, resulting in the data being disconnected from the real physical process and the prediction basis being weak. Secondly, it adopts a single model and cannot integrate the physicochemical mechanisms in the field of materials, resulting in poor adaptability to different types of data features and low interpretability and extrapolation prediction reliability of the model. Finally, the optimization process is disconnected from the prediction model, and it mostly adopts a passive search strategy, which is inefficient in the complex high-dimensional space of components and processes, and lacks a closed-loop learning mechanism that continuously improves itself using new experimental data, making it difficult to achieve intelligent and adaptive optimization of formulations. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a method for predicting the performance of repair materials and optimizing their formulations based on machine learning algorithms, thereby addressing the problems raised in the background art through the following solutions.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for performance prediction and formulation optimization of repair materials based on machine learning algorithms, comprising S1: collecting and associating component space, process, and multidimensional performance data of the repair materials; wherein, the component space data includes chemical identifiers, mass fractions, and physicochemical property vectors of the matrix, fillers, modifiers, and additives; the process data includes time-series control parameters and environmental parameter vectors of the mixing, curing, or sintering processes; the multidimensional performance data includes quantitative results of mechanical, rheological, durability, and interfacial strength; and aligning and integrating the above data using the material formulation as an index to form a structured multimodal feature library;

[0007] S2: Establish a set of heterogeneous models guided by physical mechanisms. The set includes: a first sub-model based on graph neural network to process component space data to learn component interaction relationships; a second sub-model based on temporal convolutional network to process process data to capture process-structure correlation; and a third sub-model based on gradient boosting tree to process physical properties to identify key influencing factors. A meta-learner is used to dynamically fuse the outputs of each sub-model according to the prediction confidence in different feature spaces to generate a joint prediction of multidimensional performance.

[0008] S3: Based on the trained prediction model, a closed-loop optimization process is constructed. Bayesian optimization is used to actively search for potential formulation regions in the high-dimensional solution space composed of components and processes. The prediction model is embedded into a differentiable formula simulator to calculate the gradient of the target performance with respect to the formulation parameters. Combined with the constrained optimization algorithm, the formulation is iteratively adjusted under process and cost constraints to generate an optimized solution.

[0009] S4: Automatically feed the new formula-performance data pairs obtained after the optimized formula preparation and verification to the feature library, set trigger rules based on prediction uncertainty and bias, and automatically trigger the local incremental learning of the prediction model when the new data is in the high uncertainty region or exceeds the bias threshold.

[0010] The technical effects and advantages of this invention are as follows:

[0011] 1. This invention provides a more comprehensive and physically relevant data foundation for model learning by constructing a structured multimodal feature library that integrates "component micro-interface features", "process dynamic response" and "multi-dimensional service performance". It adopts a heterogeneous integrated model guided by physical mechanisms, and uses graph neural networks, temporal convolutional networks and gradient boosting trees to accurately model component interactions, process-structure correlations and key factors respectively. Through dynamic fusion by meta-learners, the performance prediction is not only highly accurate, but also has good physical interpretability, breaking through the limitations of traditional "black box" models.

[0012] 2. This invention addresses the high-dimensional and complex formulation-process solution space by combining variational autoencoder dimensionality reduction mapping, Bayesian optimization incorporating prior knowledge, and gradient-guided differentiable simulator. It proactively and efficiently locks potential formulation regions and performs precise gradient iterative optimization, rapidly converging to the optimal solution under multiple constraints. This achieves a shift from "blind trial and error" to "directed intelligent design," significantly shortening the R&D cycle and reducing experimental costs.

[0013] 3. This invention establishes a complete self-evolving closed loop of "prediction-optimization-verification-learning". Based on the intelligent triggering rules of prediction uncertainty and bias, it automatically feeds new experimental data back to the feature library and starts local incremental learning, only updating relevant model parameters. While absorbing new knowledge, it protects existing experience, enabling the entire system to continuously improve itself and adapt to new material systems as data accumulates, thus realizing the sustainable intelligence of the R&D process. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the overall structure of the present invention;

[0015] Figure 2 This is a schematic diagram of the S1 data construction of the present invention;

[0016] Figure 3 This is a schematic diagram illustrating the S2 model construction and fusion of the present invention;

[0017] Figure 4 This is a schematic diagram of the S3 reverse optimization of the present invention;

[0018] Figure 5 This is a schematic diagram of the S4 closed-loop update triggering of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] As attached Figure 1 Appendix Figure 2The machine learning-based method for predicting the performance of repair materials and optimizing their formulations includes S1: collecting and correlating the component space, process, and multidimensional performance data of the repair materials; wherein, the component space data includes the chemical identifiers, mass fractions, and physicochemical property vectors of the matrix, fillers, modifiers, and additives; the process data includes the time-series control parameters and environmental parameter vectors of the mixing, curing, or sintering processes; and the multidimensional performance data includes the quantitative results of mechanical, rheological, durability, and interfacial strength. The above data are aligned and integrated using the material formulation as an index to form a structured multimodal feature library.

[0021] It should be specifically noted that the structured multimodal feature library is constructed by building a dual index system of "unique formula identifier + process ID". The unique formula identifier is generated by the chemical identifier, mass fraction and micro-interface feature encoding of each component, and the process ID is generated by the core process parameters and time series response feature encoding. For component micro-interface feature data, process dynamic response time series data and multidimensional performance data, different methods are used for noise filtering, namely wavelet threshold denoising and data smoothing, sliding window anomaly detection and physical constraint verification, and 3σ criterion and performance correlation verification.

[0022] The key temporal features extracted from the process data include: viscosity peak value, peak occurrence time, initial viscosity value, final viscosity value, average and maximum slope of the viscosity rising and falling segments, average viscosity and duration of the steady segment, integral area of ​​the viscosity curve, number of viscosity abrupt change points and corresponding viscosity changes.

[0023] It should be further clarified that the data selection covers the entire chain of data, including "components-process-performance-environmental adaptability," specifically:

[0024] Innovative component space data: In addition to the basic chemical identifiers (CAS number, molecular structure formula) and mass fractions of the matrix, filler, modifier and auxiliaries, new “component micro-interface characteristic data” (such as the element diffusion depth at the matrix-filler interface and the crystallinity of the interfacial phase) and “dynamic component interaction data” (such as the hydrogen bond binding energy and van der Waals force strength between components at different temperatures).

[0025] Selection criteria: Existing technologies neglect the dominant influence of component micro-interface effects on performance, while the core properties of repair materials, including mechanical strength and durability, directly depend on the interface bonding state. Dynamic interactive data can reflect the real-time interaction patterns of components during the process, providing data support for subsequent models to accurately learn component correlations.

[0026] Innovative process data: In addition to the traditional timing control parameters and environmental parameters for mixing, curing / sintering, new "process dynamic response data" (such as the viscosity change curve of the system over time during mixing, and the timing data of the system's exothermic rate during curing) and "process-equipment matching data" (such as the shear rate distribution corresponding to the type of agitator, and the temperature gradient in different areas of the sintering furnace) have been added.

[0027] Selection criteria: Existing technologies only focus on process input parameters, ignoring the dynamic response during the process and the impact of equipment characteristics on the final structure. However, key structural indicators, including the uniformity and porosity of the repair material, are determined by the dynamic process. Equipment matching data can eliminate the problem of process parameter failure caused by differences in the operating conditions of different equipment.

[0028] Multidimensional performance innovation data: In addition to conventional mechanical properties (tensile strength, flexural strength), rheological properties (viscosity, thixotropy), durability (aging resistance, corrosion resistance) and interfacial strength (interfacial peel strength), new "dynamic service performance data" (such as fatigue life under cyclic loading, performance degradation curves in different environmental media (acids, alkalis, salts)) and "microstructure performance data" (such as grain size distribution, porosity and pore size distribution, crosslinking density) have been added.

[0029] Selection criteria: The actual service performance of repair materials depends on dynamic service performance rather than static performance indicators. Microstructure data can establish a link between "component-process-microstructure-macro performance" and solve the "black box" problem of performance prediction in existing technologies.

[0030] The entire processing strategy, including "multi-source data alignment, noise filtering, feature enhancement, and structured integration," is adopted. The specific steps are as follows:

[0031] A dual indexing system of "unique formula identifier + process ID" is constructed. The unique formula identifier is generated by the chemical identifier, mass fraction and micro-interface feature encoding of each component, while the process ID is generated by the core process parameters and time sequence response feature encoding. This ensures that component, process and performance data collected from different batches and different equipment can be accurately correlated, avoiding data mismatch caused by a single index.

[0032] The core process parameters for the mixing stage include mixing speed (r / min), mixing time (min), mixing temperature (°C), agitator type code, and material addition sequence code. The time-series response characteristics include the peak value (Pa·s) and occurrence time (min) of the system viscosity time-series curve, the steady-state viscosity value (Pa·s) and the time to reach steady state (min), and the system temperature fluctuation amplitude (°C) and the time of maximum fluctuation during mixing (min). The core process parameters for the curing stage include heating rate (°C / min), peak curing temperature (°C), holding time (min), and cooling rate (°C / min). The time-series response characteristics include the peak value (W / g) of the curing reaction exothermic rate and its occurrence time (min), and the degree of curing reaching 50%. The core process parameters for the sintering stage include the sintering atmosphere type, peak sintering temperature (°C), holding time (min), and heating rate (°C / min). The time-series response characteristics include the peak value of sample mass loss rate (mg / min) and its occurrence time (min) and the peak value of sample volume shrinkage rate (% / min) and its occurrence time (min) during sintering. The core process parameters for the auxiliary process include the pre-dispersion pressure (MPa) (for nanofillers), degassing vacuum degree (MPa), and degassing time (min). The time-series response characteristics include the minimum D90 value of the system particle size distribution and its occurrence time (min) during pre-dispersion and the peak value of bubble escape rate (number / min) and its occurrence time (min) during degassing.

[0033] Differentiated filtering methods are adopted to address the noise characteristics of different types of data: For component micro-interface feature data, a "wavelet threshold denoising + adjacent batch data smoothing" method is used to eliminate random noise during electron microscopy inspection; for process dynamic response data, a "sliding window anomaly detection + physical constraint verification" method is used to remove abnormal data points caused by equipment failure, where the physical constraints are based on the rheological theory of the repair material system (such as the viscosity change rate threshold over time); for performance data, a "3σ criterion + performance correlation verification" method is used to remove outliers caused by detection operation errors, while ensuring that the removed data meets the "theoretical correlation between microstructure performance and macroscopic performance" (such as the negative correlation between porosity and tensile strength).

[0034] For component data, a "molecular fingerprint encoding + physical property fusion" approach is used to generate physicochemical property vectors. The molecular structure of the components is transformed into a 1024-dimensional Morgan molecular fingerprint, which is then combined with physical properties and weighted and fused into a 512-dimensional component feature vector through an attention mechanism. For process time-series data, a "time-series feature extraction + trend encoding" approach is used. The time-series feature extraction network (TSE-Net) is used to extract viscosity peak, peak occurrence time, initial viscosity value, final viscosity value, average slope of viscosity rise, maximum slope of viscosity rise, average slope of viscosity fall, and viscosity fall. Sixteen key temporal features, including the maximum slope of the segment, the average viscosity of the steady segment, the duration of the steady segment, the viscosity fluctuation amplitude of the steady segment, the integral area of ​​the viscosity curve, the duration of the rising segment, the duration of the falling segment, the number of viscosity abrupt change points, and the viscosity change corresponding to the abrupt change points, are fused with the process control parameter vector and transformed into a 256-dimensional process feature vector through temporal attention encoding. For performance data, "multi-scale performance fusion" is adopted, and the microstructure performance, static performance, and dynamic service performance are reduced in dimensionality through principal component analysis (PCA) and then fused with the original key performance indicators to form a 128-dimensional performance feature vector.

[0035] A feature library is built based on a hybrid architecture of relational database (MySQL) and time-series database (InfluxDB): the relational database stores structured data including unique formula identifiers, basic component information, process control parameters, static properties (tensile strength, flexural strength, hardness, initial viscosity, thixotropy), and basic microstructure parameters (grain size, porosity); the time-series database stores time-series data including process dynamic response data (viscosity time-series curves during mixing, exothermic rate time-series data during curing), performance degradation curves (aging resistance degradation curves, performance degradation curves in acid and alkali media), and dynamic service performance time-series data (fatigue life time-series data under cyclic loading); a correlation mapping between the two is established through dual indexes, forming a full-chain structured multimodal feature library of "formula-process-microstructure-macro performance-dynamic service", and supporting rapid data retrieval, updating, and expansion.

[0036] As attached Figure 3 As shown, S2: Establish a set of heterogeneous models guided by physical mechanisms. The set includes: a first sub-model based on graph neural network to process component space data to learn component interaction relationships; a second sub-model based on temporal convolutional network to process process data to capture process-structure correlation; and a third sub-model based on gradient boosting tree to process physical properties to identify key influencing factors. A meta-learner is used to dynamically fuse the outputs of each sub-model according to the prediction confidence in different feature spaces to generate a joint prediction of multidimensional performance.

[0037] It should be specifically noted that the heteroproton model set is pre-trained using historical data, and a physical mechanism calibration step is introduced. The outputs of the three sub-models are calibrated using measured interfaces combined with energy data, microstructure data, and orthogonal experimental data. An end-to-end joint training and phased fine-tuning strategy is used to train the heteroproton model set and the meta-learner.

[0038] In the first sub-model, when constructing the dynamic component interaction graph, the molecular structure of the components is used as the node, and the estimated values ​​of hydrogen bond energy, chemical bond energy and interfacial bonding force between the components under specific process conditions are used as the edge to generate a weighted directed graph.

[0039] It should be further explained that a set of heterogeneous sub-models is constructed based on "physical mechanism constraints + differentiated feature adaptation". Each sub-model is specifically designed to handle different types of feature data. The specific design is as follows:

[0040] Graph Neural Network-Physical Mechanism Fusion Sub-model (GNN-PM): Used to process component spatial data and learn component interactions; constructs a dynamic component interaction graph with component molecules as nodes and interactions between components (such as hydrogen bonds, chemical bonds, and interfacial binding forces) as edges; embeds physical mechanism constraints into the loss function of the graph neural network to ensure that the interaction relationships learned by the model conform to thermodynamic and interfacial chemistry theories; the inputs are component molecule fingerprint encoding vectors, micro-interface feature vectors, and mass fractions, and the output is a component interaction feature matrix (reflecting the contribution weight of each component to the final performance).

[0041] Temporal Convolutional-Attention Fusion Sub-model (TCN-Att): Used to process process data and capture process-structure correlation; employs a temporal convolutional network (TCN) with dilated convolution and causal convolution to enhance feature extraction capabilities for long-term process data; introduces a process stage attention mechanism to automatically identify key temporal segments in different process stages (such as viscosity abrupt changes during the curing and heating stage); uses physical mechanisms as constraints on attention weights to ensure that the model focuses on process stages that have a significant impact on microstructure formation; inputs are process control parameter vectors and process dynamic response temporal feature vectors, and outputs are process-structure correlation feature vectors (reflecting the influence of process parameters on microstructure).

[0042] The Gradient Boosting Tree-Feature Interaction Enhancement Sub-model (XGBoost-FIE) is used to process physicochemical property data and identify key influencing factors. Based on the traditional Gradient Boosting Tree (XGBoost), a new feature interaction layer is added. This layer automatically learns high-order interaction features between physicochemical properties (such as the interaction between the matrix glass transition temperature and the filler surface energy) through a polynomial kernel function. A physical constraint regularization term is introduced to restrict feature interactions that contradict known physical laws (such as the unreasonable negative correlation between density and viscosity). The input is a vector of physicochemical properties after the fusion of components and processes, and the output is a ranking and contribution vector of key influencing factors.

[0043] A meta-learner fusion strategy of "confidence weighting + performance bias correction" is adopted, combined with incremental pre-training methods to improve the model's generalization ability. Specific steps are as follows:

[0044] Based on historical data in the feature library, three sub-models were pre-trained respectively. A physical mechanism calibration step was introduced: for the GNN-PM sub-model, the component interaction feature matrix output by the model was calibrated by combining the interface with measured data; for the TCN-Att sub-model, the process-structure correlation feature vector was calibrated by measured microstructure data (such as porosity); for the XGBoost-FIE sub-model, the accuracy of key influencing factors was verified by orthogonal experimental data to ensure that the sub-model output conforms to physical laws.

[0045] When each sub-model outputs its prediction results, it simultaneously calculates the prediction confidence score. The confidence score is determined by "historical statistical values ​​of model prediction errors + feature similarity of the current input data" (the higher the feature similarity between the input data and the training data, the greater the confidence score weight). The meta-learner uses an adaptive weight network, with the prediction confidence score of each sub-model as the initial weight, and dynamically adjusts the weights in combination with the prediction deviations of different performance indicators (such as mechanical performance prediction deviation and durability prediction deviation), and outputs a joint prediction result of multi-dimensional performance.

[0046] The strategy of "end-to-end joint training + phased fine-tuning" is adopted: first, the sub-model parameters are fixed and the weight network of the meta-learner is trained; then, the sub-model parameters are released, and the parameters of the three sub-models and the meta-learner are optimized simultaneously through backpropagation with the joint prediction error of multi-dimensional performance as the objective function; an early stopping mechanism is introduced, with the performance prediction deviation on the validation set being less than 5% as the stopping condition to avoid model overfitting.

[0047] As attached Figure 4 As shown, S3: Based on the trained prediction model, a closed-loop optimization process is constructed. Bayesian optimization is used to actively search for potential formulation regions in the high-dimensional solution space composed of components and processes. The prediction model is embedded into a differentiable formula simulator to calculate the gradient of the target performance with respect to the formulation parameters. Combined with the constrained optimization algorithm, the formulation is iteratively adjusted under process and cost constraints to generate an optimized scheme.

[0048] It should be specifically noted that the differentiable allocation simulator is constructed by integrating the diffusion equation and reaction kinetic equation of the repair material with the process-structure data learned by the sub-model.

[0049] It should be further explained that, for the high-dimensional solution space (dimension ≥ 50) composed of components and processes, a three-layer innovative strategy of "Bayesian optimization - dimensionality reduction mapping - potential region screening" is adopted to improve search efficiency:

[0050] Based on the data in the feature library, the high-dimensional component-process solution space is mapped to the low-dimensional latent space (dimension ≤ 10) by the variational autoencoder (VAE), preserving the key correlation structure in the solution space. At the same time, a bidirectional mapping relationship is established between the low-dimensional latent space and the high-dimensional original space to ensure that the search results in the low-dimensional space can be accurately restored to the original component-process parameters.

[0051] An improved Bayesian optimization algorithm is adopted, with "multi-dimensional performance comprehensive score (weights set by user requirements) - process cost - environmental adaptability" as the comprehensive objective function. "Potential region prior knowledge constraint" is introduced, and the key influencing factors identified by the XGBoost-FIE sub-model are used as prior constraints to narrow the search range (such as limiting the range of changes in the mass fraction of components that contribute ≥30% to performance). Active sampling is carried out in the low-dimensional latent space, and the comprehensive objective function value of the sampling points is evaluated by the prediction model. The Gaussian process regression model is iteratively updated until the potential region with the optimal comprehensive objective function is found.

[0052] Potential regions in low-dimensional space are restored to high-dimensional component-process parameters, and preliminary verification is carried out using a differentiable partition simulator to eliminate invalid parameter combinations caused by solution space mapping. For effective potential regions, local mesh refinement sampling is adopted to further explore the optimal parameter combinations.

[0053] A differentiable formula simulator driven by both physical mechanisms and data was constructed, and combined with a gradient-guided constraint optimization algorithm, to achieve precise iterative adjustment of formula parameters.

[0054] It integrates the physicochemical process mechanisms of repair materials (such as diffusion equations and reaction kinetic equations) with data-driven models (such as the process-structure correlation characteristics of the TCN-Att sub-model); ensuring that the simulator output (microstructure, multidimensional properties) is continuously differentiable with respect to the input components and process parameters, and can directly calculate the gradient of the target performance with respect to each formulation parameter.

[0055] Guided by the gradient calculated by the differentiable fractional simulator, an improved interior point method is used for constraint optimization. Constraints include process constraints (e.g., curing temperature ≤200℃, mixing time ≤60min), cost constraints (e.g., material cost per unit mass ≤50 yuan), and performance constraints (e.g., tensile strength ≥50MPa, aging resistance ≥1000h). The component mass fraction and process control parameters are iteratively adjusted. After each iteration, the optimization effect is evaluated through a prediction model until all constraints are met and the overall objective function is optimal, generating the final optimized scheme, including detailed component formulation, process parameters, expected performance, and cost estimation.

[0056] As attached Figure 5 As shown, S4: The new formula-performance data pairs obtained after the optimized formula preparation and verification are automatically fed back to the feature library. Triggering rules based on predictability and bias are set. When the new data is in a high uncertainty region or exceeds the bias threshold, the local incremental learning of the prediction model is automatically triggered.

[0057] It should be specifically noted that the triggering rule uses the Monte Carlo Dropout method to calculate the model's prediction uncertainty for new data. When the uncertainty is greater than or equal to a first preset threshold, it is determined to be triggered. The relative deviation between the measured performance and the model's predicted performance in the new data is calculated. When the relative deviation is greater than or equal to a second preset threshold, it is determined to be triggered. The model update process is started when any of the conditions are met.

[0058] The local incremental learning specifically includes the following steps: dividing the feature library data into a basic dataset and an incremental dataset, wherein the incremental dataset includes new data and historical highly similar data; updating only the model parameters related to the new data, while freezing the remaining basic parameters; evaluating the prediction accuracy of the updated model through a validation set; if the reduction in the average prediction bias reaches or exceeds a preset improvement threshold, retaining the update results and updating and optimizing prior knowledge; otherwise, rolling back the model parameters and readjusting the incremental learning range.

[0059] It should be further explained that the specific steps for achieving automated and intelligent feedback of optimized formulation preparation validation data are as follows:

[0060] During the optimization and validation of the formulation, automated testing equipment is used to collect new formulation-performance data pairs, including actual component mass fractions, actual process control parameters, microstructure data, and static and dynamic performance data. The data is then standardized using the data processing method in S1 to generate structured data that conforms to the feature library format.

[0061] Based on the dual index of "unique formula identifier + process ID", the standardized new data is automatically associated with the corresponding formula-process entry in the feature library; at the same time, the equipment information and testing environment parameters of the data collection are recorded to form a complete data traceability chain.

[0062] Verify whether the performance deviation between the new data and the output of the prediction model is within the allowable range of detection error (e.g., mechanical performance detection error ≤ 3%); calculate the relative standard deviation (RSD) of the data by repeating the detection 3 times. Only new data with RSD ≤ 5% can be included in the feature library to ensure the reliability of the feedback data.

[0063] By employing a dual-trigger rule of "prediction uncertainty + performance bias" combined with a local incremental learning method, efficient and accurate model updates are achieved. The specific steps are as follows:

[0064] The Monte Carlo dropout method is used to calculate the prediction uncertainty of the formula-process parameters corresponding to the new data. When the uncertainty is greater than or equal to the first preset threshold (e.g., 0.2, determined by the dispersion of the feature library data), the new data is determined to be in the high uncertainty region, triggering a model update. The relative deviation between the measured performance and the model prediction performance in the new data is calculated. When the relative deviation is greater than or equal to the second preset threshold (e.g., 8%, determined by the detection accuracy and engineering requirements), a model update is triggered. When any triggering condition is met, the model update process is started to avoid invalid or delayed updates.

[0065] The historical data in the feature library is divided into a basic dataset (80%, covering the mainstream formulations and processes) and an incremental dataset (new data + historical highly similar data, accounting for 20%). Only the model parameters related to the new data are updated, namely the parameters of the corresponding component interaction relationships in the GNN-PM sub-model, the parameters of the corresponding process stages in the TCN-Att sub-model, and the weight parameters of the corresponding performance indicators in the meta-learner. The basic parameters of the model that are not related to the new data are frozen to avoid the model's prediction accuracy for the original mature formulations and processes decreasing due to incremental learning. The prediction accuracy of the updated model is evaluated through the validation set. If the average prediction bias of the updated model decreases by ≥2%, the updated result is retained; otherwise, the model parameters are rolled back to the original parameters, and the parameter range of incremental learning is re-optimized.

[0066] The new correlation patterns learned by the updated model (such as the interaction of new components and the influence of new process parameters) are transformed into structured knowledge and added to the prior constraints of Bayesian optimization to improve the efficiency of subsequent optimization processes.

[0067] Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other.

[0068] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for predicting the performance of repair materials and optimizing their formulations based on machine learning algorithms, characterized in that, include: S1: Collect and correlate the component space, process, and multidimensional performance data of the repair materials; among which, the component space data includes the chemical identifiers, mass fractions, and physicochemical property vectors of the matrix, fillers, modifiers, and additives; the process data includes the time-series control parameters and environmental parameter vectors of the mixing, curing, or sintering processes; the multidimensional performance data includes the quantitative results of mechanical, rheological, durability, and interfacial strength; align and integrate the above data using the material formulation as an index to form a structured multimodal feature library; S2: Establish a set of heterogeneous models guided by physical mechanisms. The set includes: a first sub-model based on graph neural network to process component space data to learn component interaction relationships; a second sub-model based on temporal convolutional network to process process data to capture process-structure correlation; and a third sub-model based on gradient boosting tree to process physical properties to identify key influencing factors. A meta-learner is used to dynamically fuse the outputs of each sub-model according to the prediction confidence in different feature spaces to generate a joint prediction of multidimensional performance. S3: Based on the trained prediction model, a closed-loop optimization process is constructed. Bayesian optimization is used to actively search for potential formulation regions in the high-dimensional solution space composed of components and processes. The prediction model is embedded into a differentiable formula simulator to calculate the gradient of the target performance with respect to the formulation parameters. Combined with the constrained optimization algorithm, the formulation is iteratively adjusted under process and cost constraints to generate an optimized solution. S4: Automatically feed the new formula-performance data pairs obtained after the optimized formula preparation verification to the feature library, set trigger rules based on prediction uncertainty and bias, and automatically trigger local incremental learning of the prediction model when the new data is in the high uncertainty region or exceeds the bias threshold. The structured multimodal feature library is constructed by building a dual index system of "unique formula identifier + process ID". The unique formula identifier is generated by the chemical identifier, mass fraction and micro-interface feature encoding of each component, and the process ID is generated by the core process parameters and time series response feature encoding. For component micro-interface feature data, process dynamic response time series data and multidimensional performance data, different methods are used for noise filtering, namely wavelet threshold denoising and data smoothing, sliding window anomaly detection and physical constraint verification, and 3σ criterion and performance correlation verification. The key temporal features extracted from the process data include: viscosity peak value, peak occurrence time, initial viscosity value, final viscosity value, average and maximum slope of the viscosity rising and falling segments, average viscosity and duration of the steady segment, integral area of ​​the viscosity curve, number of viscosity abrupt change points and corresponding viscosity changes. In the first sub-model, when constructing the dynamic component interaction graph, the molecular structure of the components is used as the node, and the estimated values ​​of hydrogen bond energy, chemical bond energy and interfacial bonding force between the components under specific process conditions are used as the edge to generate a weighted directed graph. The differentiable partitioning simulator is constructed by integrating the diffusion equation and reaction kinetic equation of the repair material with the process-structure data learned by the sub-model.

2. The method for predicting the performance of repair materials and optimizing their formulations based on machine learning algorithms according to claim 1, characterized in that: The heteroproton model set is pre-trained using historical data, and a physical mechanism calibration step is introduced. The outputs of the three sub-models are calibrated using measured interfaces combined with energy data, microstructure data, and orthogonal experimental data. The heteroproton model set and meta-learner are trained using an end-to-end joint training and phased fine-tuning strategy.

3. The method for predicting the performance of repair materials and optimizing their formulations based on machine learning algorithms according to claim 1, characterized in that: The triggering rules employ the Monte Carlo Dropout method to calculate the model's prediction uncertainty for new data. When the uncertainty is greater than or equal to a first preset threshold, a trigger is determined. The relative deviation between the measured performance and the model's predicted performance in the new data is calculated. When the relative deviation is greater than or equal to a second preset threshold, a trigger is determined. The model update process is initiated when any of the conditions are met.

4. The method for predicting the performance of repair materials and optimizing their formulations based on machine learning algorithms according to claim 1, characterized in that: The local incremental learning specifically includes the following steps: dividing the feature library data into a basic dataset and an incremental dataset, wherein the incremental dataset includes new data and historical highly similar data; updating only the model parameters related to the new data, while freezing the remaining basic parameters; evaluating the prediction accuracy of the updated model through a validation set; if the reduction in the average prediction bias reaches or exceeds a preset improvement threshold, retaining the update results and updating and optimizing prior knowledge; otherwise, rolling back the model parameters and readjusting the incremental learning range.