A method, system, and medium for strength prediction of a steel slag-based cementitious material

CN122551983APending Publication Date: 2026-08-11JIANGXI PROVINCIAL EXPRESSWAY INVESTMENT GRP CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种钢渣基胶凝材料的强度预测方法、系统和介质,解决了现有技术中存在的纯数据驱动的机器学习模型在进行胶凝材料配合比寻优时,因缺乏物理化学定律约束而输出违背质量守恒与化学平衡的“物理不可行配方”,从而导致极高工程返工率的底层技术问题

Benefits of technology

1、本发明通过建立包含“组分质量分数之和约束”以及“总氧化物含量约束”的物理化学约束集,并将其直接嵌入序列最小二乘规划(SLSQP)算法的迭代求解过程中。该方案改变了传统机器学习盲目进行纯数学寻优的黑盒模式,利用拉格朗日乘子与罚函数强行限制迭代游标不得越过宏观质量守恒与微观化学平衡的物理红线。与现有技术相比,该方案彻底杜绝了组分总和不闭合或氧化物比例畸变的无效配方,将工程试配返工率降至极低水平;同时,约束条件的引入不仅未损害数学拟合精度,反而像“灯塔”一样引导算法避开了抽象空间内的局部伪最优解,显著降低了配方寻优的实际平均误差。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551983A_ABST
    Figure CN122551983A_ABST
Patent Text Reader

Abstract

This invention relates to a method, system, and medium for predicting the strength of steel slag-based cementitious materials, belonging to the field of material performance prediction and formulation optimization technology. The method includes: acquiring mix proportion data of the steel slag-based cementitious materials; establishing a strength prediction model using a machine learning algorithm; establishing an objective function with the goal of minimizing the absolute value of the difference between the model's predicted strength value and the target strength value; constructing a set of physicochemical constraints including constraints on the sum of component mass fractions and total oxide content; iteratively solving the objective function using a sequential least squares programming algorithm, restricting the iteration points using the constraint set during the solution process, and finally outputting the optimal mix proportion data. This invention effectively avoids the problem of purely data-driven models outputting physically infeasible formulations by embedding the hydration mechanism and physical conservation laws of materials science as mathematical hard constraints into the bottom layer of the optimization algorithm, achieving accurate mix proportion inversion and intelligent optimization with low error.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of material performance prediction and formulation optimization technology, specifically relating to a method, system and medium for predicting the strength of steel slag-based cementitious materials. Background Technology

[0002] As a major industrial solid waste, the resource utilization of steel slag in cementitious material systems has always been a research hotspot in the building materials field. The compressive strength of steel slag-based cementitious materials is affected by the nonlinear coupling of various complex raw materials such as steel slag powder, slag powder, and desulfurized gypsum. The traditional method of trial mixing based on manual experience is time-consuming and costly, and cannot meet the needs of modern engineering for rapid response and precise quantitative mixing ratios.

[0003] In recent years, with the development of artificial intelligence technology, establishing a mapping relationship between mix proportions and macroscopic strength using machine learning algorithms, and then optimizing the mix proportions in reverse to obtain the target mix proportion, has gradually become a new technical approach. For example, Chinese patent document CN104991051A discloses a concrete strength prediction method based on a hybrid model, which predicts concrete strength by collecting mix proportion data of raw materials such as cement and slag powder and using machine learning modeling methods. This purely data-driven technical solution improves prediction efficiency to a certain extent, but with its in-depth application in industrial settings, a specific underlying algorithmic defect has gradually been exposed.

[0004] Existing prediction and optimization methods typically treat machine learning models as purely mathematical "black boxes," relying entirely on gradient descent or heuristic search algorithms to find mathematical extrema in a multidimensional feature space. Because these algorithms lack the constraints of fundamental physical and chemical laws of materials science, the optimization solvers often converge blindly in an abstract mathematical space detached from actual physical meaning. Specifically, in an attempt to force a fit to the target strength, the algorithm easily outputs mathematically optimal solutions where the sum of the mass fractions of each raw material component deviates significantly from 100% (e.g., the total of all components reaches 110%), or where the proportions of key oxides within the mixture completely violate chemical common sense. Such mathematically perfect proportions are simply impossible to achieve in the physical weighing and chemical hydration reactions of real-world mixing plants, resulting in a persistently high rate of rework during trial mixing.

[0005] Furthermore, existing technologies typically use the absolute dosage of each raw material directly when constructing model input features. This feature matrix, lacking dimensionality reduction based on materials science mechanisms, makes it difficult for algorithms to capture the potential activation patterns of steel slag and mineral slag under specific alkaline environments. Given the extremely high cost of obtaining industrial experimental data and the extremely limited sample size (e.g., hundreds of samples), simply relying on the original dosage features often leads to severe overfitting and weak generalization ability. Therefore, effectively embedding physicochemical constraints into the underlying optimization algorithms of machine learning to guide the model to meaningfully optimize within the real physical feasible domain is a specific fundamental technical challenge that urgently needs to be addressed in this field. Summary of the Invention

[0006] The purpose of this invention is to provide a method, system, and medium for predicting the strength of steel slag-based cementitious materials. This invention solves the fundamental technical problem in the prior art where purely data-driven machine learning models, when optimizing the mix proportions of cementitious materials, output "physically infeasible formulas" that violate the law of conservation of mass and chemical equilibrium due to the lack of constraints from physical and chemical laws, thus leading to extremely high engineering rework rates.

[0007] The objective of this invention can be achieved through the following technical solutions: A method for predicting the strength of steel slag-based cementitious materials includes acquiring mix proportion data of the steel slag-based cementitious materials and establishing a strength prediction model using a machine learning algorithm; the method further includes: An objective function is established, wherein the objective function aims to minimize the absolute value of the difference between the predicted value of the intensity prediction model and the preset target intensity value. Construct a physicochemical constraint set, which includes constraints on the sum of component mass fractions and total oxide content; The objective function is solved iteratively using a sequential least squares programming algorithm. During the solution process, the physicochemical constraint set is used to restrict the iteration points, and the optimal mix ratio data is output.

[0008] Furthermore, the sum of the component mass fractions in the physicochemical constraint set is constrained to ensure that the deviation between the sum of the mass fractions of steel slag powder, slag powder, desulfurized gypsum, cement and fly ash and the preset value is within a preset tolerance range; the total oxide content constraint is that the sum of the contents of calcium oxide, silicon dioxide, ferric oxide, aluminum oxide and magnesium oxide in the mixed system is within a preset range.

[0009] Furthermore, the input features of the strength prediction model include derived features constructed based on the hydration mechanism of cementitious materials, such as alkalinity index and calcium-silicon ratio.

[0010] Furthermore, the alkalinity index is obtained by dividing the sum of the contents of calcium oxide and magnesium oxide in the mixed system by the sum of the contents of silicon dioxide and aluminum oxide; the calcium-silicon ratio is obtained by dividing the calcium oxide content in the mixed system by the silicon dioxide content.

[0011] Furthermore, before establishing the intensity prediction model using machine learning algorithms, the physical consistency of the mix proportion data is also checked. When the sum of the component mass fractions or the total oxide content of the data sample deviates from the preset range, a normalization correction strategy is used to process the feature values ​​of the data sample.

[0012] Furthermore, after the dataset is partitioned, the normalization correction strategy is applied to the abnormal data samples in the training set, while only diagnostic records are processed for the abnormal data samples in the test set without any correction.

[0013] Furthermore, before performing the iterative solution, the intensity prediction model is analyzed using Shapley value theory to extract the nonlinear response range of component characteristics to intensity.

[0014] Furthermore, the nonlinear response interval is embedded as a dynamic optimization boundary into the physicochemical constraint set, and the sequential least squares programming algorithm is guided to converge by compressing the parameter search space.

[0015] A strength prediction system for steel slag-based cementitious materials includes: The data preprocessing module is used to acquire the mix proportion data of steel slag-based cementitious materials and perform physical consistency verification. The model building module is used to build intensity prediction models using machine learning algorithms; The constraint optimization module is used to construct the objective function and perform sequential least squares programming in combination with the physicochemical constraint set to output the optimal mix ratio data.

[0016] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0017] The beneficial effects of this invention are: 1. This invention establishes a set of physicochemical constraints, including constraints on the sum of component mass fractions and total oxide content, and directly embeds them into the iterative solution process of the Sequential Least Squares Programming (SLSQP) algorithm. This approach changes the traditional black-box mode of machine learning, which blindly seeks pure mathematical optimization. It uses Lagrange multipliers and penalty functions to forcibly restrict the iterative cursor from crossing the physical red lines of macroscopic mass conservation and microscopic chemical equilibrium. Compared with existing technologies, this approach completely eliminates invalid formulations with non-closed sums of components or distorted oxide ratios, reducing the rework rate of engineering trial formulations to an extremely low level. At the same time, the introduction of constraints not only does not impair the accuracy of mathematical fitting, but also acts like a "lighthouse" to guide the algorithm to avoid local pseudo-optimal solutions in the abstract space, significantly reducing the actual average error of formulation optimization.

[0018] 2. In the model feature construction stage, this invention constructs derived features such as alkalinity index and calcium-silicon ratio based on the hydration kinetics mechanism of cementitious materials, and combines this with a rigorous physical consistency verification and differentiated normalization correction mechanism. This approach explicitly transforms prior knowledge of materials science into mathematical dimensions readable by the algorithm in advance, significantly reducing the computational overhead of the decision tree model in fitting the implicit mapping of complex chemical reactions. This enables the prediction model to establish a highly robust prediction surface that conforms to the laws of the real physical world, even under the limited condition of relying only on a very small sample dataset, avoiding the overfitting phenomenon caused by deep networks forcibly fitting the original doping features.

[0019] 3. This invention further introduces the Shapley value theory analytical model to accurately extract the nonlinear response range of component characteristics to intensity, and uses this as a dynamic optimization boundary to feed back into the constrained optimization algorithm. This scheme upgrades the static "physical conservation defense line" to a dynamic "chemical activity defense line," precisely compressing the enormous parameter search space into the interior of the most active manifold of microscopic chemical reactions. This not only significantly reduces the number of iterations and computational power required for the quasi-Newton algorithm to approach the target extreme point, but also ensures that the final output mix ratio accurately falls within the golden sweet zone of the material's hydration reaction, greatly improving the early measured strength of the material. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the SLSQP optimization logic with embedded hard constraints in this invention. Figure 2 This is a flowchart illustrating the derivation and differentiation verification of the hydration mechanism of this invention. Figure 3 This is a schematic diagram of the SHAP dynamic boundary feedback optimization principle of the present invention. Detailed Implementation

[0022] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0023] Example 1 like Figure 1 As shown, a method for predicting the strength of steel slag-based cementitious materials includes acquiring mix proportion data of the steel slag-based cementitious materials and establishing a strength prediction model using a machine learning algorithm. The method further includes: establishing an objective function, wherein the objective function aims to minimize the absolute value of the difference between the predicted value of the strength prediction model and a preset target strength value; constructing a physicochemical constraint set, which includes constraints on the sum of component mass fractions and total oxide content; and iteratively solving the objective function using a sequential least squares programming algorithm, using the physicochemical constraint set to restrict the iteration points during the solution process, and outputting optimal mix proportion data.

[0024] The constraint on the sum of the mass fractions of the components in the physicochemical constraint set is that the deviation of the sum of the mass fractions of steel slag powder, slag powder, desulfurized gypsum, cement and fly ash from the preset value is within the preset tolerance range; the constraint on the total oxide content is that the sum of the contents of calcium oxide, silicon dioxide, ferric oxide, aluminum oxide and magnesium oxide in the mixed system is within the preset range.

[0025] A strength prediction system for steel slag-based cementitious materials further includes: a data preprocessing module for acquiring mix proportion data of steel slag-based cementitious materials and performing physical consistency verification; a model building module for establishing a strength prediction model using machine learning algorithms; and a constraint optimization module for constructing an objective function and performing sequential least squares programming to solve the problem in combination with a set of physicochemical constraints, and outputting optimal mix proportion data.

[0026] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, performs the steps of the method described above.

[0027] The current commercial mixing plant batching system, based on SLSQP and physicochemical hard constraints, frequently outputs invalid proportions when conventional pure data-driven models are used for formula prediction, resulting in deviations in the sum of the mass fractions of each component or severe imbalances in the proportion of microscopic oxides. Such mathematical solutions, deviating from the laws of materials science, lead to extremely high rework rates in actual batching experiments.

[0028] This embodiment addresses the production bottleneck by deploying a computing server equipped with a central processing unit, 256GB of random access memory, and a high-speed data bus to perform mix proportion optimization. The server reads 111 sets of original mix proportion data streams of steel slag-based cementitious materials and corresponding 7-day measured compressive strength data from an external steel company's experimental database via a high-bandwidth communication interface.

[0029] The central processing unit performs quantile truncation on the read data, calculates the 5th and 95th quantiles of the feature dimensions, and limits the extreme values ​​within this range, eliminating gradient calculation instability caused by sensor measurement noise. The feature vector input to the server covers five dimensions of quality score indicators. Steel slag powder is rich in free calcium oxide and has self-hardening properties; blast furnace slag powder exhibits a glassy structure and has potential cementing properties; desulfurized gypsum provides sulfate ions to stimulate the formation of ettringite crystals; cement provides an early-strength hydrated calcium silicate framework; and fly ash improves the micropore structure through morphological effects.

[0030] The central processing unit (CPU) calls its internal arithmetic logic unit to execute the extreme gradient boosting algorithm code to build the strength prediction model. In the early trial-and-error phase of algorithm selection, support vector machine regression or fully connected neural network algorithms were attempted. Support vector machine regression was extremely sensitive to the penalty coefficient hyperparameter when dealing with problems such as the nonlinear decay of compressive strength caused by drastic changes in steel slag powder content; the high-dimensional space mapping calculation of the kernel function caused the time consumed in a single parameter tuning to exceed the tolerance of industrial deployment. Fully connected neural networks frequently got stuck in local minima and suffered severe overfitting on small datasets of around 100 samples.

[0031] The extreme gradient boosting algorithm employs a gradient boosting strategy and a second-order Taylor expansion optimization method. It utilizes the first and second derivatives of the loss function to determine the optimal split point, approximating the minimum value of the loss function. This approach demonstrates significant advantages in sparse data perception and computational efficiency. The central processing unit iteratively constructs a series of basic decision tree weak learners in the cache. The outputs of these weak learners are linearly combined to fit a nonlinear mapping network between raw material components and macroscopic intensity.

[0032] Suppose the model contains in a specific iteration batch The mathematical formula for calculating the prediction strength value of a decision tree for any feature vector input to the server is defined as follows: In the formula Indicates that the central processing unit (CPU) is for the first... The predicted compressive strength assessment value output from each data sample. This represents the total number of decision trees currently embedded in the system. Indicates the first The independent spatial mapping structure of decision trees, The first step of the feed algorithm The multidimensional feature vectors of each matching sample are used. The central processing unit calculates a composite loss function including regularization, the evaluation benchmark of which is defined as: In the formula To measure the actual compressive strength tested in the laboratory With model prediction and evaluation strength Differentiable loss scale of the absolute deviation between them This is a regularization term used to control the growth depth of the model.

[0033] The central processing unit is based on the formula: Calculate the specific numerical value, where This represents the total number of leaf nodes generated by the decision tree in memory. This represents the weight vector space stored in each leaf node. and These are hyperparameters used by the central processing unit to control the splitting saturation threshold of the algorithm.

[0034] By calculating the candidate set of feature split points and restricting the L2 norm of the leaf node weights, the central processing unit serializes and stores the trained model parameters in non-volatile memory. After testing and evaluation, the model achieved a prediction performance of 0.7156 on the test set, with the root mean square error controlled within 3.27 MPa.

[0035] After completing the prediction model construction, the central processing unit (CPU) transitions to the intelligent optimization stage of the proportioning. The core technical challenge of this optimization stage lies in transforming the hydration reaction law of the cementitious material into unambiguous mathematical boundary constraints, forcibly guiding the computational cursor to move within the physically feasible manifold space. Early engineering tests evaluated a post-selection method. This method uses Monte Carlo random sampling to generate 100,000 sets of formula inputs for model prediction, retaining numerical solutions with prediction errors less than 2 MPa. The verification mechanism identified that over 98% of the numerical solutions were in the oxide imbalance dead zone; for example, the ratio of calcium oxide to silica in the mixed system deviated significantly from the reasonable range for the formation of hydrated calcium silicate gel. The remaining 2% of solutions, when converted into actual mixing station feed weight instructions, showed that the sum of the mass fractions of each component did not equal 100%, and the proportioning could not be closed. An attempt was made to use a genetic algorithm to perform constrained optimization. However, the genetic algorithm relies on a large dynamic penalty coefficient for correction when dealing with stringent equality constraints.

[0036] Test data shows that even after 500 generations of evolution, the genetic algorithm population still exhibits a total deviation of 0.02 in the proportion of some individuals, exceeding the 0.005-point engineering weighing tolerance of sensors in commercial mixing plants. To address these limitations and algorithmic bottlenecks, this embodiment employs a sequential least squares programming algorithm as the core solver. The central processing unit constructs the objective function in its logic registers.

[0037] Let the preset expected target intensity value be The specific mathematical formula for the optimization objective function is defined as follows: In the formula This is a feature vector matrix containing various raw material blending ratios to be solved. The objective function drives the underlying algorithm to find feature matching points that make the difference between the model's predicted intensity and the expected intensity approach the zero dead zone. The central processing unit dynamically allocates verification addresses in random access memory to construct a physicochemical constraint set containing three lines of defense. The first line of defense in the constraint set is the constraint of the sum of component mass fractions.

[0038] The formula for determining whether the iterative mix proportions generated by the algorithm possess the real-world mass conservation property is set as follows: In the formula The values ​​represent the mass fraction percentages of the five main materials: steel slag powder, blast furnace slag powder, desulfurization gypsum, cement, and fly ash. This represents the preset tolerance threshold for the weighing sensor's execution error in the tolerance system. The central processing unit strictly locks this tolerance threshold at 0.005. The total mass of the five main materials must converge infinitely close to 1, preventing invalid formulations from being generated with a total of 1.03 or 0.95. The second deep line of defense in the constraint set is the total oxide content constraint. The microscopic hydration product network depends on the balance of the ratio of metal and non-metal oxides.

[0039] The central processing unit sets the control formula for the total content of oxides in the mixed system as follows: In the formula This indicates the first component introduced into the system after being converted by mass from the five main raw material components. The absolute mass fractions of various oxides are defined, specifically covering calcium oxide, silicon dioxide, ferric oxide, aluminum oxide, and magnesium oxide. Controlling the total oxide content within the range of 0.9 to 1.1 allows for slight fluctuations in chemical composition between batches of industrial by-products, while simultaneously preventing algorithms from attempting to suppress the objective function error by randomly piling up highly reactive components. The third line of defense in the constraint set is non-negative mathematical constraints. The central processing unit specifies the characteristic values ​​of all components at the instruction level. The algorithm fails to generate feature parameters with quality less than zero in any gradient descent direction. Once the memory stack required for the objective function and the physicochemical constraint set is ready, the central processing unit initiates a high-frequency iterative solution using a sequential least squares programming algorithm. This sequential least squares programming algorithm is based on a quasi-Newton optimization method with limited memory, handling boundary constraints through a double-loop recursive algorithm and an activity set strategy.

[0040] In each millisecond-level solution iteration, the high-energy arithmetic unit of the central processing unit calculates the high-dimensional gradient matrix information of the objective function at the current coordinate point based on the function structure of the prediction model. The central processing unit maintains recent gradient changes and variable update data, and dynamically constructs an approximate model of the inverse of the Hessian matrix. .

[0041] The system calculates the search direction for the region where the pointing error is minimized: And perform a line search to determine the iteration step size. In the variable update phase In this system, an activity set strategy and projection operation are introduced. The newly generated mix ratio coordinate array is substituted into the constraint set for verification. When a parameter change exceeding the limit is detected due to a specific iteration direction, such as the sum of the five components reaching 1.012 and triggering a tolerance alarm, the processor immediately injects a damping cost into the Lagrange function. The system uses Lagrange multipliers combined with a penalty function to forcibly reproject the out-of-bounds point back into the physically feasible domain space that strictly satisfies the upper and lower limit conditions.

[0042] The matrix multiplication instructions for physical correction attributes are executed cyclically within the processor. The central processing unit updates the specific values ​​of the mix proportion vector based on penalty feedback. When the absolute value difference of the objective function in two consecutive iterations falls below the convergence dead zone, the algorithm reaches dynamic mathematical equilibrium. The system stops the iteration loop and outputs the optimal mix proportion data that satisfies all constraint boundary conditions and approximates the target strength to the external memory via the system bus. Unconstrained optimization, due to ignoring material properties, produces more than half of the severely non-compliant mix proportion samples during verification, resulting in a rework rate as high as 83.3%. The constrained optimization in this embodiment reduces the rework rate to zero and compresses the average strength error to 2.69 MPa.

[0043] The steel slag-based cementitious material strength prediction system architecture in this embodiment adopts a modular software design. The data preprocessing module is scheduled by the operating system kernel thread, capturing material proportion data through input / output interfaces and performing consistency checks to eliminate non-closed samples. The model building module calls a graphics processing unit (GPU) or a central processing unit (CPU) core to execute an extreme gradient boosting algorithm to assemble and verify the decision tree array. The constraint optimization module establishes a mathematical objective function for the absolute value of the prediction difference and loads the aforementioned multiple boundary defenses. It then activates a sequential least squares programming algorithm library to traverse and solve the mix proportion vector space, ultimately serializing the multidimensional parameters into a human-readable material preparation document.

[0044] The non-volatile computer-readable storage medium described in this embodiment is implemented as an enterprise-grade solid-state drive or high-density disk array unit. Binary computer program instructions are physically stored within the sectors of its internal silicon-based memory chips. After the storage medium is mounted to the server file system, the memory controller reads these instructions into the running memory for the computing core to execute them one by one, stably reproducing the entire software and hardware interaction process of establishing a predictive model, setting the objective function, attaching physical and chemical constraint defenses, and using nonlinear iterative algorithms to perform precise inversion of engineering formulas.

[0045] Example 2 like Figure 2 As shown, the input features of the strength prediction model include derived features constructed based on the hydration mechanism of cementitious materials. These derived features include the alkalinity index and the calcium-silicon ratio. The alkalinity index is obtained by dividing the sum of the calcium oxide and magnesium oxide contents in the mixed system by the sum of the silicon dioxide and aluminum oxide contents; the calcium-silicon ratio is obtained by dividing the calcium oxide content in the mixed system by the silicon dioxide content.

[0046] Before establishing the intensity prediction model using machine learning algorithms, the physical consistency of the mix proportion data is verified. When the sum of the component mass fractions or the total oxide content of the data samples deviates from a preset range, a normalization correction strategy is used to process the feature values ​​of the data samples. After the dataset is divided, the normalization correction strategy is applied to the abnormal data samples in the training set, while only diagnostic records are executed for the abnormal data samples in the test set without correction.

[0047] The batching control systems of commercial concrete mixing plants and special building material production lines typically rely on manufacturing execution systems (MES) to record historical material feeding data. Weighing belt scales in industrial sites face long-term mechanical wear and zero-point drift, and the moisture content of powder materials such as fly ash and desulfurized gypsum fluctuates non-linearly with ambient humidity.

[0048] These physical deviations between the hardware and the environment prevented the original mix proportion data directly exported from the MES system from being mathematically closed. In this embodiment, the computing workstation, via a gigabit Ethernet fiber optic link and TCP / IP protocol, directly extracted 111 sets of original weighing data and corresponding 7-day compressive strength breaking test records of steel slag-based cementitious materials from the relational database of the steel enterprise's laboratory information management system in batches. In the early trial-and-error phase of model engineering, the conventional approach is to use statistical outlier removal methods, such as the Raida criterion (3-Sigma rule) or the box plot quartile truncation method.

[0049] When the system performed 3-Sigma cleaning and verification on these 111 sets of data, it directly identified 24 sets as abnormal and discarded them. For internet recommendation algorithms with millions of samples, eliminating 20% ​​of the data is a routine operation. In the field of materials science, the cost of sample acquisition is limited by physical time. Each set of cementitious material data corresponds to a curing cycle of 7 to 28 days, including raw material homogenization, molding, standard constant temperature and humidity curing, and press molding.

[0050] With only 87 samples remaining, the decision tree encountered a shortage of leaf node samples when splitting to the third layer due to insufficient leaf node samples. The model exhibited an error of 1.2 MPa on the training set, but the root mean square error on the validation set soared to 6.8 MPa, indicating severe overfitting. Existing R&D budgets and project schedules could not support the low-cost expansion of hundreds of validation data sets through repeated experiments in the short term. The technical team attempted to use the K-nearest neighbors algorithm or mean interpolation to mathematically impute these outliers. Mean interpolation, however, completely altered the original relative proportions of steel slag and desulfurized gypsum within the samples when patching local data, resulting in a chemometrically fictitious material structure in the generated formulation.

[0051] Based on the aforementioned constraints, this scheme abandons statistical elimination and mathematical interpolation. For the 111 extremely small datasets, a physical consistency verification mechanism based on normalization correction is deployed within the central processing unit (CPU). The CPU parses the feature matrix row by row. The first verification is a physical conservation check of the macroscopic proportions. The CPU extracts the weight fractions of the five main cementitious materials—steel slag powder, blast furnace slag powder, desulfurized gypsum, cement, and fly ash—from the row vector and performs a summation calculation, defined as follows: In the formula The original measured mass proportions of the five basic raw materials in the corresponding sample. The control formula for determining whether this sum deviates from the conservation state is set as follows: In the formula To tolerate the preset threshold of dynamic error in weighing equipment. The static calibration error of powder scales in commercial mixing plants typically fluctuates between 0.2% and 0.3%. Adding the aerodynamic suspension error during the material feeding process, this solution will... It solidified to 0.005.

[0052] When the calculation results show that the sum of the sample mass is outside the limit, the central processing unit triggers a normalization correction. The correction formula is defined as follows: In the formula This is a newly corrected quality fraction feature value that is overwritten into memory.

[0053] This formula forcibly constrains the total mass of the mixture to an absolutely conserved state, while simultaneously scaling proportionally to preserve the relative dosing relationships between materials in the original experimental design. In-depth verification of the microscopic oxide content constitutes a second line of defense. The compressive strength of cementitious materials originates from the microcrystalline network generated by the hydration reaction, and the density of this network depends on the absolute content of metal and non-metal oxides. The central processing unit extracts the mass fractions of calcium oxide, silicon dioxide, ferric oxide, aluminum oxide, and magnesium oxide measured by X-ray fluorescence spectrometry (XRF) of the sample, and calculates the total oxide content: In the formula This indicates the measured mass percentage of a single oxide. XRF testing involves loss on ignition (LOI), and carbonate decomposition and moisture evaporation result in a total detection percentage that is not 100%. Some trace elements, such as titanium, phosphorus, and potassium, are not included in the core oxide list.

[0054] Forcing the sum of these five principal oxides to equal 1.0 at the algorithm level would distort the true compositional distribution of the matter. The discrimination control formula set by the central processing unit is as follows: Lower limit value Set to 0.9, maximum value Set to 1.1.

[0055] When the total oxide content of a sample exceeds the above critical range, for example, if the total is only 0.82, it indicates that the batch of raw materials contains excessive inert impurities or that there is a systematic missed detection in the testing process. The central processing unit uses the formula... Initiate micro-normalization correction. This represents the corrected proportion of individual oxides. After completing the feature matrix correction, the system uses a pseudo-random number generator to divide it into training set memory blocks and test set memory blocks at a 7:3 ratio.

[0056] During the model deployment and cross-validation phase, if the test set undergoes the same standard of normalization correction, the model evaluation stage deals with a noise-free data stream that has been arithmetically smoothed. However, the sensor data stream returned by the industrial field programmable logic controller (PLC) carries inherent measurement biases. This solution implements differentiated isolation processing within the logic control flow. For anomalous samples residing in the training set memory block, the CPU enforces a normalization correction formula, limiting the machine learning algorithm to fit the objective function within a data manifold that conforms to physicochemical laws. For anomalous samples transferred to the test set memory block, the CPU only performs normalization based on… and The offset is used to generate diagnostic error codes and write them to the hard disk log, without performing floating-point value overwrite operations. The test set maintains the original data format of the industrial field.

[0057] The root mean square error obtained from the isolation mechanism test accurately maps the robustness boundary of the prediction engine under the interference of physical weighing errors. Breaking the cognitive black box of data-driven models is the core of feature engineering. Pure ensemble algorithms lack awareness of the physical properties of input parameters. In benchmark tests, the algorithm directly receives a 10-dimensional original feature matrix containing 5 macroscopic admixtures and 5 microscopic oxides. When splitting nodes, the decision tree attempts to find the optimal dividing boundary between steel slag admixture and fly ash admixture in high-dimensional space. The algorithm consumes 50 layers of tree depth to implicitly fit the material law that "high steel slag admixture requires a specific proportion of acidic oxides to activate activity." The excessively deep tree structure caused the model to quickly fall into overfitting on 111 sets of samples. The research team attempted to introduce principal component analysis (PCA) to reduce the dimensionality of the feature space.

[0058] The PCA algorithm generates principal components (such as PC1 and PC2) with the largest variance through orthogonal transformations. However, these principal components, as linear combinations of the original chemical features, lose their interpretive significance in materials science. When the model outputs a prediction of a significant 7-day strength decline, process engineers, faced with the numerically anomalous PC1 variable, are unable to trace the specific raw material source causing the strength decrease. This solution abandons the purely mathematical dimensionality reduction approach and actively constructs derived features characterizing the hydration kinetic mechanism in memory space through algebraic mapping. The release of activity in the steel slag and blast furnace slag composite cementitious system is essentially the deagglomeration of the dense glassy network structure under alkaline conditions.

[0059] silicon-oxygen tetrahedron with aluminum-oxygen tetrahedron The network relies on a high concentration of hydroxide ions to break chemical bonds. The molar ratio of basic to acidic oxides within the system constitutes the threshold condition that dominates the volcanic ash reaction rate. The central processing unit reads measured data from the oxide dimension of the feature matrix and calculates the alkalinity index derivation formula: In the formula This represents a scalar value for alkalinity index. , , , This represents the mass fraction of the four core oxides in the data sample. The system appends this macroscopic hydration characterization parameter as a new independent column vector to the feature matrix. During the root node splitting phase, the decision tree scans the floating-point distribution of this index and calculates the Gini impurity.

[0060] Classification based on alkalinity thresholds reduces the computational cost of blind trial and error in the algorithm. The calcium-to-silica ratio of the micro-hydration products constitutes another characteristic dimension. After mixing and hydration, the core material framework providing early compressive strength in the cementitious material is calcium silicate hydrate (CSH) gel. The interlayer spacing and micro-morphological defects of the CSH gel are controlled by the concentration ratio of free calcium ions to soluble silicate ions in the liquid phase. The central processing unit uses the formula: Calculate the calcium-to-silicon ratio characteristic parameters of the sample. Where... The calcium-silicon ratio, and This represents the proportion of purified oxides. The algorithm captures... As the numerical gradient changes, a linear mapping relationship is established between the system's chemical composition and early compressive strength. After the feature matrix is ​​constructed, some feature dimensions maintain static values ​​across different batches. Retaining input parameters with low information entropy would dilute the model's weight in the feature importance evaluation stage.

[0061] The system initiates a variance threshold filtering instruction before the feature set input algorithm pipeline. The central processing unit calculates the population variance of the numerical distribution in each column of the matrix, using the following formula: In the formula For the first The variance of the column eigenvectors, This represents the number of sample rows in the training set. These are floating-point values ​​in the matrix cells. This is the arithmetic mean of the features in this column. The variance filtering dead zone is set to 0.0005 in this scheme.

[0062] Statistical scanning results show that the variance of magnesium oxide content in this batch of steel slag raw materials is less than 0.0002, and the numerical variation is within the background noise range of the XRF spectrometer. Based on the judgment instruction, the workstation logic control unit extracts the magnesium oxide feature column and its corresponding input channel from the final matrix into memory. The algorithm input retains the core variable combination corresponding to the information density.

[0063] Example 3 like Figure 3As shown, before performing the iterative solution, the intensity prediction model is analyzed using Shapley value theory to extract the nonlinear response intervals of component characteristics to intensity. These nonlinear response intervals are then embedded as dynamic optimization boundaries into the physicochemical constraint set, guiding the sequential least squares programming algorithm to converge by compressing the parameter search space.

[0064] During the experimental testing phase of the mixing plant project, the system revealed a chemical-level optimization flaw. The sequential least-squares programming solver, performing blind gradient descent within the 11-dimensional eigenvector hypercube search space, frequently output mathematical extreme solutions within chemically inert regions. When the algorithm increased the steel slag powder content to 45%, the physical sum of the proportions equaled 1, and the five major oxides fell within the range of [0.9, 1.1]. Excluding the specific alkalinity index activation environment, the 45% steel slag contained a large number of unhydrated 20-50 micrometer inert particles. These particles induced severe internal stress and microcracks in the microscopic interface transition zone of the hardened slurry, causing the actual 7-day compressive strength in the failure test to precipitate from the expected 42 MPa to 28 MPa.

[0065] In conventional cement building materials projects, the control logic is written into the programmable logic controller (PLC) by process engineers, imposing a mandatory experience dead zone, such as limiting the maximum steel slag content to no more than 30%. This static hard boundary disrupts the synergistic excitation effect between oxides within the cementitious system. When the system's alkalinity-derived characteristics (such as...) When the concentration of hydroxide ions increases significantly, the highly active hydroxide ion network is sufficient to broaden the actual activation limit of steel slag to over 38%.

[0066] This embodiment relies on industrial-grade computing nodes equipped with NVIDIA Tesla architecture GPUs. Before the sequential least squares programming solution is triggered, a computational flow is implemented to inversely transform the hidden mappings within the prediction model into dynamic optimization boundaries. The extreme gradient boosting model is composed of hundreds or thousands of decision trees with a depth of up to 6 layers, exhibiting high-dimensional nonlinear black-box characteristics.

[0067] Initially, the Partial Dependency Graph (PDP) algorithm was used to extract the feature response boundaries of univariates. When calculating the marginal effects of features, the underlying mathematical derivation of the PDP algorithm forces the assumption that the features of each component are statistically absolutely independent. The release rate of free calcium oxide in steel slag powder and the deagglomeration of the vitreous network in slag exhibit a strong physicochemical coupling, and the secondary pozzolanic reaction of fly ash is highly dependent on the concentration of calcium hydroxide released during the initial hydration of cement.

[0068] Forcibly separating feature correlations leads to severe distortion of the feature response range output by the PDP. The system sets constraints based on the distortion boundaries provided by the PDP, causing the solver to calculate a singular value in the inverse of the Hessian matrix on the 12th iteration, resulting in a complete deadlock in the optimization process. This solution abandons interpretation tools based on the independence assumption and deploys a game-theoretic Shapley value algorithm on a GPU streaming multiprocessor.

[0069] Shapley's theory views formulation prediction as a cooperative game involving multiple materials. Basic or mechanism-derived features are defined as game participants, and the prediction strength is defined as the total payoff. The algorithm is responsible for quantifying the true marginal contribution of each feature parameter to the prediction strength.

[0070] The fundamental mathematical formula for Shapley value is defined as follows: In the formula It represents the complete set of feature parameters of the input model. This means that the tested feature is not included. Any combination of feature subsets, The total dimension of the features is 11 in this scheme. The subset contains the number of features. The system only accepts subsets When some feature information is masked out and other dimensions are hidden, the expected predicted value of the output compressive strength benchmark is obtained.

[0071] Operation terms Representation will feature The difference in compressive strength caused by the re-injection of a partially mixed system. Adding steel slag first and then desulfurized gypsum, versus adding desulfurized gypsum first and then steel slag, drastically alters the early nucleation rate of ettringite crystals in the system; the difference in strength increment fluctuates non-linearly with the evolution path of the combination. The factorial coefficient term represents the mathematical probability weighting of a specific feature combination order appearing in all possible permutations. The core technical bottleneck in extracting dynamic boundaries lies in the geometrical explosion of computational complexity. The classic Shapley value permutation calculation faces an 11-dimensional composite feature matrix, requiring traversal in a single interpretation. Seed set.

[0072] Extreme gradient boosting models in industrial scenarios typically contain more than 1000 decision trees. Using the aforementioned theoretical formulas for brute-force algebraic solutions, the feature interpretation task for a single recipe requires approximately 72 hours of full-load computation on a dual-CPU system, resulting in frequent L3 cache misses and data page overflows.

[0073] Commercial concrete mixing plants require a time window of no more than 3 minutes from inputting raw material parameters to generating an optimized formula. The computational overhead of the classic SHAP algorithm cannot meet the time constraints of industrial deployment. The computing cluster scheduling kernel refactors the basic operational logic into a tree-structured Shapley value (TreeSHAP) algorithm. TreeSHAP abandons the exhaustive search of the external input feature space and directly delves into the internal topology of the prediction model.

[0074] The GPU concurrently traverses the depth path of the decision tree, no longer repeatedly calculating the predicted values ​​of occluded features. Instead, it dynamically infers the conditional mathematical expectation of the missing feature state from top to bottom within the tree by calculating the proportion of the training set samples covered by each branch node. Originally The exponential time complexity is compressed to hardware-level optimizations. polynomial level complexity ( For the number of trees, The number of leaf nodes. For feature dimension, (Maximum tree depth).

[0075] The algorithm executes hundreds of thousands of floating-point operations concurrently within the GPU's CUDA core cluster, drastically reducing the time required for single-sample contribution analysis from 72 hours to 1.2 seconds. The algorithm decomposes the total predicted resilience value of 111 samples into the sum of the model's baseline expected value and the local contribution values ​​of each feature. The computing cluster extracts the multidimensional Shapley contribution values ​​of the validation set from global GPU memory and constructs a two-dimensional dependency array of feature values ​​and their corresponding strength contribution values ​​within the logical address space.

[0076] The GPU arithmetic unit executes a calculus boundary scan algorithm to extract the nonlinear response range of component features to intensity. Taking the characteristic parameters of steel slag powder as an example, when the system scans the mass fraction gradually increasing from 0% to 18%, the nucleation effect in the early stage of hydration reaction dominates the microscopic evolution, and the Shapley contribution value increases positively and nonlinearly with the increase of admixture amount.

[0077] When the doping concentration exceeds the hidden saturation critical value of 28%, the high concentration of hydroxide ions providing the excitation environment within the system is depleted. Excess inert steel slag particles hinder the network bonding of hydrated calcium silicate gel, causing the Shapley contribution value to reverse negatively in the two-dimensional array, exhibiting an exponential drop. The core control formula for the system to perform nonlinear response interval scanning is defined as: In the formula This represents the set of optimal activity characteristic regions after purification. For continuously scanned component mass fraction characteristic scalars, The corresponding contribution value to the Shapley intensity. The preset critical value for positive contribution is set at 0.5 MPa in this scheme, eliminating ineffective doping ranges that offer no substantial strength gain to the macrostructure. (Calculus term) The first-order partial derivative of the contribution curve represents the instantaneous gradient of the contribution degree as a function of doping concentration. The upper limit of the system's tolerable contribution decay rate is set to 0.15 MPa / %. When the negative absolute value of the first-order partial derivative exceeds 0.15, the material activity enters the collapse stage. The composite algebra formula prevents the algorithm from including feature coordinates in the descent channel within the safe range. The calculus scanning algorithm locates the lower boundary coordinate values ​​that satisfy the composite condition along the feature coordinate axes. With upper boundary coordinate values These two floating-point coordinates, accurate to four decimal places, constitute the dynamic optimization boundary of the raw materials under the current chemical environment.

[0078] The GPU serializes the dynamic optimization boundaries of each dimension of features and writes them to the cache of the constraint optimization module via shared memory, feeding back into the sequential least squares programming solver engine in Example 1. Under the static constraint system, the upper and lower limits of the search for a single raw material are broadly set as the closed interval [0,1]. In this example architecture, the objective function of the sequential least squares programming algorithm maintains the minimization of the absolute difference of predictions, and the system covers and embeds a dynamically generated set of optimal boundaries in the underlying physicochemical constraint set. The solver's active set policy engine receives the dynamic boundary data stream and reallocates the Lagrange multiplier array.

[0079] The mathematical constraint formula for embedding dynamic optimization boundaries is defined as follows: as well as In the formula and For the real-time parsing of the first The lower and upper bounds of the feature interval are floating-point values. When the algorithm approximates the Hessian matrix in the multidimensional matrix space and plans the gradient descent direction, if the calculated exploration step size makes the feature variables... The value is lower than or higher Dead zone drift causes the arithmetic unit to trigger a high penalty function value in the approximate cost function. Algebraic penalties generate correction instructions, forcing the iterative cursor to change direction. The parameter search space is compressed from the macroscopic full range of [0,1] to the interior of a microscopically chemically active manifold.

[0080] Geometric dimensionality reduction of the search space reduces the number of iterations required for the quasi-Newton algorithm to approximate the extreme point, reducing the computation time for a single formula generation by 70%. The possibility of the black-box model fabricating low-reactivity mathematical proportions in the local space is eliminated by the underlying mathematical constraints. When the hardware timer determines that the gradient iteration error has fallen below the convergence dead zone after two consecutive iterations, the computation data stream is stopped, and the serial interface sends the ingredient feature vector that satisfies both macroscopic mass conservation and microscopic hydration activity to the PLC.

[0081] In the description of this specification, the references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0082] The above description is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in the claims, they should all fall within the protection scope of the present invention.

Claims

1. A method for predicting the strength of steel slag-based cementitious materials, comprising acquiring mix proportion data of the steel slag-based cementitious materials and establishing a strength prediction model using a machine learning algorithm, characterized in that, The method further includes: An objective function is established, wherein the objective function aims to minimize the absolute value of the difference between the predicted value of the intensity prediction model and the preset target intensity value. Construct a physicochemical constraint set, which includes constraints on the sum of component mass fractions and total oxide content; The objective function is solved iteratively using a sequential least squares programming algorithm. During the solution process, the physicochemical constraint set is used to restrict the iteration points, and the optimal mix ratio data is output.

2. The method of strength prediction of steel slag-based cementitious material according to claim 1, characterized in that, The constraint on the sum of the mass fractions of the components in the physicochemical constraint set is that the deviation of the sum of the mass fractions of steel slag powder, slag powder, desulfurized gypsum, cement and fly ash from the preset value is within the preset tolerance range; the constraint on the total oxide content is that the sum of the contents of calcium oxide, silicon dioxide, ferric oxide, aluminum oxide and magnesium oxide in the mixed system is within the preset range.

3. The method of claim 1, wherein the steel slag-based cementitious material is a steel slag-based cement. 2 The input features of the strength prediction model include derived features constructed based on the hydration mechanism of cementitious materials, including alkalinity index and calcium-silicon ratio.

4. The method of strength prediction of steel slag-based cementitious materials according to claim 3, characterized in that, The alkalinity index is obtained by dividing the sum of the contents of calcium oxide and magnesium oxide in the mixed system by the sum of the contents of silicon dioxide and aluminum oxide; the calcium-silicon ratio is obtained by dividing the calcium oxide content in the mixed system by the silicon dioxide content.

5. The method of claim 1, wherein the steel slag-based cementitious material is a steel slag-based cement. Before establishing the intensity prediction model using machine learning algorithms, the physical consistency of the mix proportion data is also checked. When the sum of the component mass fractions or the total oxide content of the data sample deviates from the preset range, a normalization correction strategy is used to process the feature values ​​of the data sample.

6. The method of strength prediction of steel slag-based cementitious materials according to claim 5, characterized in that, After the dataset is partitioned, the normalization correction strategy is applied to the abnormal data samples in the training set, while only diagnostic records are processed for the abnormal data samples in the test set without any correction.

7. The method for predicting the strength of steel slag-based cementitious materials according to claim 1, characterized in that, Before performing the iterative solution, the intensity prediction model is analyzed using Shapley value theory to extract the nonlinear response range of component characteristics to intensity.

8. The method for predicting the strength of steel slag-based cementitious materials according to claim 7, characterized in that, The nonlinear response interval is embedded as a dynamic optimization boundary into the physicochemical constraint set, and the sequential least squares programming algorithm is guided to converge by compressing the parameter search space.

9. A strength prediction system for steel slag-based cementitious materials, characterized in that, include: The data preprocessing module is used to obtain the mix proportion data of steel slag-based cementitious materials and perform physical consistency verification. The model building module is used to build intensity prediction models using machine learning algorithms; The constraint optimization module is used to construct the objective function and perform sequential least squares programming in combination with the physicochemical constraint set to output the optimal mix ratio data.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for predicting concrete strength based on hybrid model

    CN104991051A