Cement-based grouting material proportion optimization method and system based on deep reinforcement learning
Through a deep reinforcement learning method, combined with neural network and reinforcement learning, the problem of inefficient design of traditional grouting materials is solved, and accurate prediction and optimized design of grouting materials performance are achieved, which improves the scientificity and reliability of the design.
Patent Information
- Application Number
- CN202510127206.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-01
AI Technical Summary
The research and design of traditional grouting materials rely on frequent formula adjustments and parameter optimization under laboratory conditions, resulting in complex design processes, inefficient efficiency, and safety risks. At the same time, existing machine learning methods have shortcomings in algorithm interpretability and material proportion optimization design.
A method based on deep reinforcement learning is adopted to combine neural networks with reinforcement learning to establish a unified material performance prediction and proportion optimization design framework. By learning the nonlinear mapping relationship between material ratio and performance, accurate prediction of grouting material performance is achieved, and an ε-greedy strategy is introduced to iteratively adjust the initial material ratio.
It realizes high-precision prediction of grouting material performance, reduces errors in human intervention and subjective judgment, improves the scientificity and reliability of material research and development design, broadens the search space, and improves the accuracy and efficiency of grouting material optimization design.
Smart Images

Figure CN120072135A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of cement-based material optimization, and specifically to a method and system for optimizing the proportion of cement-based grouting materials based on deep reinforcement learning. Background Art
[0002] The statements in this section merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] As an effective underground engineering measure, grouting is widely used in the treatment of water inrush during tunnel construction. The main function of grouting is to block the flow of groundwater and improve the stability of the tunnel face and its surrounding strata. The proportion of the grouting material is the key and direct factor determining its material properties and greatly affects the grouting treatment effect. Therefore, in the treatment of water inrush, it is particularly important to reasonably design the proportion of the grouting material. Traditional research and design of grouting materials highly rely on frequent formula adjustments and parameter optimizations under laboratory conditions. This method is not only time-consuming and laborious, making the entire design process complex and inefficient, but also may pose safety risks when operating toxic materials.
[0004] With the rapid development of information technology, machine learning technology has gradually become a new tool for the research and development of grouting materials. Algorithms such as support vector machines, random forests, and neural networks have been used for performance prediction, structural analysis, and proportion optimization design of grouting materials. Using these algorithms, the internal relationship between material properties and formulas can be mined from existing data to achieve intelligent prediction of material properties and proportion optimization. However, although existing machine learning methods have promoted the intelligent process of grouting material research to a certain extent, there are still obvious deficiencies in the interpretability of the algorithms. Especially when exploring how different material proportions specifically affect grouting performance, the lack of an intuitive explanation mechanism and visualization tool has become a bottleneck restricting further development. For the optimization design of material proportions, the performance-oriented reverse design method faces higher technical challenges and relies on more complex customized algorithms, such as autoencoders and generative adversarial networks, to generate a reasonable crystal structure of a specific material system from a microscopic perspective.
[0005] Currently, the methods for predicting and analyzing the performance of grouting materials mainly include single-factor experiments, response surface analysis, polynomial regression, and machine learning algorithms. On this basis, combined with multi-objective optimization strategies and multi-index comprehensive evaluation techniques such as the entropy weight method, the optimization or optimal design of the proportion of grouting materials can be achieved. However, performance prediction and proportion optimization design usually require separately constructing different algorithm models, resulting in a long and complex overall process. Summary of the Invention
[0006] To solve the above problems, the present disclosure proposes a method and system for optimizing the mix ratio of cement-based grouting materials based on deep reinforcement learning, which combines the neural network backpropagation mechanism with reinforcement learning to establish a unified framework for material property prediction and mix ratio optimization design. The neural network learns the non-linear mapping relationship between the material mix ratio and properties to accurately predict the properties of the grouting material. The ε-greedy strategy of reinforcement learning is introduced to iteratively adjust the initial material mix ratio. By setting reasonable reward functions and action selection strategies, the optimization process can be dynamically updated according to the prediction error and gradually approach the grouting performance required by the design.
[0007] According to some embodiments, the present disclosure adopts the following technical solutions:
[0008] A method for optimizing the mix ratio of cement-based grouting materials based on deep reinforcement learning, comprising:
[0009] Obtaining the original grouting material mix ratio data;
[0010] Constructing a deep reinforcement learning network, inputting the original grouting material mix ratio data as the initial input features into the deep reinforcement learning network, and using the prediction module to output the initial predicted grouting parameters;
[0011] Introducing a Q-Learning reinforcement update module into the deep reinforcement learning network, given the optimal target output of the grouting parameters, designing a dynamic environment adaptation update strategy in the Q-Learning reinforcement update module, realizing the reverse optimization of the input features through the feedback mechanism iteration adjustment, selecting the input feature update mode and update direction in the fine-tuning actions of the input features, inputting the updated features into the prediction module, and dynamically iteratively adjusting according to the difference between the prediction output of the prediction module and the target output, making the prediction output of the prediction module gradually approach the target output during the iteration, and realizing the optimization of the initial input features, that is, obtaining the optimal grouting material mix ratio input.
[0012] According to some embodiments, the present disclosure adopts the following technical solutions:
[0013] A system for optimizing the mix ratio of cement-based grouting materials based on deep reinforcement learning, comprising:
[0014] A data acquisition module for obtaining the original grouting material mix ratio data;
[0015] A reinforcement learning reverse optimization module for constructing a deep reinforcement learning network, inputting the original grouting material mix ratio data as the initial input features into the deep reinforcement learning network, and using the prediction module to output the initial predicted grouting parameters;
[0016] Introduce a Q-Learning reinforcement update module into the deep reinforcement learning network. Given the optimal target output of grouting parameters, design a dynamic environment adaptation update strategy in the Q-Learning reinforcement update module. Through the feedback mechanism, iteratively adjust to achieve the reverse optimization of input features. In the fine-tuning actions of input features, select the input feature update mode and update direction. Input the updated features into the prediction module, and dynamically iteratively adjust according to the difference between the prediction output of the prediction module and the target output. During the iteration, make the prediction output of the prediction module gradually approach the target output, realizing the optimization of the initial input features, that is, obtaining the optimal input of grouting material ratio.
[0017] According to some embodiments, the present disclosure adopts the following technical solutions:
[0018] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method for optimizing the ratio of cement-based grouting materials based on deep reinforcement learning as described above.
[0019] According to some embodiments, the present disclosure adopts the following technical solutions:
[0020] A non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, it implements the method for optimizing the ratio of cement-based grouting materials based on deep reinforcement learning as described above.
[0021] According to some embodiments, the present disclosure adopts the following technical solutions:
[0022] An electronic device includes: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the method for optimizing the ratio of cement-based grouting materials based on deep reinforcement learning as described above.
[0023] Compared with the prior art, the beneficial effects of the present disclosure are:
[0024] The method for optimizing the ratio of cement-based grouting materials based on deep reinforcement learning of the present disclosure starts from the perspective of data-driven, uses machine learning methods to extract valuable features from a large amount of historical data, establishes a CNN neural network model, and realizes high-precision grouting performance prediction through effective hyperparameter tuning, reducing the errors caused by human intervention and subjective judgment in the process of grouting material research and development and performance testing, and improving the scientificity and reliability of material research and development design.
[0025] The disclosed method for optimizing the mix proportion of cement-based grouting materials based on deep reinforcement learning calculates the gradient of the loss function with respect to the input features using the backpropagation mechanism of the neural network in the exploitation mode, thereby achieving local gradient optimization. In the exploration mode, it explores new regions with the reinforcement learning ε-greedy strategy. The two modes are used randomly, which can not only achieve local gradient optimization to ensure that each optimization moves towards a better mix proportion direction, but also explore different update paths of input features during the optimization process, thus jumping out of local traps and enhancing the convergence speed. This method greatly broadens the search space and is more flexible than traditional optimization methods due to relatively fixed and limited optimization strategies, improving the accuracy of the grouting material optimization design and meeting diverse engineering requirements with higher efficiency and better effects.
[0026] The disclosed method for optimizing the mix proportion of cement-based grouting materials based on deep reinforcement learning realizes the automation and intelligence of performance prediction and material optimization design on the basis of the traditional grouting material design process. Based on the joint optimization of reinforcement learning and neural network, it can not only effectively shorten the optimization time, but also adapt to the multi-objective optimization of different engineering requirements, improve the efficiency and applicability of the grouting material mix proportion design, provide a reliable material selection scheme for the fissured rock mass grouting project, and has important theoretical and practical values. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the disclosure. The illustrative embodiments and descriptions thereof of the disclosure are used to explain the disclosure and do not constitute an improper limitation of the disclosure.
[0028] Figure 1 is the training flow chart of the deep reinforcement learning network according to the embodiment of the disclosure;
[0029] Figure 2 is the flow chart of iteratively optimizing the input features according to the embodiment of the disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0031] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the disclosure belongs.
[0032] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0033] Embodiment 1
[0034] In an embodiment of the present disclosure, a method for optimizing the proportion of cement-based grouting materials based on deep reinforcement learning is provided, including the following steps:
[0035] Step 1: Obtain the original grouting material proportion data;
[0036] Step 2: Construct a deep reinforcement learning network, input the original grouting material proportion data as the initial input features into the deep reinforcement learning network, and use the prediction module to output the initial predicted grouting parameters;
[0037] Introduce a Q-Learning reinforcement update module into the deep reinforcement learning network. Given the optimal target output of the grouting parameters, design a dynamic environment adaptation update strategy in the Q-Learning reinforcement update module. Through the feedback mechanism, iteratively adjust to achieve the reverse optimization of the input features. Select the input feature update mode and update direction in the fine-tuning action of the input features. Input the updated features into the prediction module, and dynamically iteratively adjust according to the difference between the predicted output of the prediction module and the target output. In the iteration, make the predicted output of the prediction module gradually approach the target output, realize the optimization of the initial input features, that is, obtain the optimal grouting material proportion input.
[0038] As an embodiment, the method for optimizing the proportion of cement-based grouting materials based on deep reinforcement learning in the present disclosure utilizes the learning ability of the neural network to establish a unified framework for material performance prediction and proportion optimization design. Specifically, the machine learning neural network realizes the accurate prediction of the performance of grouting materials by training and learning the non-linear mapping relationship between the material proportion and performance. On the basis of the trained model, introduce the ε-greedy strategy of reinforcement learning, iteratively adjust the initial material proportion, and through setting reasonable reward functions and action selection strategies, make the optimization process dynamically update according to the prediction error and gradually approach the grouting performance required by the design. The specific implementation process is as follows:
[0039] Step 1: Obtain the original data of the grouting material proportion and grouting performance parameters, and preprocess and divide the original data to obtain the required data set;
[0040] Specifically, 1) Use indoor tests to obtain the original data of the grouting material ratio and grouting parameters. The original data includes the grouting parameters characterizing the grouting performance, such as the slurry strength, fluidity, and setting time under different material ratios and their conditions.
[0041] 2) Preprocess the original data, including abnormal data cleaning, data interpolation and supplementation, and feature selection.
[0042] Specifically, abnormal data can be found by establishing existing data variation indicators, such as range, standard deviation, coefficient of variation, etc., and cleaning is completed by deleting abnormal values and mean correction.
[0043] For data interpolation and supplementation, existing methods such as linear interpolation and Lagrange interpolation can be selected to ensure that the interpolated data can maintain the overall trend and structure of the data.
[0044] Furthermore, for feature selection of the cement-based material ratio data in the original data, selectable methods include principal component analysis, linear discriminant analysis, etc., to retain the key input features that are most useful for model prediction.
[0045] Furthermore, the preprocessed original data should be divided into three parts: training set, validation set, and test set. The training set is used for model training, optionally accounting for 70 - 80% of the total data set. The validation set is used for hyperparameter tuning and model selection of the model, optionally accounting for 10 - 15% of the total data set. The test set is used to evaluate the final performance of the model, optionally accounting for 10 - 15% of the total data set.
[0046] Step 2: Establish a construction deep reinforcement learning network for optimizing the grouting material ratio, and use the training set for model training.
[0047] 1) The deep reinforcement learning network includes a prediction module. The prediction module uses a machine learning model to predict grouting parameters, trains the machine learning model using the training set, and adjusts the hyperparameters with the help of the validation set.
[0048] Specifically, the establishment of the machine learning model includes selecting a suitable network framework, configuring the network framework, defining a loss function, and selecting an optimizer.
[0049] First, select a suitable network framework. Convolutional neural network (CNN), recurrent neural network (RNN), multi-layer perceptron (MLP), etc. can be selected. As a preferred option, select the convolutional neural network (CNN).
[0050] Further configure each layer of the network framework. The configuration of each layer of the network framework includes setting the number of input layer nodes, selecting an appropriate number of layers and the number of neurons in each hidden layer, and setting the output layer.
[0051] Among them, the number of nodes in the input layer is the same as the number of input features. Assuming that the number of input features retained after feature engineering is n, the dimension of the input layer is (n, 1); the hidden layer includes a convolutional layer, a pooling layer, and a fully connected layer. Here, 2 convolutional layers are set. Convolutional layer 1 uses 16 filters, the kernel size is 1×3, the activation layer uses the ReLU function, convolutional layer 2 uses 32 filters, the kernel size is 1×3, and the activation layer uses the ReLU function; a max pooling layer is set after each convolutional layer, and the window size is 1×2; finally, a flattening layer and a fully connected layer are set. The fully connected layer contains 64 neuron nodes and is activated using the ReLU function; the output layer requires 3 neuron nodes to match the form of the prediction target.
[0052] Define the loss function and select the optimizer. The loss function can be the mean squared error (MSE) or the mean absolute error (MAE). The calculation formula is:
[0053]
[0054]
[0055] where n is the number of training samples, y i is the true output corresponding to the training data x i , and is the predicted value corresponding to the training data x i ; the optimizer can be the Adam algorithm, stochastic gradient descent, etc.
[0056] Preferably, the loss function is selected as the mean squared error (MSE), and the optimizer is selected as the Adam algorithm.
[0057] The training process of the model includes forward propagation, loss calculation, backpropagation, algorithm optimization, and setting appropriate training epochs and batch sizes. The specific process is as follows:
[0058] Forward propagate the initial input features through the prediction module to output the predicted grouting parameters;
[0059] Set the loss function to calculate the loss value between the current output and the true target; calculate the gradient of the loss function with respect to the model weights and update the weights through the optimizer during backpropagation;
[0060] Input the validation dataset to tune the hyperparameters of the model. Optional hyperparameter tuning methods include grid search, random search, Bayesian optimization, etc.;
[0061] Input the test set to test and evaluate the model, which should include metric evaluation or cross-validation. Appropriate evaluation metrics to select include accuracy, precision, recall, F1-score, etc.
[0062] 2) The deep reinforcement learning network also includes a Q-Learning reinforcement update module. The Q-Learning reinforcement update module is set behind the machine learning prediction model to establish a dynamic environment adaptation-based reinforcement learning model, initialize the model parameters, and set the reinforcement learning dynamic environment.
[0063] Improve and optimize the Q-Learning method in reinforcement learning, design a dynamic environment adaptation update strategy, and iteratively adjust through a feedback mechanism to achieve reverse optimization of the neural network input features of the prediction module. This method can be regarded as constructing an optimization process on the basis of a trained neural network, selecting the update mode and update direction in the fine-tuning action of the input features, inputting the updated features into the prediction module, and adjusting the update direction according to the difference between the neural network prediction output and the target output of the prediction module, so that the neural network prediction output gradually approaches the target output during iteration.
[0064] Specifically, the setting of the reinforcement learning dynamic environment includes the reward function, the Q-function update strategy, and the action set setting.
[0065] Among them, 1) The reward function is used to evaluate the effect of each operation. Based on the difference between the actual output result and the target output result, the negative mean square error is used as the reward. Then the reward R obtained in the t-th round of loop t The calculation formula is:
[0066] R t = -(y pred - y target ) 2
[0067] Among them, y pred is the neural network prediction output, and y target is the target output.
[0068] 2) The Q-function update strategy is designed to be dynamic environment adaptation-based. The Q-function is a function used in reinforcement learning to evaluate the expected effect of optimization operations. In the Q-Learning algorithm, the Q-function is gradually updated with iteration, and the influence of new information on old information is controlled by the learning rate α. Its update formula is:
[0069]
[0070] Among them, x t is the input feature in the t-th round of loop; Δx t is the update direction of the input x t in the t-th round of loop; Q(x t , Δx t ) is the Q-value in the t-th round of loop; αt is the learning rate, used to control the degree of iterative update; R tis the reward obtained in the t-th round of the loop, also known as the immediate reward; is the maximum Q value that can be obtained from the fine-tuning actions in the action set for the (t + 1)-th round of the loop, also known as the future reward; γ t is the discount factor, which is used to balance the immediate reward and the future reward, and its value range is [0, 1].
[0071] Specifically, during the training process, the learning rate α and the discount factor γ in the Q value are dynamically adjusted according to the changes in the state and action space. For the learning rate α, if the reward does not increase significantly after the Q value is updated within a certain period, the learning rate can be increased. Conversely, if the reward gradually increases, the learning rate is decreased, and the adjustment formula is:
[0072]
[0073] where λ α is a hyperparameter that controls the adjustment speed of the learning rate, R t is the reward obtained in the t-th round of the loop, and μ and σ are the average and standard deviation of the rewards within a certain fixed time window,
[0074]
[0075] where N is the window size, which can be set to 100. For the discount factor γ, by observing the prediction accuracy of the model at different time steps, when the prediction of the long-term reward by the model is stable, the value of γ is gradually increased to pay more attention to the future reward, and the adjustment formula is:
[0076]
[0077] where λ γ is a hyperparameter that controls the increasing rate of the γ value, ΔQ is the change in the Q value, ΔQ = Q(x t+1 , Δx t+1 ) - Q(x t , Δx t ), and Q max is the current maximum Q value.
[0078] 3) Action set setting. The action is a small adjustment to the input features. All the fine-tuning combinations of the input features are combined into an action set. The specific setting process is as follows: For each input feature, fine-tuning actions with three fixed adjustment amplitudes of increasing by 5%, decreasing by 5%, and remaining unchanged are set. Then, for n input features, the action set will include 3 n types of actions.
[0079] Furthermore, the model parameter initialization includes initializing the Q function, the initial input features, and the target output.
[0080] Specifically, the Q function can be initially set to 0 or a random value. The initial input x0 , that is, the initial grouting material ratio, can be selected based on historical data or prior knowledge. The target output y target is the specific value of the target slurry performance parameters of compressive strength, viscosity, and setting time, which is selected based on the grouting design requirements.
[0081] Step 3: Achieve the reverse optimization of the input features through iterative adjustment of the feedback mechanism. Select the input feature update mode and update direction in the fine-tuning action of the input features, input the updated features into the prediction module, and perform dynamic iterative adjustment according to the difference between the prediction output of the prediction module and the target output. During the iteration, make the prediction output of the prediction module gradually approach the target output, and achieve the optimization of the initial input features, that is, obtain the optimal grouting material ratio input.
[0082] Select the input feature update mode and update direction in the fine-tuning action of the input features, including: randomly select the backpropagation algorithm of the neural network model or the reinforcement learning exploration strategy to iteratively update the initial input value. The backpropagation algorithm is used for local search optimization of the input features, and the reinforcement learning exploration strategy is used for global search, and finally obtain the optimized grouting material ratio based on the target grouting parameters.
[0083] Specifically, the iterative update includes 3 steps: forward propagation, selecting the input feature update mode, and judging the iteration process.
[0084] Step 1) Forward propagation, input the initial input feature x 0 into the trained machine learning model (prediction module) for prediction to obtain the current output prediction y'. pred ;
[0085] Step 2) Select the input feature update mode, randomly select backpropagation update or use the reinforcement learning ε-greedy policy exploration mechanism to guide the direction of the new input features. Backpropagation can provide a clear gradient direction to accelerate optimization convergence, while the exploration of the input feature direction by reinforcement learning can help the model find possible good directions in the early stage. This combined mode can explore new input directions with a certain probability, rather than completely relying on the gradient direction, to help jump out of the local optimum and make the prediction output obtained by the input features through neural network calculation approach the target output.
[0086] Specifically, first, as Figure 2 shown, set an appropriate exploration rate, with a value range of [0, 1]. At the same time, randomly generate a number r, with a value range of [0, 1], and compare it with the exploration rate ε. If r < ε, enter the exploration mode; if r > ε, then enter the exploitation mode.
[0087] Among them, the exploration mode is to select the update direction Δx according to the ε-greedy policy t , and the calculation formula is:
[0088]
[0089] Among them, x t is the input feature in the t-th round of loop; random action represents randomly selecting a fine-tuning action from the action set as the update direction; argmax Δx Q(x t , Δx) represents selecting the fine-tuning action that maximizes the Q value from the action set as the update direction; ε is the exploration rate.
[0090] Next, adjust and update the input feature according to the selected update direction. The formula is:
[0091] x t+1 = x t + Δx t
[0092] Then, input the updated input feature x t+1 into the prediction module to obtain a new prediction output y pre d .
[0093] Furthermore, use the method of updating the input feature by means of the backpropagation mechanism. Calculate the loss L between the current output prediction and the target output. The calculation formula is:
[0094]
[0095] Furthermore, calculate the gradient of the loss L with respect to the input feature x. The calculation formula is:
[0096]
[0097] Among them, x is the input feature, is the gradient of the loss L with respect to the prediction output y′ pred , is the gradient of the prediction output y′ pred with respect to the input feature x.
[0098] Use the adaptive learning rate Adam to perform backpropagation update on the input feature to obtain the input feature x grad in the gradient direction. The update formula is:
[0099]
[0100] Among them, η is the learning rate.
[0101] Input the updated input feature x grad into the neural network to obtain a new prediction output y pred .
[0102] Step 3) Determine the iteration process. After updating the input features and predicted output through the exploration mode or exploitation mode, calculate the current predicted output y pred and the target output y target The loss between them, and the loss function L is:
[0103]
[0104] Check whether the loss function L and the number of iterations t reach the threshold. If they have reached, stop the iteration. If not, the reward value R needs to be calculated t , and use the update strategy and update function to optimize the Q function. Iteratively repeat the steps of forward propagation, selecting the input feature update mode, and determining the iteration process, continuously switch between exploration and exploitation, update the input features until the loss function converges to the preset requirements or the number of iterations reaches the threshold, and finally output the updated input features, that is, optimize the grouting material ratio.
[0105] Example 2
[0106] In one embodiment of the present disclosure, a system for optimizing the ratio of cement-based grouting materials based on deep reinforcement learning is provided, including:
[0107] A data acquisition module for acquiring the original grouting material ratio data;
[0108] A reinforcement learning reverse optimization module for constructing a deep reinforcement learning network, inputting the original grouting material ratio data as the initial input features into the deep reinforcement learning network, and using the prediction module to output the initial predicted grouting parameters;
[0109] Introduce a Q-Learning reinforcement update module into the deep reinforcement learning network. Given the optimal target output of the grouting parameters, design a dynamic environment adaptation update strategy in the Q-Learning reinforcement update module, and realize the reverse optimization of the input features through iterative adjustment by the feedback mechanism. Select the input feature update mode and update direction in the fine-tuning action of the input features, input the updated features into the prediction module, and perform dynamic iterative adjustment according to the difference between the predicted output of the prediction module and the target output. In the iteration, make the predicted output of the prediction module gradually approach the target output, and realize the optimization of the initial input features, that is, obtain the optimal input of the grouting material ratio.
[0110] Example 3
[0111] In one embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the method for optimizing the ratio of cement-based grouting materials based on deep reinforcement learning.
[0112] Example 4
[0113] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the method for optimizing the proportion of cement-based grouting materials based on deep reinforcement learning is implemented.
[0114] Example 5
[0115] In one embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the method for optimizing the proportion of cement-based grouting materials based on deep reinforcement learning.
[0116] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate computer-implemented processing, and thus the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0118] Although the specific implementation manners of the present disclosure have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.
Claims
1. A cement-based grouting material ratio optimization method based on deep reinforcement learning, characterized in that: include: Obtain original grouting material ratio data; Construct a deep reinforcement learning network, input the original grouting material ratio data into the deep reinforcement learning network as the initial input feature, and use the prediction module output to obtain the initial predicted output grouting parameters; The Q-Learning reinforcement update module is introduced into the deep reinforcement learning network. Given the optimal grouting parameter target output, a dynamic environment adaptation update strategy is designed in the Q-Learning reinforcement update module. The reverse optimization of the input features is achieved through iterative adjustment of the feedback mechanism. The input feature update mode and update direction are selected in the fine-tuning action of the input features. The updated features are input into the prediction module, and dynamic iterative adjustment is performed according to the difference between the predicted output of the prediction module and the target output. During the iteration, the predicted output of the prediction module is gradually close to the target output, thereby optimizing the initial input features, that is, obtaining the optimal grouting material ratio input.
2. The cement-based grouting material ratio optimization method based on deep reinforcement learning according to claim 1, characterized in that: The output grouting parameters include slurry strength, fluidity and setting time. After obtaining the original grouting material ratio data, it is preprocessed, including abnormal data cleaning, data interpolation and feature selection, to retain the most critical input features for predicting the output.
3. The cement-based grouting material ratio optimization method based on deep reinforcement learning according to claim 1, characterized in that: The prediction module of the deep reinforcement learning network is a machine learning network model. Each layer of the network framework of the machine learning network model is configured. The configuration of each layer of the network framework includes setting the number of input layer nodes, selecting an appropriate number of layers and the number of neurons in each layer in the hidden layer, and setting the output layer. The number of input layer nodes is the same as the number of input features. If the number of input features is n, the input layer dimension is (n, 1). The machine learning network model forward propagates the initial input features through the model and outputs the predicted grouting parameters.
4. The cement-based grouting material ratio optimization method based on deep reinforcement learning according to claim 1, characterized in that: The Q-Learning reinforcement update module is a reinforcement learning dynamic environment. The setting of the reinforcement learning dynamic environment includes the reward function, the Q function update strategy and the action set setting. The reward function evaluates the effect of each operation. Based on the difference between the output result and the target output result, the negative mean square error is used as the reward. The reward R obtained in the tth round of the cycle is t The calculation formula is: R t =-(y pred -y target ) 2 Among them, y pred is the neural network prediction output, y target The target output.
5. The cement-based grouting material ratio optimization method based on deep reinforcement learning according to claim 1, characterized in that: The Q-Learning reinforcement update module uses the Q function update strategy to evaluate the expected effect of the optimization operation, introduces the learning rate and discount factor, and the Q function is gradually updated with the iteration. The learning rate controls the influence of new information on old information, and dynamically adjusts the learning rate and discount factor in the Q value according to the changes in the state and action space. For the learning rate, if the reward does not increase significantly after the Q value is updated within a certain period, the learning rate is increased. On the contrary, if the reward gradually increases, the learning rate is reduced.
6. The cement-based grouting material ratio optimization method based on deep reinforcement learning according to claim 1, characterized in that: In the fine-tuning action of the input features, the input feature update mode and update direction are selected, including: randomly selecting back-propagation update or using reinforcement learning ε-greedy strategy exploration mechanism to guide the new input feature direction, iteratively updating the initial input value, the back-propagation algorithm is used for local search optimization of the input features, and the reinforcement learning exploration strategy is used for global search, so that the predicted output obtained by the input feature calculated by the neural network approaches the target output, and finally the optimized grouting material ratio based on the target grouting parameters is obtained.
7. A cement-based grouting material ratio optimization system based on deep reinforcement learning, characterized in that: include: A data acquisition module is used to obtain original grouting material ratio data; The reinforcement learning reverse optimization module is used to construct a deep reinforcement learning network, input the original grouting material ratio data into the deep reinforcement learning network as the initial input feature, and use the prediction module output to obtain the initial predicted output grouting parameters; The Q-Learning reinforcement update module is introduced into the deep reinforcement learning network. Given the optimal grouting parameter target output, a dynamic environment adaptation update strategy is designed in the Q-Learning reinforcement update module. The reverse optimization of the input features is achieved through iterative adjustment of the feedback mechanism. The input feature update mode and update direction are selected in the fine-tuning action of the input features. The updated features are input into the prediction module, and dynamic iterative adjustment is performed according to the difference between the predicted output of the prediction module and the target output. During the iteration, the predicted output of the prediction module is gradually close to the target output, thereby optimizing the initial input features, that is, obtaining the optimal grouting material ratio input.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for optimizing the proportion of cement-based grouting materials based on deep reinforcement learning described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the cement-based grouting material ratio optimization method based on deep reinforcement learning as described in any one of claims 1-6 is implemented.
10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes the cement-based grouting material ratio optimization method based on deep reinforcement learning as described in any one of claims 1-6.
Citation Information
Patent Citations
Q learning-based deep neural network adaptive back-off strategy implementation method and system
CN111867139A
Method for predicting house prices by mining space-time relevance of house prices based on neural network
CN112418939A
Working condition monitoring and control model building method and application method and device thereof
CN114519291A
Deep Q learning bearing fault diagnosis method based on Bayesian optimization
CN117171508A
Manufacturing process technological parameter optimization method and system based on reinforcement learning
CN117420800A
Cited By
Friction plate formula design method and system based on deep reinforcement learning
CN121093761A
Grouting construction scheme intelligent decision-making method and system based on deep learning
CN121145745A