H-shaped steel rolling uncertainty self-adaptive dynamic optimization method and system

By constructing a probabilistic digital twin and performing adversarial domain randomized training, combined with recursive Bayesian filtering and adaptive policy optimization, the uncertainty and adaptability issues in the H-beam rolling process were resolved, achieving efficient and reliable rolling control and improving rolling efficiency and product quality.

CN121857296APending Publication Date: 2026-04-14MCC (SHANGHAI) STEEL STRUCTURE TECHNOLOGY CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MCC (SHANGHAI) STEEL STRUCTURE TECHNOLOGY CORP LTD
Filing Date
2025-12-22
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing H-beam billet rolling control methods are difficult to adapt to the uncertainties in the rolling process, resulting in control decisions relying on incomplete model information and model-reality discrepancies. Furthermore, existing digital twin and reinforcement learning technologies are insufficient in terms of robustness and adaptability, making it difficult to achieve high-precision and high-reliability rolling production.

Method used

A probabilistic digital twin containing structured uncertainty representations is constructed. A robust reinforcement learning control strategy is trained using a domain randomization method based on adversarial perturbation. The strategy is then adjusted online through a recursive Bayesian filtering and adaptive policy optimization framework. Combined with real-time data acquisition from a multimodal sensor network, this enables comprehensive perception and dynamic optimization of the complex rolling process.

Benefits of technology

It significantly improves rolling efficiency and product quality, reduces performance degradation, and enables rapid adaptation to special working conditions such as new steel grades and specifications, meeting the flexibility and adaptability requirements of modern industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121857296A_ABST
    Figure CN121857296A_ABST
Patent Text Reader

Abstract

The invention discloses an H-shaped steel rolling uncertainty self-adaptive dynamic optimization method. The method comprises the following steps: constructing probability digital twin bodies containing structured uncertainty characterization in an H-shaped steel rolling process; a domain randomization method based on confrontation disturbance is adopted, and a reinforcement learning robust control strategy is trained in the probabilistic digital twinborn body; in the actual rolling process of the H-shaped steel, real-time sensing data are obtained, digital twinborn parameters are dynamically updated through recursive Bayesian filtering, and a pre-training strategy is adjusted online through a self-adaptive strategy optimization framework; and judging whether the real-time sensing data reaches a termination condition or not, if not, returning, and if yes, outputting an optimization result. According to the method, an adversarial domain randomization training strategy is adopted, so that the reinforcement learning strategy actively learns to cope with the worst working condition disturbance, and the robustness of the control strategy is improved. In the face of a real and uncertain rolling environment, the problems that in a traditional method, the occurrence rate of folding cracks is high and the yield is low due to model mismatching are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control technology for metal rolling, and in particular to an adaptive dynamic optimization method and system for uncertainty in H-beam rolling. Background Technology

[0002] With the rapid development of industrial automation and intelligence, the H-beam rolling process faces increasingly complex challenges. Traditional rolling control methods mainly rely on preset fixed roll pass sequences and reduction rate adjustment rules based on temperature thresholds. While this method improves product quality to some extent, it has significant limitations.

[0003] First, traditional methods often rely on static and offline process parameters, making it difficult to adapt to uncertainties arising from fluctuations in incoming materials, equipment wear, and operational condition drift during rolling. Second, these methods have limited perception of process states, and control decisions depend primarily on incomplete model information, leading to a "model-reality discrepancy." Strategies that perform well in digital twins may fail in real rolling processes. To address these issues, the combination of digital twins and reinforcement learning techniques has provided new insights for optimizing H-beam billet rolling in recent years. However, ensuring that reinforcement learning strategies are robust to model uncertainties and can continuously adapt in real-world environments remains a critical technical bottleneck. Simply applying standard reinforcement learning algorithms often results in policy performance degradation due to insufficient model accuracy, failing to meet the requirements of high-precision, high-reliability rolling production.

[0004] Regarding solutions to the problem of improving the robustness and economic efficiency of H-beam billet rolling control strategies, the following existing technologies are included.

[0005] For example, CN120146482A discloses a multi-objective collaborative optimization scheduling method and system for industrial parks based on artificial intelligence. This method constructs a digital twin model and a dynamic topology graph, utilizes graph neural networks to model resource flow relationships, combines a multi-model prediction framework for resource demand analysis, and employs a multi-agent reinforcement learning algorithm to generate scheduling strategies. However, this method still has room for improvement in the structure and algorithm of the graph neural network, making it difficult to fully enhance the accuracy and efficiency of relationship modeling.

[0006] CN119180189A discloses a dynamic optimization method and system for digital twins based on machine learning. This method acquires real-time data streams from the manufacturing environment to obtain real-time performance data of the digital twin model. Based on the real-time data streams, performance data, or real-time task requirements, it determines whether to adjust the network structure of the digital twin model, and optimizes the network structure and updates parameters when necessary. However, this method still has room for improvement in its network structure optimization algorithm, making it difficult to achieve more efficient and accurate optimization results.

[0007] It is evident that existing technologies have the following drawbacks in practical implementation:

[0008] 1. Traditional H-beam billet rolling control methods mainly rely on preset fixed roll pass sequences and reduction rate adjustment rules based on temperature thresholds. The process parameters of this method are often static and offline, making it difficult to adapt to the uncertainties caused by material fluctuations, equipment wear, and operating condition drift during the rolling process. This leads to control decisions relying on incomplete model information, causing the "model-reality discrepancy" problem.

[0009] 2. Existing digital twin and reinforcement learning technologies still have shortcomings in the optimization of H-beam billet rolling. They are difficult to make reinforcement learning strategies robust to model uncertainties, cannot continuously adapt in real environments, and are prone to policy performance degradation due to insufficient model accuracy.

[0010] 3. Current H-beam billet rolling control methods still have limitations in their ability to perceive and process complex rolling processes, making it difficult to effectively integrate and analyze multi-dimensional data in real time, which affects the further improvement of rolling efficiency and product quality.

[0011] 4. Existing algorithms and methods for network structure optimization and parameter updating still have room for improvement, making it difficult to fully enhance the accuracy and efficiency of digital twin models, thus affecting the optimization effect of the rolling process.

[0012] 5. Traditional H-beam billet rolling control methods lack the ability to quickly adapt to special working conditions such as new steel grades and specifications, making it difficult to meet the requirements of modern industrial production for flexibility and adaptability.

[0013] Therefore, there is an urgent need for a dynamic optimization method for H-beam billet rolling that can explicitly handle model uncertainties and possess online adaptive capabilities. This method should effectively bridge the gap between digital twins and the real world, achieving comprehensive perception and dynamic optimization of complex rolling processes, thereby significantly improving rolling efficiency, product quality, and system reliability. Summary of the Invention

[0014] To address the aforementioned problems in the existing technology, the purpose of this invention is to provide an adaptive dynamic optimization method, system, and medium for H-beam rolling uncertainty.

[0015] Another objective of this invention is to provide an adaptive dynamic optimization system for H-beam rolling uncertainty.

[0016] To address the above problems, the present invention adopts the following technical solution: a dynamic optimization method for adaptive uncertainty in H-beam rolling, the method comprising the following steps:

[0017] Step 1: Construct a probabilistic digital twin that includes structured uncertainty representations during the H-beam rolling process;

[0018] Step 2: Using a domain randomization method based on adversarial perturbation, a reinforcement learning robust control policy is trained in the probabilistic digital twin;

[0019] Step 3: During the actual rolling process of H-beams, real-time sensor data is acquired, the digital twin parameters are dynamically updated through recursive Bayesian filtering, and the pre-trained strategy is adjusted online using an adaptive strategy optimization framework.

[0020] Step 4: Determine whether the real-time sensing data mentioned in Step 3 has reached the H-beam rolling termination condition. If not, return to Step 3; if yes, proceed to Step 5.

[0021] Step 5: Output the optimization results.

[0022] Furthermore, step 1 includes:

[0023] Step 101: Establish the basic physical model of the H-beam billet rolling process, including the heat conduction model, the material rheology model, and the rolling mechanics model;

[0024] Step 102: Construct a structured uncertainty characterization model for key process parameters in the H-beam rolling process, and model the structured uncertainty parameters as random variables with a probability distribution that fluctuates within the range of ±15% to ±25% of the nominal value;

[0025] Step 103: Integrate the basic physical model, uncertainty characterization, and multi-physics coupling relationship based on the key process parameters to construct a probabilistic digital twin.

[0026] Furthermore, the domain randomization method based on adversarial perturbation described in step 2 is implemented through a minimax game between a policy network and an adversarial network, and its objective function for optimization is as follows:

[0027] (1)

[0028] in:

[0029] : For policy networks Parameters;

[0030] : Parameters for adversarial networks;

[0031] : for from Parameterized probability distribution The environmental disturbance vector sampled in the middle;

[0032] : Expected cumulative reward;

[0033] : Calculated for expectation;

[0034] : In a disturbed environment Next, strategy The expected cumulative reward obtained;

[0035] The goal of adversarial networks is to generate perturbations. To maximize the loss of the policy, the objective of the policy network is to address the perturbation. Minimize the loss to enable the policy network to learn robustness over a wide range of uncertainties.

[0036] Furthermore, step 2 specifically includes:

[0037] Step 201: Set up the strategy network. The input data includes the billet temperature field and the current posterior estimate of the uncertainty parameters. The output is the roll gap setting value, rolling speed setting value, and other relevant control parameters for each stand of the H-beam billet.

[0038] Step 202: Set up an adversarial network to generate the most unfavorable perturbation combination within the feasible region of the structured uncertainty parameters;

[0039] Step 203: Through minimax game between the policy network and the adversarial network, the policy network learns a robust control policy within a wide range of uncertainty in the actual working conditions.

[0040] Furthermore, step 3 includes:

[0041] Step 301: Collect process parameter data such as rolling force, torque, temperature, and billet shape in real time through a multimodal sensor network deployed on the production line;

[0042] Step 302: Using the recursive Bayesian filtering algorithm, the posterior probability distribution of the process parameters is dynamically updated using the real-time collected data as the observations.

[0043] Step 303: Based on the updated probabilistic digital twin, the pre-trained robust control strategy is adjusted online using an adaptive strategy optimization framework to adapt and optimize the process parameters to the current real-time operating conditions.

[0044] Furthermore, in step 302, the recursive Bayesian filtering algorithm dynamically updates the probability distribution of process parameters through prediction and updating. Its core recursive equation is as follows:

[0045] The prediction step specifically includes:

[0046] (2)

[0047] The update steps are as follows:

[0048] (3)

[0049] In the formula:

[0050] In order to be in The parameters to be calibrated at any given time;

[0051] In order to be in Sensor measured data at any given time;

[0052] From time 1 to time 2 All observation data at any given time;

[0053] The prior probability distribution of the current state is based on all past observation data;

[0054] This is a state transition model to describe the evolution of key process parameters during the H-beam rolling process;

[0055] To observe the likelihood model, at the current time k, the key process parameters... Sensor readings predicted by physical models The probability density of occurrence;

[0056] It is the optimal probability estimate of unmeasurable key process parameters in the H-beam rolling process after integrating all sensor observations at the current moment;

[0057] use New observational data at time For the prior distribution obtained from the prediction step After correction, the posterior distribution is obtained. .

[0058] Furthermore, step 302 specifically involves,

[0059] Step 3021: Initialize the parameters and state of the Bayesian filter, and set the initial prior distribution of the key process parameters to a Gaussian distribution centered at the nominal value with a standard deviation of 0.1;

[0060] Step 3022: At the current time k, predict the parameter distribution based on the state transition model p(xk|xk-1) to obtain the prior distribution p(xk|z1:k-1);

[0061] Step 3023: Input the parameter samples in the prior distribution into the probability digital twin, calculate the corresponding sensor prediction value z^k through its forward simulation function, and construct the observation likelihood p(zk|xk);

[0062] Step 3024: Combining the actual observation data zk with the observation likelihood, update the parameter distribution according to Bayes' theorem to obtain the posterior distribution p(xk|z1:k);

[0063] Step 3025: Use the posterior distribution as the calibration result at the current time, update the probabilistic digital twin, and then execute step 303.

[0064] Furthermore, step 303, based on the updated probabilistic digital twin, utilizes an adaptive policy optimization framework to adjust the pre-trained robust control strategy online, so that the process parameters adapt to and optimize the current real-time operating conditions, specifically includes:

[0065] Step 3031: Input the pre-trained robust control policy network parameters into the adaptive optimizer as the initial policy for online adjustment;

[0066] Step 3032: Based on the posterior distribution p(xk|z1:k) of the key process parameters output in step 302 and the real-time operating conditions of the current pass, construct a dynamic optimization objective for H-beam dimensional accuracy and rolling stability.

[0067] Step 3033: Within the neighborhood of the pre-trained policy parameters, a lightweight online learning algorithm is used to adaptively adjust the policy network;

[0068] Step 3034: In the updated probabilistic digital twin, the robustness of the adaptive strategy is verified based on the posterior distribution;

[0069] Step 3035: If the adaptive strategy meets the performance improvement requirements and complies with the security constraints, then accept the update; otherwise, retain the original strategy.

[0070] Step 3036: Determine whether the online optimization time limit or performance convergence threshold has been reached; if so, terminate the optimization, output the current adaptive strategy, and proceed to the strategy deployment and execution steps.

[0071] Furthermore, it also includes executing step 304, specifically:

[0072] Step 3041: Convert the optimized strategy into executable code;

[0073] Step 3042: Download the code to the edge controller;

[0074] Step 3043: Configure rolling parameters and execute rolling operation.

[0075] This application also provides a system for implementing the dynamic optimization method for adaptive uncertainty in H-beam rolling, including a sensing layer comprising an infrared thermal imager, a strain sensor, a mechanical sensor, and a visual sensor, for collecting real-time sensing data during the H-beam rolling process, providing multimodal real-time observation input for subsequent data processing and control decisions;

[0076] It includes an edge layer, which is used to preprocess the real-time data collected by the perception layer and execute real-time policy instructions; it uploads the preprocessed real-time data to the cloud / server layer, receives optimized control instructions and sends them to the execution layer;

[0077] It includes an execution layer, which drives the rolling mill to complete the actual rolling action based on the control instructions issued by the edge layer;

[0078] Including the cloud / server layer, including:

[0079] The online Bayesian calibration and adaptive strategy optimization module dynamically updates the probability distribution of key process parameters based on real-time observation data using a recursive Bayesian filtering algorithm, and adjusts the pre-trained strategy online in conjunction with the current operating conditions.

[0080] The probabilistic digital twin module is used to construct simulation models that integrate multi-physics coupling relationships, characterize structured uncertainties, and support policy performance evaluation and robustness verification.

[0081] The adversarial domain randomized training module trains a robust control policy under simulated worst-case perturbation conditions through a minimax game between the policy network and the adversarial network.

[0082] The edge layer uploads preprocessed real-time data to the cloud / server layer, triggering the online Bayesian calibration and adaptive policy optimization module; the calibrated parameters are fed back to the probabilistic digital twin module; the policy performance feedback is used to guide the adversarial domain randomization training module to continuously optimize the policy network; the optimized policy parameters are sent down to the execution layer via the edge layer.

[0083] Compared with the prior art, the beneficial technical effects of the present invention are as follows:

[0084] 1. This application constructs a probabilistic digital twin containing structured uncertainty representations and employs an adversarial domain randomization training strategy, enabling the reinforcement learning strategy to actively learn to cope with the most severe working condition disturbances, thereby improving the robustness of the control strategy. When facing real, uncertain rolling environments, performance degradation is significantly reduced, effectively solving the problems of high folding crack incidence and low yield caused by model mismatch in traditional methods.

[0085] 2. Through the online Bayesian model calibration provided in this application, the digital twin is no longer a static model that deviates from reality, but a dynamic model that can continuously self-correct using real-time data, maintaining high fidelity. The dynamic evolution mechanism effectively solves the problem of deviation between digital twins and the real world in the prior art.

[0086] 3. This application adopts a new industrial AI model of "pre-training-fine-tuning" to complete the time-consuming deep reinforcement learning training in an offline simulation environment, and only performs lightweight and rapid fine-tuning online, which not only ensures the advancement of the strategy, but also meets the stringent real-time requirements of industrial sites.

[0087] 4. This application achieves comprehensive perception and dynamic analysis of complex rolling processes through real-time data acquisition and recursive Bayesian filtering of a multimodal sensor network, effectively integrating multi-dimensional data such as rolling force, torque, temperature, and billet shape, and significantly improving rolling efficiency and product quality.

[0088] 5. Based on the updated probabilistic digital twin, this application utilizes an adaptive policy optimization framework to rapidly fine-tune the pre-trained robust policy online, enabling rapid adaptation to special working conditions such as new steel grades and specifications, thus meeting the requirements of modern industrial production for flexibility and adaptability. Attached Figure Description

[0089] Figure 1 This is a system architecture diagram of an embodiment of this application;

[0090] Figure 2 This is a flowchart of the method described in the embodiments of this application;

[0091] Figure 3 This is a diagram illustrating the probabilistic digital twin architecture of an embodiment of this application.

[0092] Figure 4 This is a schematic diagram of the adversarial threshold randomization training framework structure according to an embodiment of this application;

[0093] Figure 5 For online adaptive calibration and fine-tuning timing diagrams;

[0094] Figure 6 This is a comparison chart of uncertainty characterization and model calibration results;

[0095] Figure 7 Workflow diagram for adaptive strategy optimization framework. Detailed Implementation

[0096] The technical solution of the present invention will be further described clearly and in detail below with reference to the embodiments and accompanying drawings.

[0097] Example

[0098] like Figure 1 As shown, this application provides an adaptive dynamic optimization system for H-beam rolling uncertainty. The system includes a four-layer architecture, specifically:

[0099] It includes a sensing layer, which includes an infrared thermal imager, strain sensor, mechanical sensor and vision sensor, used to collect real-time sensing data during the H-beam rolling process, providing multimodal real-time observation input for subsequent data processing and control decisions;

[0100] It includes an edge layer, which is used to preprocess the real-time data collected by the perception layer and execute real-time policy instructions; it uploads the preprocessed real-time data to the cloud / server layer, receives optimized control instructions and sends them to the execution layer;

[0101] It includes an execution layer, which drives the rolling mill to complete the actual rolling action based on the control instructions issued by the edge layer;

[0102] Including the cloud / server layer, including:

[0103] The online Bayesian calibration and adaptive strategy optimization module dynamically updates the probability distribution of key process parameters based on real-time observation data using a recursive Bayesian filtering algorithm, and adjusts the pre-trained strategy online in conjunction with the current operating conditions.

[0104] The probabilistic digital twin module is used to construct simulation models that integrate multi-physics coupling relationships, characterize structured uncertainties, and support policy performance evaluation and robustness verification.

[0105] The adversarial domain randomized training module trains a robust control policy under simulated worst-case perturbation conditions through a minimax game between the policy network and the adversarial network.

[0106] The edge layer uploads preprocessed real-time data to the cloud / server layer, triggering the online Bayesian calibration and adaptive policy optimization module; the calibrated parameters are fed back to the probabilistic digital twin module; the policy performance feedback is used to guide the adversarial domain randomization training module to continuously optimize the policy network; the optimized policy parameters are sent down to the execution layer via the edge layer.

[0107] Figure 2 As shown, this invention provides an uncertainty-aware adaptive dynamic optimization method for H-beam billet rolling, and the specific implementation steps are as follows:

[0108] Step 1: Construct a probabilistic digital twin containing structured uncertainty representations, such as... Figure 3 As shown, it includes:

[0109] Step 101: Establish the basic physical model of the H-beam billet rolling process. First, a heat conduction model is established to describe the heat conduction process inside the billet, using the finite element method to simulate the temperature field. Second, a material rheology model is constructed to describe the rheological behavior of the billet at high temperatures, using the Johnson-Cook material model. Finally, a rolling mechanics model is established to describe the mechanical response during the rolling process, using the Bauschinger model. These models together constitute the basic physical model of the H-beam billet rolling process.

[0110] Step 102: Construct structured uncertainty characterization models for key process parameters. For key process parameters such as friction coefficient, material rheological stress, and heat transfer coefficient, probability distribution models are used to describe their uncertainties. Specifically, the friction coefficient is modeled as a uniform distribution within the range of 0.3 ± 0.06; the material rheological stress is modeled as a triangular distribution within the range of 400-600 MPa; and the heat transfer coefficient is modeled as a linear distribution within the range of 50-100 W / (m·K). These probability distribution models collectively form the basis of the probabilistic digital twin.

[0111] Step 103: Establish the coupling relationship between process parameters. Using system identification technology and based on a large amount of historical production data, establish the coupling relationship between various process parameters and construct a probabilistic digital twin. Specifically, the coupling relationship between the temperature field and the reduction rate and rolling speed is described by a polynomial mapping model; the coupling relationship between billet deformation and the temperature field and stress field is described by a convolutional neural network.

[0112] Step 2: Employ an adversarial domain randomization training strategy, such as... Figure 4 As shown:

[0113] The adversarial domain randomization training strategy is implemented through a minimax game between the policy network and the adversarial network, and its objective function is as shown in formula (1): The goal of adversarial networks is to generate perturbations. To maximize the loss of the strategy (i.e. minimize the reward) The goal of the policy network is to minimize the loss under the worst-case perturbation, making policy learning robust across a wide range of uncertainties.

[0114] Step 201: Design the strategy network. A deep neural network is used as the strategy network. The input layer includes a 64×64 temperature field, 16 nodes of size and morphology features, and 5 nodes of uncertainty parameter estimates. The output layer includes 2 nodes, representing the reduction rate and rolling speed, respectively. The hidden layer uses the ReLU activation function and has 2 nodes.

[0115] Step 202: Design an adversarial network. A generative adversarial network is used as the adversarial network. The generator has 5 convolutional layers and the discriminator has 4 convolutional layers. The sample size generated by the generator is 64×64.

[0116] Step 203: Perform adversarial training. First, fix the policy network and train the adversarial network; then fix the adversarial network and train the policy network. Through 2 million steps of adversarial training, the policy network is forced to learn a robust optimal policy within a wide range of uncertainties covering 99% of actual working conditions.

[0117] Step 3: Online adaptive calibration and strategy fine-tuning, such as... Figure 5 As shown:

[0118] Step 301, Data Sensing Stage: Deploy a multimodal sensor network. During the rolling process, deploy infrared thermal imagers, strain sensors, mechanical sensors, and vision sensors to collect data such as temperature field, deformation field, rolling force, torque, and billet shape in real time.

[0119] Step 302, Model Calibration Stage: Perform Recursive Bayesian filtering. Using the data collected in Step 301 as observations, the Recursive Bayesian filtering algorithm is applied, and the update steps are as follows:

[0120] The recursive Bayesian filtering algorithm dynamically updates the probability distribution of key process parameters through two steps: prediction and update. Figure 6 As shown, its core recursive equation is as follows:

[0121] For the prediction step, see formula (2): The update step is shown in formula (3): process utilization New observational data at time For the prior distribution obtained from the prediction step Make corrections to obtain a more accurate posterior distribution. Specifically,

[0122] Step 3021: Initialize the parameters and state of the Bayesian filter, and set the initial prior distribution of the key process parameters to a Gaussian distribution centered at the nominal value with a standard deviation of 0.1;

[0123] Step 3022: At the current time k, predict the parameter distribution based on the state transition model p(xk|xk-1) to obtain the prior distribution p(xk|z1:k-1);

[0124] Step 3023: Input the parameter samples in the prior distribution into the probability digital twin, calculate the corresponding sensor prediction value z^k through its forward simulation function, and construct the observation likelihood p(zk|xk);

[0125] Step 3024: Combining the actual observation data zk with the observation likelihood, update the parameter distribution according to Bayes' theorem to obtain the posterior distribution p(xk|z1:k);

[0126] Step 3025: Use the posterior distribution as the calibration result at the current time, update the probabilistic digital twin, and then execute step 303.

[0127] Step 303, strategy optimization, based on the updated probabilistic digital twin, such as Figure 7 As shown, an adaptive policy optimization framework is used to adjust the pre-trained robust control strategy online, so that the process parameters adapt to and optimize the current real-time operating conditions. Specifically, this includes:

[0128] Step 3031: Input the pre-trained robust control policy network parameters into the adaptive optimizer as the initial policy for online adjustment;

[0129] Step 3032: Based on the posterior distribution p(xk|z1:k) of the key process parameters output in step 302 and the real-time operating conditions of the current pass, construct a dynamic optimization objective for H-beam dimensional accuracy and rolling stability.

[0130] Step 3033: Within the neighborhood of the pre-trained policy parameters, a lightweight online learning algorithm is used to adaptively adjust the policy network; specifically, candidate policies are generated and evaluated in the following manner:

[0131] a) Define the strategy search space: Set a limited adjustment range (e.g., reduction rate ±20%, rolling speed ±10%) around the output instructions of the current strategy network (such as reduction rate, rolling speed).

[0132] b) Generate and evaluate candidate policies: Within the search space, generate a set of candidate policy parameters using methods such as random sampling or gradient guidance. Utilize the dynamic environment constructed in step 3032 to perform a fast simulation rollout for each candidate policy and calculate its expected cumulative reward;

[0133] Step 3034: In the updated probabilistic digital twin, the robustness of the adaptive strategy is verified based on the posterior distribution;

[0134] Step 3035: If the adaptive strategy meets the performance improvement requirements and complies with the security constraints, then accept the update; otherwise, retain the original strategy.

[0135] Step 3036: Determine whether the online optimization time limit or performance convergence threshold has been reached (e.g., reaching the maximum optimization rounds of 500, the performance improvement is less than the set threshold, or the optimization time has reached the real-time limit); if so, terminate the optimization, output the current adaptive strategy, and proceed to the strategy deployment and execution steps.

[0136] If the termination condition is not met, return to step 3033 and continue iterative optimization based on the new current strategy;

[0137] Step 304: Control Execution Phase, Execution Strategy Transformation and Deployment:

[0138] Step 3041: Convert the optimized strategy into executable code;

[0139] Step 3042: Download the code to the edge controller;

[0140] Step 3043: Configure rolling parameters and execute rolling operation until the rolling termination condition is met.

[0141] Example 2

[0142] This invention provides an adaptive dynamic optimization method for H-beam rolling uncertainty, the specific implementation steps of which are as follows:

[0143] Step 1: Construct a probabilistic digital twin containing structured uncertainty representations:

[0144] Step 101: Establish the basic physical model of the H-beam billet rolling process. First, a heat conduction model is established using the finite element software Abaqus to create a temperature field model. Second, a material rheological model is constructed using the Zener-Hollomon parametric model to describe the rheological behavior of the billet at high temperatures. Finally, a rolling mechanics model is established using the LSD model to describe the mechanical response during the rolling process. These models together constitute the basic physical model of the H-beam billet rolling process.

[0145] Step 102: Construct structured uncertainty characterization models for key process parameters. For key process parameters such as friction coefficient, material rheological stress, and heat transfer coefficient, probability distribution models are used to describe their uncertainties. Specifically, the friction coefficient is modeled as a uniform distribution within the range of 0.35±0.07; the material rheological stress is modeled as a triangular distribution within the range of 450-550 MPa; and the heat transfer coefficient is modeled as a linear distribution within the range of 60-80 W / (m·K). These probability distribution models collectively form the basis of the probabilistic digital twin.

[0146] Step 103: Establish the coupling relationship between process parameters. Using machine learning algorithms and based on a large amount of historical production data, establish the coupling relationship between various process parameters and construct a probabilistic digital twin. Specifically, the coupling relationship between the temperature field and the reduction rate and rolling speed is described by a multinomial regression model; the coupling relationship between billet deformation and the temperature field and stress field is described by a convolutional neural network.

[0147] Step 2: Employ an adversarial domain randomization training strategy:

[0148] The adversarial domain randomization training strategy is implemented through a minimax game between the policy network and the adversarial network, and its objective function is as shown in formula (1): The goal of adversarial networks is to generate perturbations. To maximize the loss of the strategy (i.e. minimize the reward) The goal of the policy network is to minimize the loss under the worst-case perturbation, thereby forcing the policy to learn robustness over a wide range of uncertainties.

[0149] Step 201: Design the policy network. A deep reinforcement learning network is used as the policy network. The input layer includes a 128×128 temperature field, 32 nodes for size and shape features, and 6 nodes for uncertainty parameter estimates. The output layer includes 3 nodes, representing the reduction rate, rolling speed, and lubricant addition amount, respectively. The hidden layers use the LeakyReLU activation function and have 3 nodes.

[0150] Step 202: Design the adversarial network. A generative adversarial network (GAN) is used as the adversarial network. The generator uses 6 residual convolutional layers, and the discriminator uses 5 residual convolutional layers. The sample size generated by the generator is 128×128.

[0151] Step 203: Perform adversarial training. First, fix the policy network and train the adversarial network; then fix the adversarial network and train the policy network. Through 3 million steps of adversarial training, the policy network is forced to learn a robust optimal policy within a wide range of uncertainties covering 99.5% of real-world operating conditions.

[0152] Step 3: Online Bayesian model calibration and strategy fine-tuning:

[0153] Step 301: Deploy a multimodal sensor network. During the rolling process, deploy FLIR thermal imagers, sheet strain gauges, Kistler mechanical sensors, and Keyence vision sensors to collect data such as temperature field, deformation field, rolling force, torque, and billet shape in real time.

[0154] Step 302: Perform recursive Bayesian filtering. Using the collected data as observations, the recursive Bayesian filtering algorithm is applied, and the update steps are as follows:

[0155] The recursive Bayesian filtering algorithm dynamically updates the probability distribution of key process parameters through two steps: prediction and update. Its core recursive equation is as follows:

[0156] The prediction step is shown in formula (2): The update step is shown in formula (3): This process utilizes New observational data at time For the prior distribution obtained from the prediction step Make corrections to obtain a more accurate posterior distribution. .

[0157] Step 3021: Initialize the filter parameters and state, set the initial error covariance to 0.05, and the initial transition matrix to the Toeplitz matrix;

[0158] Step 3022: Calculate the measurement model to obtain the predicted values. The predicted values ​​are calculated from the physical model;

[0159] Step 3023: Calculate the residuals to obtain the difference between the observed and predicted values;

[0160] Step 3024: Update the filter parameters and status;

[0161] Step 3025: Determine whether the convergence condition has been met. If not, return to step 3022. If yes, proceed to step 303.

[0162] In step 303, the adaptive strategy optimization framework achieves rapid online adjustment through the following specific steps:

[0163] Step 3031, Policy Adaptive Initialization: Load the offline trained robust policy network as the initial policy for online optimization, and set optimization parameters (such as learning rate, search range, etc.).

[0164] Step 3032: Construct a dynamic optimization objective: Based on the probabilistic digital twin updated by recursive Bayesian filtering, construct a simulation environment for the current rolling pass and define a reward function for evaluating the strategy performance (this function comprehensively considers indicators such as cross-sectional dimension accuracy, rolling force stability, and energy consumption).

[0165] Step 3033: Perform policy parameter adaptation: Within the parameter neighborhood of the initial policy, use an efficient online learning algorithm to search for and optimize the policy. Specifically, candidate policies are generated and evaluated in the following way:

[0166] a) Define the strategy search space: Set a limited adjustment range (e.g., reduction rate ±18%, rolling speed ±8%) around the output instructions of the current strategy network (such as reduction rate, rolling speed).

[0167] b) Generate and evaluate candidate policies: Within the search space, generate a set of candidate policy parameters using methods such as random sampling or gradient guidance. Utilize the dynamic environment constructed in step 3032 to perform a fast simulation rollout for each candidate policy and calculate its expected cumulative reward.

[0168] Step 3034: Update the optimal policy: Compare the performance of all candidate policies. If there is a candidate with better performance than the current policy, update the policy network parameters and set the candidate policy with the best performance as the new current policy.

[0169] Step 3035, Convergence Judgment and Loop: Determine whether the optimization process meets the preset termination conditions (e.g., reaching the maximum optimization rounds of 500 rounds, the performance improvement is less than the set threshold, or the optimization time reaches the real-time limit).

[0170] If the termination condition is not met, return to step 3033 and continue iterative optimization based on the new current strategy;

[0171] If the termination condition is met, the online optimization process will terminate, the current optimal strategy will be output, and the strategy deployment and execution steps will begin.

[0172] Step 304: Execution Strategy Transformation and Deployment:

[0173] Step 3041: Convert the optimized strategy into executable code;

[0174] Step 3042: Download the code to the edge controller;

[0175] Step 3043: Configure rolling parameters and execute rolling operation until the rolling termination condition is met.

[0176] Finally, it should be pointed out that the above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A dynamic optimization method for adaptively addressing uncertainties in H-beam rolling, characterized in that, The method includes the following steps: Step 1: Construct a probabilistic digital twin that includes structured uncertainty representations during the H-beam rolling process; Step 2: Using a domain randomization method based on adversarial perturbation, a reinforcement learning robust control policy is trained in the probabilistic digital twin; Step 3: During the actual rolling process of H-beams, real-time sensor data is acquired, the digital twin parameters are dynamically updated through recursive Bayesian filtering, and the pre-trained strategy is adjusted online using an adaptive strategy optimization framework. Step 4: Determine whether the real-time sensing data mentioned in Step 3 has reached the H-beam rolling termination condition. If not, return to Step 3; if yes, proceed to Step 5. Step 5: Output the optimization results.

2. The adaptive dynamic optimization method for H-beam rolling uncertainty according to claim 1, characterized in that, Step 1 includes: Step 101: Establish the basic physical model of the H-beam billet rolling process, including the heat conduction model, the material rheology model, and the rolling mechanics model; Step 102: Construct a structured uncertainty characterization model for key process parameters in the H-beam rolling process, and model the structured uncertainty parameters as random variables with a probability distribution that fluctuates within the range of ±15% to ±25% of the nominal value; Step 103: Integrate the basic physical model, uncertainty characterization, and multi-physics coupling relationship based on the key process parameters to construct a probabilistic digital twin.

3. The adaptive dynamic optimization method for H-beam rolling uncertainty according to claim 1, characterized in that, The domain randomization method based on adversarial perturbation described in step 2 is implemented through a minimax game between a policy network and an adversarial network, and its objective function is as follows: (1) in: : For policy networks Parameters; : Parameters for adversarial networks; : for from Parameterized probability distribution The environmental disturbance vector sampled in the middle; : Expected cumulative reward; : Calculated for expectation; : In a disturbed environment Next, strategy The expected cumulative reward obtained; The goal of adversarial networks is to generate perturbations. To maximize the loss of the policy, the objective of the policy network is to address the perturbation. Minimize the loss to enable the policy network to learn robustness over a wide range of uncertainties.

4. The adaptive dynamic optimization method for H-beam rolling uncertainty according to claim 1, characterized in that, Step 2 specifically involves: Step 201: Set up the strategy network. The input data includes the billet temperature field and the current posterior estimate of the uncertainty parameters. The output is the roll gap setting value, rolling speed setting value, and other relevant control parameters for each stand of the H-beam billet. Step 202: Set up an adversarial network to generate the most unfavorable perturbation combination within the feasible region of the structured uncertainty parameters; Step 203: Through minimax game between the policy network and the adversarial network, the policy network learns a robust control policy within a wide range of uncertainty in the actual working conditions.

5. The adaptive dynamic optimization method for H-beam rolling uncertainty according to claim 1, characterized in that, Step 3 includes: Step 301: Collect process parameter data such as rolling force, torque, temperature, and billet shape in real time through a multimodal sensor network deployed on the production line; Step 302: Using the recursive Bayesian filtering algorithm, the posterior probability distribution of the process parameters is dynamically updated using the real-time collected data as the observations. Step 303: Based on the updated probabilistic digital twin, the pre-trained robust control strategy is adjusted online using an adaptive strategy optimization framework to adapt and optimize the process parameters to the current real-time operating conditions.

6. The adaptive dynamic optimization method for H-beam rolling uncertainty according to claim 5, characterized in that, In step 302, the recursive Bayesian filtering algorithm dynamically updates the probability distribution of process parameters through prediction and updating. Its core recursive equation is as follows: The prediction step specifically includes: (2) The update steps are as follows: (3) In the formula: In order to be in The parameters to be calibrated at any given time; In order to be in Sensor measured data at any given time; From time 1 to time 2 All observation data at any given time; This is the prior probability distribution of the current state based on all past observation data; This is a state transition model to describe the evolution of key process parameters during the H-beam rolling process; To observe the likelihood model, at the current time k, the key process parameters... Sensor readings predicted by physical models The probability density of occurrence; It is the optimal probability estimate of unmeasurable key process parameters in the H-beam rolling process after integrating all sensor observations at the current moment; use New observational data at time For the prior distribution obtained from the prediction step After correction, the posterior distribution is obtained. .

7. The adaptive dynamic optimization method for H-beam rolling uncertainty according to claim 6, characterized in that, Specifically, step 302 is as follows: Step 3021: Initialize the parameters and state of the Bayesian filter, and set the initial prior distribution of the key process parameters to a Gaussian distribution centered at the nominal value with a standard deviation of 0.1; Step 3022: At the current time k, predict the parameter distribution based on the state transition model p(xk|xk-1) to obtain the prior distribution p(xk|z1:k-1); Step 3023: Input the parameter samples in the prior distribution into the probability digital twin, calculate the corresponding sensor prediction value z^k through its forward simulation function, and construct the observation likelihood p(zk|xk); Step 3024: Combining the actual observation data zk with the observation likelihood, update the parameter distribution according to Bayes' theorem to obtain the posterior distribution p(xk|z1:k); Step 3025: Use the posterior distribution as the calibration result at the current time, update the probabilistic digital twin, and then execute step 303.

8. The adaptive dynamic optimization method for H-beam rolling uncertainty according to claim 1, characterized in that, Step 303, based on the updated probabilistic digital twin, utilizes an adaptive policy optimization framework to adjust the pre-trained robust control strategy online, so that the process parameters adapt to and optimize the current real-time operating conditions, specifically includes: Step 3031: Input the pre-trained robust control policy network parameters into the adaptive optimizer as the initial policy for online adjustment; Step 3032: Based on the posterior distribution p(xk|z1:k) of the key process parameters output in step 302 and the real-time operating conditions of the current pass, construct a dynamic optimization objective for H-beam dimensional accuracy and rolling stability. Step 3033: Within the neighborhood of the pre-trained policy parameters, a lightweight online learning algorithm is used to adaptively adjust the policy network; Step 3034: In the updated probabilistic digital twin, the robustness of the adaptive strategy is verified based on the posterior distribution; Step 3035: If the adaptive strategy meets the performance improvement requirements and complies with the security constraints, then accept the update; otherwise, retain the original strategy. Step 3036: Determine whether the online optimization time limit or performance convergence threshold has been reached; if so, terminate the optimization, output the current adaptive strategy, and proceed to the strategy deployment and execution steps.

9. The adaptive dynamic optimization method for H-beam rolling uncertainty according to claim 5, characterized in that, This also includes executing step 304, which is as follows: Step 3041: Convert the optimized strategy into executable code; Step 3042: Download the code to the edge controller; Step 3043: Configure rolling parameters and execute rolling operation.

10. A system for implementing the adaptive dynamic optimization method for H-beam rolling uncertainty as described in claim 1, characterized in that, It includes a sensing layer, which includes an infrared thermal imager, strain sensor, mechanical sensor and vision sensor, used to collect real-time sensing data during the H-beam rolling process, providing multimodal real-time observation input for subsequent data processing and control decisions; It includes an edge layer, which is used to preprocess real-time data collected by the perception layer and execute real-time policy instructions; The preprocessed real-time data is uploaded to the cloud / server layer, and optimized control commands are received and sent to the execution layer. It includes an execution layer, which drives the rolling mill to complete the actual rolling action based on the control instructions issued by the edge layer; Including the cloud / server layer, including: The online Bayesian calibration and adaptive strategy optimization module dynamically updates the probability distribution of key process parameters based on real-time observation data using a recursive Bayesian filtering algorithm, and adjusts the pre-trained strategy online in conjunction with the current operating conditions. The probabilistic digital twin module is used to construct simulation models that integrate multi-physics coupling relationships, characterize structured uncertainties, and support policy performance evaluation and robustness verification. The adversarial domain randomized training module trains a robust control policy under simulated worst-case perturbation conditions through a minimax game between the policy network and the adversarial network. The edge layer uploads preprocessed real-time data to the cloud / server layer, triggering the online Bayesian calibration and adaptive policy optimization module; the calibrated parameters are fed back to the probabilistic digital twin module; the policy performance feedback is used to guide the adversarial domain randomization training module to continuously optimize the policy network; the optimized policy parameters are sent down to the execution layer via the edge layer.

Citation Information

Patent Citations

  • Digital twinning dynamic optimization method and system based on machine learning

    CN119180189A

  • Industrial park multi-target collaborative optimization scheduling system and method based on artificial intelligence

    CN120146482A