Method for controlling surface quality by coupling speed and tension of stainless steel recoiling unit
By using a Markov decision process model and a multi-layer reward function established through deep reinforcement learning, combined with a deep Q-network to optimize the speed and tension coupling control of a stainless steel rewinding unit, the problem of insufficient adaptive capability of traditional methods under complex working conditions is solved, and high-precision surface quality control is achieved.
Patent Information
- Application Number
- CN202510969493.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-11-21
AI Technical Summary
The existing speed and tension coupling control method of stainless steel rewinding units is difficult to adapt to complex working conditions, resulting in low surface quality control accuracy and insufficient adaptive capability. Furthermore, traditional control algorithms are difficult to meet the tension control requirements under high-speed operating conditions.
By employing deep reinforcement learning techniques, a Markov decision process model is established, a multi-layer reward function is designed, and a deep Q-network and a target network are combined. Through experience playback and noise simulation, the speed and tension coupling control strategy is optimized to achieve adaptive and precise control for different working conditions.
It significantly improved the level of surface quality control, reduced the defect rate by about 30%, and improved the control accuracy by about 25%, ensuring the consistency and stability of the surface quality of the product in different speed ranges.
Smart Images

Figure CN120993723A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of stainless steel processing equipment control, and specifically to a method for controlling the surface quality of a stainless steel rewinding unit by coupling speed and tension. Background Technology
[0002] With the rapid development of the stainless steel industry, stainless steel rewinding units, as key equipment in stainless steel production lines, directly affect the surface quality and production efficiency of products through the performance of their control systems. In the stainless steel rewinding process, the coupled control of speed and tension is a crucial factor in ensuring product surface quality.
[0003] Currently, the control systems of stainless steel rewinding mills mainly employ traditional methods such as PID control and fuzzy control. These methods have limitations in handling the complex coupling relationship between speed and tension, and are difficult to adapt to the control requirements under different operating conditions. Especially under high-speed operating conditions, traditional control methods struggle to maintain constant tension, easily leading to defects such as scratches and ripples on the product surface, affecting product quality.
[0004] With the development of artificial intelligence technology, deep reinforcement learning, as an effective control method, has been applied in many fields. However, its application in the control of stainless steel processing equipment still faces the following problems: 1. Traditional control algorithms struggle to effectively handle the complex coupling relationship between speed and tension in stainless steel rewinding units, leading to insufficient surface quality control and defects such as scratches or ripples; 2. Existing control methods lack adaptability when facing complex, multi-variable, and time-varying systems, making it difficult to maintain high control performance under changing operating conditions; 3. While existing deep reinforcement learning methods have achieved good results in other fields, they have not been optimized for the specific needs of stainless steel rewinding units, making them difficult to directly apply to the speed and tension coupling control of stainless steel rewinding units; 4. Existing reward function designs do not fully consider multi-objective optimization problems such as surface quality and winding uniformity during stainless steel rewinding, making it difficult to achieve precise control of product quality; 5. There is a lack of comprehensive adaptability to tension control requirements within different speed ranges, making it difficult to meet control requirements under complex operating conditions.
[0005] Therefore, there is an urgent need for a control method for stainless steel rewinding units that can effectively handle the coupling relationship between speed and tension, improve the accuracy of surface quality control, and have good self-adaptive capabilities. Summary of the Invention
[0006] To address the problem that existing stainless steel rewinding machines suffer from insufficient speed and tension coupling control, resulting in low surface quality control accuracy and a lack of adaptability to tension control requirements across different speed ranges, this invention provides a method for controlling surface quality using speed and tension coupling in stainless steel rewinding machines.
[0007] The technical solution adopted by this invention to solve its technical problem is: to provide a method for controlling the surface quality of a stainless steel rewinding unit by coupling speed and tension, comprising the following steps:
[0008] Step 1: Establish a Markov decision process model for the stainless steel rewinding unit, including:
[0009] Step 101: Define state variables, including parameters such as speed, tension, surface quality, winding uniformity, and running speed;
[0010] Step 102: Establish a state transition model to describe the transition relationships between different states;
[0011] Step 103: Design the reward function, including:
[0012] Step 104: The first-level reward function uses the opposite of the absolute value of the angle between the vertical rod of the inverted pendulum and the vertical direction as the angle reward.
[0013] Step 105: Second-level reward function: When the distance 0 < d < 0.05, add 3 to the current reward to improve the balance control accuracy;
[0014] Step 106: The third-level reward function: when the distance d is stable within 0.05m, the second-level reward is used as the reward for the later stage.
[0015] Step 2: Design control strategies, including:
[0016] Step 201: Initialize the deep Q-network and use experience replay and target network techniques to improve learning efficiency and stability;
[0017] Step 202: Input the states such as velocity, tension, and surface quality into the deep Q-network and output the optimal control strategy;
[0018] Step 203: Execute corresponding speed and tension control according to the adopted control strategy, and feed the feedback to the deep Q network in real time for learning;
[0019] Step 3: Determine if the convergence condition has been met. If yes, proceed to Step 4; otherwise, return to Step 2 to continue learning.
[0020] Step 4: Output the optimal control strategy, including:
[0021] Step 401: Save the learned optimal policy to the policy library;
[0022] Step 402: Select the most suitable control strategy from the strategy library based on the current operating conditions;
[0023] Step 403: Perform noise simulation on the strategy to improve its reliability.
[0024] The beneficial effects of this invention are as follows:
[0025] By establishing a Markov decision process model incorporating states such as speed, tension, and surface quality, and designing a corresponding reward function, precise control of the speed-tension coupling relationship was achieved. This effectively solved the problem that traditional control algorithms struggle to handle complex coupling relationships, significantly improving surface quality control and reducing the surface quality defect rate by approximately 30%. The use of Deep Q-Network (DQN) to learn the optimal control strategy, combined with experience playback and target network technology, improved learning efficiency and stability, enabling the control strategy to adapt to speed-tension coupling relationships under different operating conditions in real time. This effectively overcomes the shortcomings of traditional PID controllers in terms of insufficient adaptive capability in complex systems. Through online learning and optimization mechanisms, the accuracy of speed and tension control was continuously improved, increasing control accuracy by approximately 25% and ensuring the stability of speed and tension control, thereby significantly improving the surface quality and accuracy of the product. Adaptive adjustment of the tension control strategy within different speed ranges achieved comprehensive adaptation to different operating conditions, effectively solving the problem of difficult tension control at high speeds and ensuring consistent surface quality of the product under different speed conditions. The innovative application of deep reinforcement learning technology to the speed-tension coupling control of stainless steel rewinding units provides a new solution for resolving complex coupling relationships and has significant practical application value. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating the steps of a method for controlling surface quality using speed and tension coupling in a stainless steel rewinding unit according to the present invention.
[0027] Figure 2 This is a flowchart of step 1 of the method for controlling surface quality by speed and tension coupling in a stainless steel rewinding unit according to the present invention.
[0028] Figure 3 This is a flowchart of step 2 of the method for controlling surface quality by speed and tension coupling in a stainless steel rewinding unit according to the present invention.
[0029] Figure 4 This is a flowchart of step 4 of the method for controlling surface quality by speed and tension coupling in a stainless steel rewinding unit according to the present invention.
[0030] Figure 5 This is a flowchart of step 102 of the method for controlling surface quality by speed and tension coupling of a stainless steel rewinding unit according to the present invention.
[0031] Figure 6 This is a flowchart of steps 104, 105, and 106 of the method for controlling surface quality by speed and tension coupling of a stainless steel rewinding unit according to the present invention.
[0032] Figure 7 This is a flowchart of step 201 of the method for controlling surface quality by speed and tension coupling in a stainless steel rewinding unit according to the present invention.
[0033] Figure 8 This is a flowchart of step 202 of the method for controlling surface quality by speed and tension coupling in a stainless steel rewinding unit according to the present invention.
[0034] Figure 9 This is a flowchart of step 401 of the method for controlling surface quality by speed and tension coupling in a stainless steel rewinding unit according to the present invention.
[0035] Figure 10 This is a flowchart of step 402 in the method for controlling surface quality by speed and tension coupling of a stainless steel rewinding unit according to the present invention.
[0036] Figure 11 This is a flowchart of step 403 in the method for controlling surface quality by speed and tension coupling of a stainless steel rewinding unit according to the present invention. Detailed Implementation
[0037] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0038] like Figure 1-11 As shown, a method for controlling the surface quality of a stainless steel rewinding mill by coupling speed and tension is presented. This method utilizes deep reinforcement learning technology to achieve coupled control of the speed and tension of the stainless steel rewinding mill, thereby improving the surface quality of the product. The specific implementation steps are as follows:
[0039] Step 1: Establish a Markov decision process model for the stainless steel rewinding unit.
[0040] Step 101: Define state variables, including parameters such as speed, tension, surface quality, winding uniformity, and running speed.
[0041] In this embodiment, the state variables are a set of key parameters describing the current operating state of the stainless steel rewinding unit. These state variables include: speed V, which represents the operating speed of the stainless steel rewinding unit, in meters per minute (m / min), with a value range of 0-1200 m / min; tension T, which represents the tension output of the stainless steel rewinding unit, in kilonewtons (kN), with a value range of 5-50 kN; surface quality Q, which represents the surface quality score of the final product, using a scoring system of 0-100, where 100 represents perfect and defect-free; winding uniformity U, which represents the winding uniformity index, using a scoring system of 0-100, where 100 represents complete uniformity; and operating speed V', which represents the rate of change of the current operating speed, in meters per minute², with a value range of -20 to 20 meters per minute².
[0042] These state variables constitute the state space S, i.e., S = {V, T, Q, U, V'}. The state of the system at any time t can be represented as St = {Vt, Tt, Qt, Ut, V't}.
[0043] Step 102: Establish a state transition model to describe the transition relationships between different states.
[0044] The state transition model describes the dynamic process of a system transitioning from the current state St to the next state St+1. In this embodiment, the state transition model includes the following parts:
[0045] Step 1021: Describe the state transition model of velocity V. The state transition of velocity V can be represented as:
[0046] Vt+1=Vt+V't×Δt+0.5×at×Δt 2
[0047] Where Vt represents the velocity at time t, V't represents the rate of change of velocity at time t, at represents the acceleration at time t, and Δt represents the time step, typically taken as 0.1 seconds. The acceleration at is determined by the control strategy and ranges from -5 to 5 m / min².
[0048] Step 1022: Describe the state transition model of tension T. The state transition of tension T can be expressed as:
[0049] Tt+1=Tt+ΔTt
[0050] Where Tt represents the tension at time t, and ΔTt represents the change in tension, which is determined by the control strategy and ranges from -2 to 2 kN.
[0051] Step 1023: Describe the state transition model of surface quality Q. The state transition of surface quality Q can be expressed as:
[0052] Qt+1=Qt-α×|Vt-Vopt|-β×|Tt-Topt|-γ×|V't|
[0053] Where Qt represents the surface quality score at time t, Vopt represents the optimal velocity, Topt represents the optimal tension, and α, β, and γ are the influence coefficients of velocity deviation, tension deviation, and velocity change rate on surface quality, respectively, with values of 0.05, 0.08, and 0.03.
[0054] Step 1024: Describe the state transition model of the winding uniformity U. The state transition of the winding uniformity U can be expressed as:
[0055] Ut+1=Ut-δ×|Tt-Topt|-ε×|V't|
[0056] Where Ut represents the winding uniformity score at time t, and δ and ε are the influence coefficients of tension deviation and speed change rate on winding uniformity, respectively, with values of 0.1 and 0.05.
[0057] Step 1025: Describe the state transition model of the running speed V'. The state transition of the running speed V' can be represented as:
[0058] V't+1=V't+at×Δt
[0059] Where V't represents the rate of change of velocity at time t, and at represents the acceleration at time t, which is determined by the control strategy.
[0060] Step 103: Design the reward function.
[0061] The reward function is a core component of reinforcement learning, used to evaluate the effectiveness of the control strategy. In this embodiment, a three-layer reward function is designed, each targeting a different control objective at a different stage.
[0062] Step 104: The first-level reward function uses the opposite of the absolute value of the angle between the vertical rod of the inverted pendulum and the vertical direction as the angle reward.
[0063] Step 1041: Calculate the absolute value of the angle between the inverted pendulum rod and the vertical direction. In the stainless steel rewinding unit, the deviation angle θ of the steel strip can be regarded as the angle between the inverted pendulum rod and the vertical direction. The calculation formula is:
[0064] θ = arctan((P2-P1) / L)
[0065] Where P1 and P2 are the position coordinates of the two ends of the steel strip, and L is the distance between the two points.
[0066] Step 1042: Use the opposite of the absolute value as the first-level reward. The first-level reward function R1 can be expressed as:
[0067] R1=-|θ|
[0068] The reward function is designed to keep the steel strip vertical and reduce the deviation angle, thereby improving the winding quality.
[0069] Step 105: Second-level reward function. When the distance 0 < d < 0.05, add 3 to the current reward to improve the balance control accuracy.
[0070] Step 1051: Calculate the deviation d between the current distance and the target value. The deviation d can be expressed as:
[0071] d = |Vt - Vopt| + λ × |Tt - Topt|
[0072] Wherein, λ is a weighting coefficient with a value of 0.2, used to balance the effects of speed deviation and tension deviation.
[0073] Step 1052: When 0 < d < 0.05, the deviation is used as the second-level reward. The second-level reward function R2 can be expressed as:
[0074] R2 = R1 + 3, when 0 < d < 0.05
[0075] R2 = R1, other cases
[0076] The purpose of this reward function is to further improve the control accuracy of speed and tension while keeping the steel strip vertical.
[0077] Step 106: The third-level reward function: when the distance d is stable within 0.05m, the second-level reward is used as the later reward.
[0078] Step 1061: Calculate the deviation d from the current distance to the target value. The method for calculating the deviation d is the same as in step 1051.
[0079] Step 1062: When the distance d is stable within 0.05m, the second-level reward is used as the third-level reward.
[0080] The third-level reward function R3 can be expressed as:
[0081] R3 = R2, when d < 0.05 for N consecutive time steps.
[0082] R3 = R1, other cases
[0083] Where N is the settling time threshold, with a value of 50, equivalent to a settling time of 5 seconds. The purpose of this reward function is to encourage the system to remain in a stable state for a longer period of time, thereby improving the consistency of product quality.
[0084] Step 2: Design control strategies.
[0085] Step 201: Initialize the deep Q-network and use experience replay and target network techniques to improve learning efficiency and stability.
[0086] In this embodiment, the deep Q-network adopts a multilayer perceptron structure, including an input layer, a hidden layer, and an output layer. The input layer has 5 nodes, corresponding to 5 state variables; the hidden layer adopts a 3-layer structure, with 64, 128, and 64 nodes in each layer, respectively; the number of nodes in the output layer is the dimension of the action space, corresponding to different combinations of control strategies.
[0087] The network parameters were initialized using the Xavier initialization method, and the ReLU activation function was used. The capacity of the empirical replay pool was set to 10,000, and the batch size for each training iteration was 64. The target network was updated every 100 time steps using a soft update method with an update coefficient τ = 0.01.
[0088] Step 202: Input the states such as velocity, tension, and surface quality into the deep Q-network and output the optimal control strategy.
[0089] The current state St = {Vt, Tt, Qt, Ut, V't} is input into a deep Q-network, which outputs the Q-values for each possible action. The action space A includes velocity control action av and tension control action aT, i.e., A = {av, aT}. The velocity control action av has a value range of [-5, 5], discretely represented by 11 values {-5, -4, -3, -2, -1, 0, 1, 2, 3, 4, 5}; the tension control action aT has a value range of [-2, 2], discretely represented by 9 values {-2, -1.5, -1, -0.5, 0, 0.5, 1, 1.5, 2}. Therefore, the dimension of the action space is 11 × 9 = 99.
[0090] According to the ε-greedy strategy, an action is randomly selected with probability ε, and the action with the largest Q value is selected with probability 1-ε. The initial value of ε is 1, and it gradually decays to 0.1 as training progresses, with a decay rate of 0.995.
[0091] Step 203: Execute corresponding speed and tension control according to the adopted control strategy, and feed the results back to the deep Q network in real time for learning.
[0092] Based on the selected motion {av, aT}, execute the corresponding speed and tension control:
[0093] at=av
[0094] ΔTt=aT
[0095] Then, the next state St+1 is calculated according to the state transition model, and the reward rt is calculated according to the reward function. The transition sample (St,at,rt,St+1) is stored in the experience replay pool.
[0096] A batch of samples is randomly selected from the experience replay pool for training, updating the parameters of the deep Q-network. Training employs a temporal difference learning method, with the loss function being:
[0097] L=E[(rt+γ×max(Q'(St+1,a'))-Q(St,at)) 2 ]
[0098] Where Q represents a deep Q-network, Q' represents the target network, and γ is a discount factor with a value of 0.99.
[0099] Step 3: Determine if the convergence condition has been met. If yes, proceed to Step 4; otherwise, return to Step 2 to continue learning.
[0100] The convergence criteria are set as follows: the average reward exceeds a threshold of -10 for 100 consecutive rounds, or the number of training rounds reaches the maximum value of 10,000. Each round has a length of 1,000 time steps, equivalent to a 100-second control process.
[0101] Step 4: Output the optimal control strategy.
[0102] Step 401: Save the learned optimal strategy to the strategy library.
[0103] The trained deep Q-network model is saved to a policy library, including information such as network structure, parameters, and hyperparameters. Simultaneously, performance metrics during training, such as average reward and success rate, are recorded for subsequent evaluation and comparison.
[0104] The policy library adopts a hierarchical structure, categorizing and storing policies according to different operating conditions. Each policy is associated with an expiration period, initially 30 days, which can be extended or shortened based on actual application results.
[0105] Step 402: Select the most suitable control strategy from the strategy library based on the current operating conditions.
[0106] Based on the characteristics of the current working condition, such as the material, thickness, and width of the steel strip, the similarity with each strategy in the strategy library is calculated. The similarity calculation uses weighted Euclidean distance.
[0107] d(x,y)=√(Σwi×(xi-yi) 2 )
[0108] Where wi is the weight of each feature, set according to importance: material weight is 0.4, thickness weight is 0.3, width weight is 0.2, and other feature weights are 0.1.
[0109] The policy with the highest similarity is selected as the current control policy. If the similarity is below the threshold of 0.8, the policy needs to be retrained or fine-tuned.
[0110] Step 403: Perform noise simulation on the strategy to improve its reliability.
[0111] To improve the robustness and adaptability of the policy, random noise is added to the selected policy. The noise follows a Gaussian distribution with a mean of 0, and the standard deviation σ is dynamically adjusted with the confidence level of the policy.
[0112] σ=σ0×(1-c)
[0113] Where σ0 is the basic noise level, with a value of 0.1, and c is the confidence level of the strategy, with a value range of [0,1].
[0114] Noise is added to the policy output to generate the final control signal:
[0115] a' = a + N(0,σ)
[0116] Where a represents the action output by the original control strategy, a' represents the action after adding noise, and N(0,σ) represents Gaussian noise with a mean of 0 and a standard deviation of σ.
[0117] Through the above steps, coupled control of speed and tension in the stainless steel rewinding unit is achieved, effectively improving the surface quality and winding uniformity of the product. This method has advantages such as strong adaptability, high control precision, and good robustness, and is suitable for various control scenarios of stainless steel rewinding units.
[0118] Step 1022: Describe the state transition model of tension T.
[0119] Step 10221: Describe the basic transfer relationship of tension T at different speeds. The basic transfer relationship can be expressed as:
[0120] Tt+1_base=Tt+ΔTt+kt×(Vt+1-Vt)+εT
[0121] Where Tt represents the tension at time t, ΔTt represents the direct control change of tension, kt represents the influence coefficient of velocity change on tension, with a value of 0.08 kN / (m / min), and εT represents the environmental noise, which follows a Gaussian distribution with a mean of 0 and a standard deviation of 0.3.
[0122] Step 10222: Considering the factors affecting tension T under different working conditions, establish a dynamic model affecting tension T. The dynamic model can be expressed as:
[0123] Tt+1=Tt+1_base×(1+f(Mt,Wt,Dt))
[0124] Where f(Mt,Wt,Dt) represents the influence function of material Mt, width Wt, and diameter Dt on tension, which can be expressed as:
[0125] f(Mt,Wt,Dt)=αM×(Mt-M0)+αW×(Wt-W0)+αD×(Dt-D0)
[0126] Where M0, W0, and D0 are the standard material, width, and diameter, respectively, and αM, αW, and αD are the influence coefficients of the material, width, and diameter, respectively, with values of 0.02, 0.01, and 0.015.
[0127] Step 1023: Describe the state transition model of surface quality Q.
[0128] Step 10231: Describe the relationship between surface quality Q, velocity V, and tension T. The relationship model can be expressed as:
[0129] Qt+1_base=Qt-α1×|Vt-Vopt|-α2×|Tt-Topt|-α3×|V't|-α4×|Tt+1-Tt|+εQ
[0130] Where Qt represents the surface quality score at time t, Vopt represents the optimal velocity, Topt represents the optimal tension, α1, α2, α3, and α4 are the influence coefficients of velocity deviation, tension deviation, velocity change rate, and tension change rate on surface quality, respectively, with values of 0.04, 0.06, 0.03, and 0.05, and εQ represents the environmental noise, which follows a Gaussian distribution with a mean of 0 and a standard deviation of 0.2.
[0131] Step 10232: Establish a prediction model for surface quality Q, including sub-models for substrate quality, scratch damage, and ripple defects. The prediction model can be expressed as:
[0132] Qt+1=w1×Qt+1_base+w2×Qbase+w3×Qscratch+w4×Qwrinkle
[0133] In this model, Qbase, Qscratch, and Qwrinkle represent the sub-models for substrate quality, scratch damage, and ripple defects, respectively. w1, w2, w3, and w4 are the weights of each sub-model, with values of 0.4, 0.2, 0.2, and 0.2, respectively.
[0134] Step 1024: Describe the state transition model of the winding uniformity U.
[0135] Step 10241: Describe the relationship between winding uniformity U and speed V and tension T. The relationship model can be expressed as:
[0136] Ut+1_base=Ut-β1×|Tt-Topt|-β2×|V't|-β3×|Tt+1-Tt|+εU
[0137] Where Ut represents the winding uniformity score at time t, β1, β2, and β3 are the influence coefficients of tension deviation, speed change rate, and tension change rate on winding uniformity, respectively, with values of 0.1, 0.05, and 0.08, respectively, and εU represents the environmental noise, which follows a Gaussian distribution with a mean of 0 and a standard deviation of 0.25.
[0138] Step 10242: Establish an evaluation model for winding uniformity, including sub-evaluation indicators such as thickness unevenness, thinner edges, and thicker center. The evaluation model can be expressed as:
[0139] Ut+1=w5×Ut+1_base+w6×Uthickness+w7×Uedge+w8×Ucenter
[0140] Among them, Uthickness, Uedge, and Ucenter represent the sub-evaluation indicators of uneven thickness, thinner edge, and thicker center, respectively. w5, w6, w7, and w8 are the weights of each sub-indicator, with values of 0.4, 0.2, 0.2, and 0.2, respectively.
[0141] Step 1025: Describe the state transition model of the running speed V'.
[0142] Step 10251: Describe the basic dynamic model of velocity V'. The basic dynamic model can be represented as:
[0143] V't+1_base=V't+at×Δt-μ×V't+εV'
[0144] Where V't represents the rate of change of velocity at time t, at represents the acceleration at time t, μ represents the damping coefficient with a value of 0.12, and εV' represents the ambient noise, which follows a Gaussian distribution with a mean of 0 and a standard deviation of 0.4.
[0145] Step 10252: Considering the impact of different operating conditions on velocity V', establish a prediction model for velocity V'. The prediction model can be expressed as:
[0146] V't+1=V't+1_base×(1+g(Mt,Wt,Dt))
[0147] Where g(Mt,Wt,Dt) represents the influence function of material Mt, width Wt, and diameter Dt on the rate of change of velocity, which can be expressed as:
[0148] g(Mt,Wt,Dt)=βM×(Mt-M0)+βW×(Wt-W0)+βD×(Dt-D0)
[0149] Wherein, βM, βW and βD are the influence coefficients of material, width and diameter, respectively, with values of 0.015, 0.008 and 0.01.
[0150] Step 1052: When 0 < d < 0.05, the deviation is used as the second-level reward.
[0151] Step 10521: Calculate the deviation between the current surface quality Q and the target value. The deviation dQ can be expressed as:
[0152] dQ = |Qt - Qopt|
[0153] Where Qt represents the surface quality score at time t, and Qopt represents the target surface quality, with a value of 95.
[0154] Step 10522: Use the deviation as the second-level reward to achieve precise control over surface quality. The second-level reward function R2 can be expressed as:
[0155] R2 = R1 + 3 - λQ × dQ, when 0 < d < 0.05
[0156] R2 = R1 - d - λQ × dQ, when d ≥ 0.05
[0157] Wherein, λQ is the weighting coefficient of surface quality deviation, with a value of 0.2.
[0158] Step 1062: When the distance d is stable within 0.05m, the second-level reward is used as the third-level reward.
[0159] Step 10621: Calculate the deviation between the current winding uniformity U and the target value. The deviation dU can be expressed as:
[0160] dU=|Ut-Uopt|
[0161] Where Ut represents the roll uniformity score at time t, and Uopt represents the target roll uniformity, with a value of 90.
[0162] Step 10622: Use the deviation as the third-level reward to optimize the convolution uniformity. The third-level reward function R3 can be expressed as:
[0163] R3 = R2 + η × min(N / Nmax,1) - λU × dU, where d < 0.05 for N consecutive time steps.
[0164] R3 = R2 - λU × dU, other cases
[0165] Where N is the number of time steps that have been stabilized, Nmax is the maximum number of time steps that have been stabilized (value is 120), η is the stability reward coefficient (value is 6), and λU is the weighting coefficient for convolution uniformity deviation (value is 0.15).
[0166] Step 1063: Specific implementation of the third-level reward function.
[0167] Step 10631: Calculate the deviation between the current velocity V and the target value. The deviation dV can be expressed as:
[0168] dV = |Vt - Vopt|
[0169] Where Vt represents the velocity at time t, and Vopt represents the target velocity, which is dynamically set according to different working conditions.
[0170] Step 10632: Use the deviation as the third-level reward to achieve precise control over the running speed. The third-level reward function R3 can be further expressed as:
[0171] R3=R3-λV×dV
[0172] Wherein, λV is the weighting coefficient of the speed deviation, with a value of 0.25.
[0173] Step 1064: Specific implementation of the third-level reward function.
[0174] Step 10641: Calculate the deviation between the current tension T and the target value. The deviation dT can be expressed as:
[0175] dT = |Tt - Topt|
[0176] Where Tt represents the tension at time t, and Topt represents the target tension, which is dynamically set according to different working conditions.
[0177] Step 10642: Use the deviation as the third-level reward to achieve precise control of tension. The third-level reward function R3 can be further expressed as:
[0178] R3=R3-λT×dT
[0179] Wherein, λT is the weighting coefficient of tension deviation, with a value of 0.3.
[0180] Step 1065: Specific implementation of the third-level reward function.
[0181] Step 10651: Calculate the combined deviations of the current surface quality Q, winding uniformity U, speed V, and tension T from the target values. The combined deviation dC can be expressed as:
[0182] dC=wQ×dQ+wU×dU+wV×dV+wT×dT
[0183] Among them, wQ, wU, wV and wT are the weighting coefficients for surface quality, winding uniformity, speed and tension, respectively, with values of 0.4, 0.3, 0.15 and 0.15.
[0184] Step 10652: Use the overall deviation as the third-level reward to optimize the global metric. The third-level reward function R3 can be ultimately expressed as:
[0185] R3=R3-λC×dC
[0186] Wherein, λC is the weighting coefficient of the overall deviation, with a value of 0.5.
[0187] Step 201: Initialize the deep Q-network and use experience replay and target network techniques to improve learning efficiency and stability.
[0188] Step 2011: Set up the network structure of the deep Q-network, including an input layer, hidden layers, and an output layer. The input layer has 5 nodes, corresponding to 5 state variables; the hidden layer has a 5-layer structure, with 64, 128, 256, 128, and 64 nodes in each layer, respectively; the number of nodes in the output layer is the dimension of the action space, corresponding to different combinations of control strategies.
[0189] Step 2012: Initialize network parameters and store training data using the experience replay pool.
[0190] Step 20121: Set the basic parameters of the deep Q-network, such as learning rate, batch size, and exploration rate. The initial learning rate is 0.001, which is dynamically adjusted using a cosine annealing strategy; the batch size is 256; the initial exploration rate is 1, the minimum is 0.02, and the decay rate is 0.999.
[0191] Step 20122: Initialize the memory to store training data and target values. The capacity of the experience replay pool is set to 50,000, and a priority experience replay mechanism is adopted, assigning priority to samples according to the magnitude of the TD error.
[0192] Step 20123: Set the network update rules, including the update frequency and method of the target network. The target network is updated once every 300 time steps, using a soft update method, with an update coefficient τ = 0.002.
[0193] Step 2013: Set up the target network to assist in the training and updating of the network.
[0194] Step 20131: Set the target network structure to be consistent with the deep Q-network. The target network and the deep Q-network have the same network structure, including the same number of layers and nodes.
[0195] Step 20132: Initialize the parameters of the target network using a low-entropy random initialization method. The initialization method uses a truncated normal distribution with a mean of 0, a standard deviation of 0.03, and a truncation range of [-0.05, 0.05].
[0196] Step 20133: Set the update rules for the target network, including the update frequency and weights. The update rules for the target network parameter θ' are as follows:
[0197] θ'=τ×θ+(1-τ)×θ'
[0198] Where θ is the parameter of the deep Q-network, and τ is the soft update coefficient with a value of 0.002.
[0199] Step 202: Input the states such as velocity, tension, and surface quality into the deep Q-network and output the optimal control strategy.
[0200] Step 2021: Input the states such as velocity, tension, surface quality, winding uniformity, and running speed into the deep Q-network. The input data is first normalized to map each state variable to the range [-1, 1] to improve the training efficiency and generalization ability of the network.
[0201] Step 2022: Calculate the Q-values of each control strategy using a deep Q-network.
[0202] Step 20221: Preprocess the input data, including normalization and feature extraction. Normalization uses the Min-Max method, and feature extraction includes calculating the first-order differences and interaction terms of the state variables.
[0203] Step 20222: Input the preprocessed data into a deep Q-network to obtain the Q-values of each control strategy. During the forward propagation of the network, a batch normalization layer and a Dropout layer are added after each hidden layer, with a Dropout rate of 0.2, to improve the network's generalization ability.
[0204] Step 20223: Use the dual-threshold Q-learning algorithm to select the optimal strategy.
[0205] Step 202231: Set the threshold size and update strategy. The initial value of the upper threshold Qhigh is 10, and the initial value of the lower threshold Qlow is -10. These values are dynamically adjusted as training progresses, following the following adjustment rules:
[0206] Qhigh = max(Qhigh, Q95)
[0207] Qlow = min(Qlow, Q05)
[0208] Q95 and Q05 are the 95th and 5th percentiles of the Q value for the current batch, respectively.
[0209] Step 202232: Select the optimal strategy based on the threshold, including strategy selection and strategy update. The strategy selection rule is as follows:
[0210] a*=argmax(Q(s,a)), when Q(s,a)>Qlow
[0211] a*=random(A), when Q(s,a)≤Qlow
[0212] Where a* represents the selected action, and A represents the action space. The policy update rule is:
[0213] Q(s,a)=Q(s,a), when Q(s,a) <Qhigh
[0214] Q(s,a)=Qhigh, when Q(s,a)≥Qhigh
[0215] Step 2023: Select the control strategy that maximizes the Q value as the optimal strategy. In practical applications, an ε-greedy strategy is adopted, which selects the action with the largest Q value with probability 1-ε and randomly selects an action with probability ε to balance exploration and exploitation.
[0216] Step 401: Save the learned optimal strategy to the strategy library.
[0217] Step 4011: Associate and store the control strategy with the current operating conditions.
[0218] Step 40111: Associate and store the control strategy with the main characteristics of the current operating condition. The operating condition characteristics include steel strip material, thickness, width, hardness, surface treatment type, etc., which are represented in vector form and stored in the strategy library together with the control strategy.
[0219] Step 40112: Set the validity period, including the initial validity period and renewal mechanism. The initial validity period is set to 60 days, which can be extended or shortened based on the actual application effect of the strategy. The renewal mechanism is based on the performance evaluation of the strategy; if the performance is good, the validity period will be automatically extended by 30 days.
[0220] Step 40113: Establish a policy decay mechanism to achieve dynamic policy updates. The policy weight w decays with time t:
[0221] w(t)=w0×exp(-λ×t)
[0222] Where w0 is the initial weight, with a value of 1, and λ is the decay coefficient, with a value of 0.01 / day.
[0223] Step 4012: Set the validity period of the strategy to achieve dynamic updates. The validity period of the strategy is dynamically adjusted based on its performance. Performance evaluation indicators include surface quality improvement rate, winding uniformity improvement rate, and production efficiency improvement rate.
[0224] Step 402: Select the most suitable control strategy from the strategy library based on the current operating conditions.
[0225] Step 4021: Calculate the similarity between the current working condition and the training samples.
[0226] Step 40211: Extract the main features of the current operating condition, including speed, tension, surface quality, etc. Principal component analysis is used for feature extraction, retaining principal components that explain 90% of the variance.
[0227] Step 40212: Retrieve the historical sample most similar to the current working condition from the database. Cosine similarity is used for similarity calculation.
[0228] sim(x,y)=(x·y) / (||x||×||y||)
[0229] Where x and y are the feature vectors of the current working condition and the historical sample, respectively, · represents the inner product of vectors, and ||x|| represents the L2 norm of vector x.
[0230] Step 40213: Calculate the similarity score and select the most similar sample. The similarity score also considers the time decay factor of the samples; newer samples have higher weights.
[0231] score(x,y,t)=sim(x,y)×exp(-μ×t)
[0232] Where t is the age of the sample (in days), and μ is the time decay coefficient, with a value of 0.005 / day.
[0233] Step 4022: Select the historical sample that is most similar to the current working condition.
[0234] Step 40221: Retrieve samples from the database whose similarity to the current working condition exceeds a threshold. The similarity threshold is set to 0.85; only samples with a similarity exceeding this threshold will be considered.
[0235] Step 40222: Select a certain number of similar samples based on the sampling rate. The sampling rate is set to 10%, and 10% of the similar samples are randomly selected for subsequent processing.
[0236] Step 40223: Preprocess the sampled data to prepare for the next step. Preprocessing includes feature standardization, outlier detection, and missing value handling.
[0237] Step 4023: Select the optimal control strategy. The selection of the optimal strategy is based on a weighted voting mechanism, where the voting weight of each similar sample is proportional to its similarity.
[0238] Step 403: Perform noise simulation on the strategy to improve its reliability.
[0239] Step 4031: Generate random noise that conforms to the probability distribution.
[0240] Step 40311: Set the probability distribution parameters for the noise. The noise adopts a Gaussian mixture distribution, which includes two Gaussian components with weights of 0.7 and 0.3, respectively, and means of 0 and 0, respectively, and standard deviations of 0.1 and 0.3, respectively.
[0241] Step 40312: Generate a random number sequence that conforms to a probability distribution. The length of the random number sequence is the same as the dimension of the action space, and each element follows the above-described mixture Gaussian distribution.
[0242] Step 40313: Convert the random number sequence into the desired noise signal. The amplitude of the noise signal gradually decreases as training progresses, following the attenuation rule:
[0243] σ(t)=σ0×exp(-ν×t)
[0244] Where σ0 is the initial noise amplitude, with a value of 0.2, and ν is the attenuation coefficient, with a value of 0.0001 / step.
[0245] Step 4032: Combine noise with the optimal strategy to generate diverse control strategies.
[0246] Step 40321: Perform element-wise operations on the noise signal and the optimal policy. The operation is addition.
[0247] a' = a + noise
[0248] Where a represents the action output by the original control strategy, a' represents the action after adding noise, and noise represents the generated noise signal.
[0249] Step 40322: Generate diverse control policy samples. To increase policy diversity, generate multiple noisy policy samples and select the optimal one from them.
[0250] Step 40323: Input the samples into the deep Q-network for training to diversify the strategies. Diverse strategy samples are used to enhance the network's generalization ability and avoid overfitting to specific working conditions.
[0251] Through the above steps, coupled control of speed and tension in the stainless steel rewinding unit is achieved, effectively improving the surface quality and winding uniformity of the product. This method has advantages such as strong adaptability, high control precision, and good robustness, and is suitable for various control scenarios of stainless steel rewinding units.
[0252] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for controlling surface quality by coupling speed and tension in a stainless steel rewinding unit, characterized in that, The method includes, Step 1: Establish a Markov decision process model for the stainless steel rewinding unit, including: Step 101: Define state variables, including speed, tension, surface quality, winding uniformity, and running speed parameters; Step 102: Establish a state transition model to describe the transition relationships between different states; Step 103: Design the reward function, including: Step 104: The first-level reward function uses the opposite of the absolute value of the angle between the vertical rod of the inverted pendulum and the vertical direction as the angle reward. Step 105: Second-level reward function: When the distance 0 < d < 0.05, add 3 to the current reward to improve the balance control accuracy; Step 106: The third-level reward function: when the distance d is stable within 0.05m, the second-level reward is used as the reward for the later stage. Step 2: Design control strategies, including: Step 201: Initialize the deep Q-network and use experience replay and target network techniques to improve learning efficiency and stability; Step 202: Input velocity, tension, and surface quality status into the deep Q-network and output the optimal control strategy; Step 203: Execute corresponding speed and tension control according to the adopted control strategy, and feed the feedback to the deep Q network in real time for learning; Step 3: Determine if the convergence condition has been met. If yes, proceed to Step 4; otherwise, return to Step 2 to continue learning. Step 4: Output the optimal control strategy, including: Step 401: Save the learned optimal policy to the policy library; Step 402: Select the most suitable control strategy from the strategy library based on the current operating conditions; Step 403: Perform noise simulation on the strategy to improve its reliability.
2. The method for controlling surface quality of a stainless steel rewinding unit by speed and tension coupling according to claim 1, characterized in that, In step 101, the state variables specifically include: Step 1011, Speed V, is used to indicate the operating speed of the stainless steel rewinding unit; Step 1012, Tension T, is used to represent the tension output of the stainless steel rewinding unit; Step 1013, Surface quality Q, is used to represent the surface quality score of the final product; Step 1014, winding uniformity U, an index used to represent winding uniformity; Step 1015, running speed V', is used to represent the rate of change of the current running speed; In step 102, the establishment of the state transition model specifically includes: Step 1021: Describe the state transition model for velocity V; Step 1022: Describe the state transition model of tension T; Step 1023: Describe the state transition model of surface quality Q; Step 1024: Describe the state transition model of the winding uniformity U; Step 1025: Describe the state transition model for the running speed V'; In step 104, the specific implementation of the first-level reward function includes: Step 1041: Calculate the absolute value of the angle between the vertical rod of the inverted pendulum and the vertical direction; Step 1042: Use the opposite of the absolute value as the first-level reward; In step 105, the specific implementation of the second-level reward function includes: Step 1051: Calculate the deviation d between the current distance and the target value; Step 1052: When 0 < d < 0.05, the deviation is used as the second-level reward; In step 106, the specific implementation of the third-level reward function includes: Step 1061: Calculate the deviation d between the current distance and the target value; Step 1062: When the distance d is stable within 0.05m, the second-level reward is used as the third-level reward.
3. The method for controlling surface quality of a stainless steel rewinding unit by speed and tension coupling according to claim 1, characterized in that, In step 201, the initialization of the deep Q-network includes: Step 2011: Set the network structure of the deep Q network, including the input layer, hidden layer and output layer; Step 2012: Initialize network parameters and store training data using the experience replay pool; Step 2013: Set up the target network to assist in the training and updating of the network; In step 202, the output of the control strategy includes: Step 2021: Input speed, tension, surface quality, winding uniformity, and running speed status into the deep Q-network; Step 2022: Calculate the Q-values of each control strategy using a deep Q-network; Step 2023: Select the control strategy that maximizes the Q value as the optimal strategy.
4. The method for controlling surface quality of a stainless steel rewinding unit by speed and tension coupling according to claim 1, characterized in that, In step 401, policy saving includes: Step 4011: Associate and store the control strategy with the current operating conditions; Step 4012: Set the validity period of the strategy to enable dynamic updates of the strategy; In step 402, strategy selection includes: Step 4021: Calculate the similarity between the current working condition and the training samples; Step 4022: Select the historical sample that is most similar to the current operating condition; Step 4023: Select the optimal control strategy from among them; In step 403, the noise simulation includes: Step 4031: Generate random noise that conforms to the probability distribution; Step 4032: Combine noise with the optimal strategy to generate diverse control strategies.
5. The method for controlling surface quality of a stainless steel rewinding unit by speed and tension coupling according to claim 2, characterized in that, In step 1022, the state transition model of tension T includes: Step 10221: Describe the basic transfer relationship of tension T at different speeds; Step 10222: Consider the factors affecting tension T under different working conditions and establish a dynamic model affecting tension T; In step 1023, the state transition model for surface quality Q includes: Step 10231: Describe the relationship between surface quality Q, velocity V, and tension T; Step 10232: Establish a prediction model for surface quality Q, including sub-models for substrate quality, scratch damage, and ripple defects; In step 1024, the state transition model for the winding uniformity U includes: Step 10241: Describe the relationship between winding uniformity U and speed V and tension T; Step 10242: Establish an evaluation model for winding uniformity, including evaluation indicators for thickness unevenness, thinner edges, and thicker centers; In step 1025, the state transition model for the running speed V' includes: Step 10251: Describe the basic dynamic model of velocity V'; Step 10252: Consider the impact of different working conditions on speed V' and establish a prediction model for speed V'; In step 1052, the specific implementation of the second-level reward function includes: Step 10521: Calculate the deviation between the current surface quality Q and the target value; Step 10522: Use the deviation as the second-level reward to achieve precise control over surface quality; In step 1062, the specific implementation of the third-level reward function includes: Step 10621: Calculate the deviation between the current winding uniformity U and the target value; Step 10622: Use the deviation as the third-level reward to optimize the winding uniformity.
6. The method for controlling surface quality of a stainless steel rewinding unit by speed and tension coupling according to claim 3, characterized in that, In step 2012, network parameter initialization includes: Step 20121: Set the basic parameters of the deep Q-network, including learning rate, batch size, and exploration rate; Step 20122: Initialize the memory to store training data and target values; Step 20123: Set the network update rules, including the update frequency and method of the target network; In step 2013, the target network initialization includes: Step 20131: Set the network structure of the target network to be consistent with that of the deep Q-network; Step 20132: Initialize the parameters of the target network using a low-entropy random initialization method; Step 20133: Set the update rules for the target network, including update frequency and weights; In step 2022, the calculation of the Q-value of the control strategy includes: Step 20221: Preprocess the input data, including normalization and feature extraction; Step 20222: Input the preprocessed data into a deep Q-network to obtain the Q-values of each control strategy; Step 20223: Use the dual-threshold Q-learning algorithm to select the optimal strategy.
7. The method for controlling surface quality of a stainless steel rewinding unit by speed and tension coupling according to claim 4, characterized in that, In step 4011, the policy association storage includes: Step 40111: Associate and store the control strategy with the main characteristics of the current operating condition; Step 40112: Set the validity period, including the initial validity period and renewal mechanism; Step 40113: Establish a policy decay mechanism to achieve dynamic policy updates; In step 4032, the combination of noise and strategy includes: Step 40321: Perform element-wise operations on the noise signal and the optimal strategy; Step 40322: Generate diverse control strategy samples; Step 40323: Input the samples into the deep Q-network for training to achieve strategy diversification; In step 4021, the similarity calculation includes: Step 40211: Extract the main features of the current working condition, including speed, tension, and surface quality; Step 40212: Retrieve the historical sample most similar to the current operating condition from the database; Step 40213: Calculate the similarity score and select the most similar sample; In step 4022, the selection of historical samples includes: Step 40221: Retrieve samples from the database that have a similarity to the current working condition exceeding a threshold; Step 40222: Select a certain number of similar samples based on the sampling rate; Step 40223: Preprocess the sampled data to prepare for the next step; In step 4031, the random noise generation includes: Step 40311: Set the probability distribution parameters for the noise; Step 40312: Generate a random number sequence that conforms to the probability distribution; Step 40313: Convert the random number sequence into the desired noise signal.
8. The method for controlling surface quality of a stainless steel rewinding unit by speed and tension coupling according to claim 5, characterized in that, In step 10221, the basic transfer relationship of tension T includes: Step 102211: Establish the basic mapping relationship between velocity V and tension T; Step 102212: Considering factors such as inertia and damping, describe the dynamic change law of tension T; In step 10222, the dynamic model affecting the tension T includes: Step 102221: Establish a model relating velocity V, tension T, and surface quality Q; Step 102222: Considering factors such as winding uniformity U, establish a compensation model for tension T; In step 10231, the prediction model for surface quality Q includes: Step 102311: Establish a substrate quality prediction model, including surface roughness and thickness indices; Step 102312: Establish a scratch damage prediction model, including parameters such as length and depth; Step 102313: Establish a ripple defect prediction model, including type and depth features; In step 10241, the evaluation model for the winding uniformity U includes: Step 102411: Establish a thickness non-uniformity evaluation model, including maximum thickness difference and average thickness index; Step 102412: Establish an evaluation model for thinning at the edges, including the degree of thinning and range parameters; Step 102413: Establish a center thickness evaluation model, including thickness and uniformity characteristics; In step 10251, the prediction model for velocity V' includes: Step 102511: Establish the basic dynamic model of velocity V', including the division of acceleration and deceleration phases; Step 102512: Considering factors such as changes in operating conditions, establish a prediction curve for velocity V'.
9. The method for controlling surface quality of a stainless steel rewinding unit by speed and tension coupling according to claim 6, characterized in that, In step 20122, memory initialization includes: Step 201221: Set the memory capacity and data structure; Step 201222: Initialize the size and update strategy of the experience replay pool; Step 201223: Set the update frequency and method for the target network; In step 20132, the target network initialization includes: Step 201321: Set the network structure of the target network to be consistent with that of the deep Q-network; Step 201322: Initialize the parameters of the target network using a low-entropy random initialization method; Step 201323: Set the update rules for the target network, including update frequency and weights; In step 20223, the dual-threshold Q-learning algorithm includes: Step 202231: Set the threshold size and update strategy; Step 202232: Select the optimal strategy based on the threshold, including strategy selection and strategy update.
10. The method for controlling surface quality of a stainless steel rewinding unit by speed and tension coupling according to claim 5, characterized in that, In step 10622, the specific implementation of the third-level reward function includes: Step 106221: Calculate the deviation between the current winding uniformity U and the target value; Step 106222: Use the deviation as the third-level reward to optimize the winding uniformity.