A multi-stage drawing control system and method for fine copper tube processing

By using a multi-stage drawing control system, combined with a multi-attribute coupling model and a Markov process decision framework, the nonlinear and time-varying characteristics of copper tubes in the multi-stage drawing process were solved, achieving high-precision and low-loss copper tube production.

CN122131599APending Publication Date: 2026-06-02青岛金泰宇铜业有限公司

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
青岛金泰宇铜业有限公司
Filing Date
2026-03-04
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Traditional multi-stage drawing control methods for copper tubes are difficult to cope with the nonlinear and time-varying characteristics caused by the coupling of multiple attributes such as differences in copper composition, mold wear, and changes in cooling conditions, resulting in problems such as dimensional deviations, surface defects, and abnormal equipment wear.

Method used

A multi-stage drawing control system is adopted, which establishes a multi-attribute coupling model through a framework module, a mapping module, and a control output module. Physical constraints and dynamic virtual impedance are introduced, and combined with a Markov process decision framework and deep learning algorithms, cross-track coordinated control is achieved.

Benefits of technology

It improves the precision and stability of the multi-stage drawing process of copper tubes, reduces energy consumption, extends equipment life, and enhances the robustness and precision of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122131599A_ABST
    Figure CN122131599A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-stage drawing control system and method for fine copper tube processing, relating to the field of copper tube processing control technology. The mapping module introduces dynamic virtual impedance into a Markov process decision framework, establishing a mapping relationship between multi-objective outputs and virtual parameters in the dynamic virtual impedance. Parameter calibration is performed for different sub-scenarios. The control output module utilizes an improved deep deterministic strategy gradient algorithm and a long short-term memory network to perform offline learning on historical multi-attribute process data, generating a transferable pre-trained parameter set and state prediction capability. The pre-trained parameter set and virtual parameters are injected into an online deep decision optimization unit to perform real-time rolling solution of the Markov process decision framework, outputting the current optimal control action. This control system, through multi-attribute deep coupling modeling combined with a Markov process decision framework, effectively improves the accuracy and stability of the multi-stage drawing process control for fine copper tubes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of copper tube processing control technology, specifically to a multi-stage drawing control system and method for processing fine copper tubes. Background Technology

[0002] In the multi-stage drawing process of fine copper tubes, it is necessary to precisely coordinate servo speed, tension, and lubrication between different passes to ensure dimensional accuracy, reduce energy consumption, and extend equipment life. However, traditional control methods are difficult to cope with the nonlinear and time-varying characteristics caused by the coupling of multiple attributes such as differences in copper composition, die wear, and changes in cooling conditions. Especially in high-speed continuous production, problems such as dimensional deviations, surface defects, and abnormal equipment wear are prone to occur. Therefore, there is an urgent need for a systematic solution that can integrate multi-attribute information of materials, dies, and equipment, embed physical constraints, and achieve cross-pass collaborative control through intelligent optimization to meet the needs of high-precision, low-loss, and intelligent fine copper tube drawing production. Summary of the Invention

[0003] The purpose of this invention is to provide a multi-stage drawing control system and method for processing fine copper tubes, so as to solve the problems in the prior art.

[0004] To achieve the above objectives, the present invention provides the following technical solution: a multi-stage drawing control system for processing thin copper tubes, comprising a frame establishment module, a mapping module, and a control output module: Framework establishment module: By weighted fusion and cross-coupling mapping of multi-source parameters, a multi-attribute coupling model is established to establish the relationship between the deformation rate and dimensional accuracy of each pass. After the physical conservation and mechanical equilibrium equations are embedded as constraints into the multi-attribute coupling model, the multi-stage drawing process is abstracted into a Markov process decision framework. Mapping module: Introducing dynamic virtual impedance into the Markov process decision framework, establishing the mapping relationship between multi-objective output and virtual parameters in dynamic virtual impedance, and using the Newton-Raphson joint iterative method to calibrate parameters for different sub-scenarios in the offline stage; Control output module: Utilizing an improved deep deterministic policy gradient algorithm and long short-term memory network, it performs offline learning on historical multi-attribute process data to generate a transferable pre-trained parameter set and state prediction capability. The pre-trained parameter set and virtual parameters are injected into the online deep decision optimization unit to perform real-time rolling solution of the Markov process decision framework and output the current optimal control action.

[0005] Preferably, the control output module injects the pre-trained parameter set and virtual parameters into the online deep decision optimization unit to perform real-time rolling solution of the Markov process decision framework and output the current optimal control action: The real-time status of the current process is input into the long short-term memory network to obtain the predicted value of the process status at the next moment. Target network A receives the predicted value of the process state and outputs the value assessment of the control action, which is then passed to target network C to calculate the optimization objective; The target network C iteratively updates the parameters of the target network A and the target network C based on the deviation between the actual quality feedback and the predicted target, and outputs the current optimal control action.

[0006] Preferably, the mapping module introduces dynamic virtual impedance in the Markov process decision framework. The dynamic virtual impedance includes virtual damping coefficient and virtual inertia coefficient. The virtual damping coefficient is the difference between the adaptive damping and the actual system friction damping, and the virtual inertia coefficient is the difference between the adaptive inertia and the actual traction wheel set and pipe mass distribution. The multi-objective output includes energy consumption reference value, dimensional accuracy reference value and equivalent control speed modulus value.

[0007] Preferably, the control output module utilizes an improved deep deterministic policy gradient algorithm and a long short-term memory network to perform offline learning on historical multi-attribute process data, generating a transferable pre-trained parameter set and state prediction capability. The network reads a batch of serialized data in the order of channels. The cell states and hidden states inside the long and short temporal memory network transmit and compress information from past channels. The attention mechanism layer performs importance evaluation and weighted aggregation of the LSTM hidden states of all past channels in parallel. Initialize target network A and target network B, and create a copy of the target network for each of target network A and target network B to stabilize the training process; After the target network A takes the state vector as input, it outputs a control action vector. The target network B takes the state and the corresponding control action generated by the target network A as input, and outputs a scalar value to evaluate the long-term expected benefit of performing the control action under the state vector.

[0008] Preferably, the control output module injects the pre-trained parameter set and virtual parameters into the online deep decision optimization unit to perform real-time rolling solution of the Markov process decision framework and output the current optimal control action: In each control cycle, real-time status data of the current pass is collected. After the real-time data is combined into the current status vector, it is sent into the decision process and input into the long short-term memory network. The long short-term memory network predicts the process status prediction value of the next moment. The predicted process status values ​​are input into the target network A in the online system. The target network A evaluates the predicted process status values ​​and outputs preliminary control action suggestions and their value assessment. The proposed control action is transmitted to the target network C, which calculates the target value, i.e., the expected Q value of the control action. The target network C obtains actual feedback from the quality inspection process and compares the actual quality feedback with the predicted target calculated based on the predicted state to calculate the deviation between the two. The target network C uses the bias to backpropagate, update its own network parameters, provide gradient information to the target network A, and drive the target network A to adjust its policy parameters. After several internal iterations and updates, the current target network A outputs the optimal control action under the current information conditions.

[0009] Preferably, the mapping module introduces dynamic virtual impedance into the Markov process decision framework to establish a mapping relationship between multi-objective outputs and virtual parameters in the dynamic virtual impedance: For each given pair of virtual parameters, simulations are performed on the Markov process decision framework, and the overall performance under disturbances is recorded to calculate the corresponding multi-objective output values. By employing multivariate function approximation techniques, we learn nonlinear mapping functions from the virtual parameter space to the multi-objective output space.

[0010] Preferably, the mapping module uses the Newton-Raphson joint iterative method to calibrate parameters for different sub-scenes in the offline phase: Set an ideal target vector containing all multi-target output reference values, initialize virtual parameter values, and enter the iteration loop: Call the mapping function to calculate the actual multi-target output vector corresponding to the current virtual parameters, and calculate the error vector between the actual multi-target output vector and the ideal target vector; Determine whether the magnitude of the error vector is less than the preset convergence threshold. If it is, the iteration terminates and the current parameter is the desired fitting parameter. If not, calculate the Jacobian matrix of the mapping function at the current parameter point. Incrementally correct the virtual parameters along the direction of the fastest error reduction, and repeat the iterative process using the corrected new parameter values ​​to find the globally optimal joint parameter solution.

[0011] Preferably, the framework building module establishes a multi-attribute coupled model for the relationship between the deformation rate and dimensional accuracy of multiple sources by weighted fusion and cross-coupling mapping of multi-source parameters: Each attribute and its sub-features are assigned a weight coefficient, and the multi-dimensional sub-features are aggregated into a comprehensive evaluation index for each attribute by weighted summation. By analyzing the interaction effects between different categories of attributes, using existing production data, and employing a nonlinear regression model in machine learning for pattern mining, the quantitative correlation between cross-coupling effects and the output variables of trace deformation rate and dimensional accuracy is extracted. Integrate all relationships to complete the construction of a multi-attribute coupling model.

[0012] Preferably, the multi-source parameters include the chemical composition and grain structure of the copper material, the initial size and hardness of the billet, the geometric parameters of the mold cavity, the viscosity and distribution of the lubricant, and the ambient temperature and cooling conditions.

[0013] This application also provides a multi-stage drawing control method for processing thin copper tubes, the control method including the following steps: By weighted fusion and cross-coupling mapping of multi-source parameters, a multi-attribute coupling model is established to establish the relationship between the deformation rate and dimensional accuracy of each pass. By introducing physical conservation and mechanical equilibrium equations as constraints and embedding them into a multi-attribute coupled model, the multi-stage drawing process is abstracted into a Markov process decision framework. In the Markov process decision framework, a dynamic virtual impedance is introduced to establish a mapping relationship between multi-objective output and virtual parameters in the dynamic virtual impedance. The Newton-Raphson joint iterative method was used to calibrate the parameters of different sub-scenes in the offline stage. By utilizing an improved deep deterministic policy gradient algorithm and a long short-term memory network with attention mechanism, we can perform offline learning on historical multi-attribute process data to generate a transferable pre-trained parameter set and state prediction capability. The pre-trained parameter set and virtual parameters are injected into the online deep decision optimization unit to perform real-time rolling solution of the Markov process decision framework and output the current optimal control action.

[0014] The technical effects and advantages provided by the present invention in the above technical solution are as follows: This application achieves accurate abstraction and global optimization of complex drawing processes by deeply coupling multiple attributes of materials, molds and equipment, and introducing physical conservation and mechanical equilibrium equations as hard constraints. Combined with the Markov process decision framework, it effectively improves the model's fit to real process laws and its generalization ability. This application introduces dynamic virtual impedance and performs offline multi-scenario parameter calibration, enabling the system to adaptively match the ideal dynamic response for different raw material hardness, mold wear and cooling conditions, which significantly enhances the robustness and accuracy of cross-working condition control. This application integrates an improved Markov process decision framework with a long short-term memory network, which can capture long-term time-series characteristics across multiple passes and predict potential risks in advance. It can also efficiently search for the optimal control strategy in the state-control action space, forming a transferable pre-trained parameter set and real-time prediction capability. This enables precise look-ahead control based on real-time state in online rolling optimization, ultimately achieving comprehensive improvements in reducing energy consumption, minimizing dimensional deviations, and delaying equipment wear. This effectively enhances the accuracy and stability of the multi-stage drawing process control for fine copper tubes. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on the drawings.

[0016] Figure 1 This is a framework diagram of the control system of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example: This example provides a multi-stage drawing control system for processing thin copper tubes. Please refer to [link / reference]. Figure 1 As shown, it includes a framework building module, a mapping module, and a control output module: Framework Establishment Module: For the multi-stage drawing process of fine copper tubes, attributes such as the chemical composition and grain structure of copper, the initial size and hardness of the billet, the geometric parameters of the die cavity, the viscosity and distribution of lubricant, and the ambient temperature and cooling conditions are mapped to the correlation between the deformation rate and dimensional accuracy per pass through weighted fusion and cross-coupling, establishing a multi-attribute coupled model of material-die-equipment. Physical conservation and mechanical equilibrium equations (such as the conservation of plastic deformation energy and the tension-velocity compatibility equation) are introduced as hard constraints embedded in the model, ensuring it both fits experimental data and conforms to the fundamental laws of metal forming. The entire multi-stage drawing process is abstracted into a Markov process decision framework (MDP). The system state is a vector of real-time dimensions, tension, temperature, lubrication status, etc. for each pass. Decision-making and control actions include servo speed, tension compensation, and lubrication dosage adjustment. State transitions are governed by a multi-attribute coupling model and physical constraints. The reward function is defined by the negative value of the comprehensive loss (energy consumption, dimensional deviation, equipment wear). The Markov process decision framework is sent to the mapping module and the control output module.

[0019] Mapping Module: A dynamic virtual impedance is introduced into the Markov process decision framework. Virtual damping coefficients (the difference between adaptive damping and actual system friction damping) and virtual inertia coefficients (the difference between adaptive inertia and the actual traction wheel assembly and pipe mass distribution) are defined. A mapping relationship is established between the system's multi-objective outputs (energy consumption reference value, dimensional accuracy reference value, and equivalent control speed modulus) and the virtual parameters. To solve for the adaptive parameters, the Newton-Raphson joint iterative method is used to calibrate parameters in sub-scenarios with different raw material hardness, mold wear, and cooling conditions during the offline stage. This ensures that the model can approximate the ideal dynamic response under various operating conditions. The virtual parameters are then sent to the control output module.

[0020] Output control module: Utilizing an improved Deep Deterministic Policy Gradient (DDPG) algorithm and an Att-LSTM network with attention mechanism, offline learning is performed on historical multi-attribute process data. Att-LSTM captures long-term time-series features across stages and predicts key parameters for the next state (such as the trend of the next stage's inlet diameter and potential instability risk). DDPG explores and approximates the optimal control strategy distribution in the state-control action space, generating a transferable pre-trained parameter set and state prediction capability. The obtained pre-trained parameter set and virtual parameters are injected into an online deep decision optimization unit to perform real-time rolling solutions on the Markov process decision framework. The real-time status of the current pass (diameter, wall thickness, tension, temperature, lubrication, etc.) is input into the Att-LSTM prediction network to obtain the predicted process status value for the next time step. The target network A (Actor) receives the predicted process status value and outputs the value assessment of the control action, which is then passed to the target network C (Critic) to calculate the optimization target. The target network C iteratively updates the parameters of the target network A and the target network C based on the deviation between the actual quality feedback (such as diameter tolerance, surface defect detection results) and the predicted target, and outputs the current optimal control action (speed curve, tension compensation, lubrication frequency).

[0021] This embodiment also provides a multi-stage drawing control method for processing thin copper tubes. Please refer to [link to relevant documentation]. Figure 1 As shown, the control method includes the following steps: For the multi-stage drawing process of fine copper tubes, the chemical composition and grain structure of copper, the initial size and hardness of the billet, the geometric parameters of the die cavity, the viscosity and distribution of lubricant, and the ambient temperature and cooling conditions are mapped to the correlation between the deformation rate and dimensional accuracy of each pass through weighted fusion and cross-coupling, establishing a multi-attribute coupled model of material-die-equipment. Physical conservation and mechanical equilibrium equations (such as the conservation of plastic deformation energy and the tension-velocity compatibility equation) are introduced as hard constraints embedded in the model, making it both fit experimental data and conform to the basic laws of metal forming.

[0022] The entire multi-stage drawing process is abstracted into a Markov process decision framework (MDP): The system state is a vector of real-time dimensions, tension, temperature, lubrication status, etc. for each pass. Decision-making and control actions include servo speed, tension compensation, and lubrication dosage adjustment. The state transition is governed by a multi-attribute coupling model and physical constraints. The reward function is defined by the negative value of the comprehensive loss (energy consumption, dimensional deviation, equipment wear).

[0023] Dynamic virtual impedance is introduced into the Markov process decision framework. Virtual damping coefficients (the difference between adaptive damping and actual system friction damping) and virtual inertia coefficients (the difference between adaptive inertia and the actual traction wheel assembly and pipe mass distribution) are defined. A mapping relationship is established between the system's multi-objective outputs (energy consumption reference value, dimensional accuracy reference value, and equivalent control speed modulus) and the virtual parameters. To solve for the adaptive parameters, the Newton-Raphson joint iterative method is used to calibrate parameters in sub-scenarios with different raw material hardness, mold wear, and cooling conditions during the offline stage, ensuring that the model can approximate the ideal dynamic response under various operating conditions.

[0024] We utilize an improved Deep Deterministic Policy Gradient (DDPG) algorithm and an Att-LSTM network with attention mechanism to perform offline learning on historical multi-attribute process data. Att-LSTM captures long-term time-series features across tracks and predicts key parameters for the next state (such as the trend of track entrance diameter and potential instability risk). DDPG explores and approximates the optimal control strategy distribution in the state-control action space, generating a transferable pre-trained parameter set and state prediction capability.

[0025] The obtained pre-trained parameter set and virtual parameters are injected into the online deep decision optimization unit to perform real-time rolling solution of the Markov process decision framework: The real-time status of the current pass (diameter, wall thickness, tension, temperature, lubrication, etc.) is input into the Att-LSTM prediction network to obtain the predicted process status value for the next time step. The target network A (Actor) receives the predicted process status value and outputs the value assessment of the control action, which is then passed to the target network C (Critic) to calculate the optimization target. The target network C iteratively updates the parameters of the target network A and the target network C based on the deviation between the actual quality feedback (such as diameter tolerance, surface defect detection results) and the predicted target, and outputs the current optimal control action (speed curve, tension compensation, lubrication frequency).

[0026] This embodiment provides a detailed description of each functional module of the control system, and the specific implementation steps are as follows: The framework building module systematically transforms the complex and continuous physical process of multi-stage drawing of thin copper tubes into a structured decision model that can be analyzed and solved by computational intelligent algorithms. Its construction process follows a progressive logic from physical entity attributes to an abstract mathematical framework.

[0027] In one embodiment disclosed in this application, the complex initial attributes affecting the drawing process are standardized to eliminate differences in dimensions and units, laying the foundation for subsequent deep coupling. Five categories of attributes—intrinsic material attributes (copper chemical composition, grain structure), initial billet state attributes (initial dimensions, hardness), die geometric attributes (cavity geometric parameters), process medium attributes (lubricant viscosity and distribution), and environmental condition attributes (ambient temperature, cooling conditions)—are deconstructed and quantified one by one. For example, for chemical composition, principal component analysis can be performed based on its contribution to material strength and plasticity to extract the weighted influence factors of key elements; for grain structure, image recognition and statistical analysis can be used to quantify it into average grain size and distribution uniformity indices; lubricant distribution can be discretized into estimated oil film thickness values ​​at key locations such as the die inlet, working area, and outlet. A weighted fusion logic is adopted: Each attribute and its sub-features are assigned a weight coefficient based on expert experience and prior experimental data. This coefficient reflects the relative influence of the attribute on the final pull-out result. Through weighted summation, the multi-dimensional sub-features are aggregated into a comprehensive evaluation index for each attribute. This process is not a simple addition, but rather provides a standardized input vector for subsequent cross-coupling.

[0028] The properties of the copper tube blank used in a certain drawing process were quantified and integrated. Three core properties were selected for analysis, including the chemical composition in the intrinsic properties of the material, the hardness in the initial state properties of the blank, and the key cavity parameters in the geometric properties of the die.

[0029] Based on expert knowledge and experimental calibration, weight coefficients reflecting the relative magnitude of their influence are assigned to these three types of attributes: In this example, the chemical composition, which determines the fundamental mechanical properties of the material, is weighted at 0.5; the billet hardness directly affects the deformation resistance, with a weight of 0.3; and the mold cavity parameters determine the deformation constraints, with a weight of 0.2. Subsequently, each attribute is deconstructed and quantified: after principal component analysis, the weighted influence factors of key elements affecting its plasticity (such as phosphorus (P) and sulfur (S)) are extracted, totaling 0.8; the initial hardness of the billet is measured to be 150 HV using a hardness tester; and through mold CAD model analysis, the ratio of the length to diameter of the core cavity sizing zone is determined to be 5.0.

[0030] A comprehensive evaluation index for each attribute type was obtained. Using a weighted fusion calculation logic, the three standardized indices were aggregated into a single billet-die feature vector that comprehensively characterizes the overall difficulty of this drawing task. The calculation formula is as follows: The comprehensive evaluation index = (weighted influence factor of chemical composition × its weight) + (initial hardness × its weight) + (cavity geometric parameters × their weight). Substituting the values ​​into the calculation, we can obtain: The comprehensive evaluation index = (0.8 × 0.5) + (150 × 0.3) + (5.0 × 0.2) = 0.4 + 45 + 1.0 = 46.4. This value of 46.4 serves as a standardized input vector component, eliminating the differences in the original data in terms of dimensions (such as HV, dimensionless ratio) and numerical range.

[0031] In one embodiment disclosed in this application, after obtaining comprehensive evaluation indicators for various attributes, these indicators are mapped to the core contradictions of the drawing process. A cross-coupling mapping method is employed. Analyze the interaction effects between different attribute categories. For example, analyze how the initial hardness of the billet and the geometric parameters of the mold cavity jointly affect the degree of local deformation, or study how lubricant viscosity and ambient temperature couple to determine the actual coefficient of friction. Using existing production data, employ nonlinear regression models in machine learning (such as Gaussian process regression) for pattern mining to extract the quantitative correlation between cross-coupling effects and the two core process output variables: deformation rate and dimensional accuracy per pass. Integrate all such correlations and, with necessary simplifications, construct a unified material-mold-equipment multi-attribute coupling model that describes the dynamic interaction between material properties, mold constraints, equipment capabilities, and process conditions. This model is essentially a high-dimensional nonlinear function; its input is all fused attribute vectors, and its output is the predicted deformation rate and dimensional accuracy.

[0032] To ensure that the constructed model not only possesses data fitting capabilities but also robust physical consistency and extrapolation reliability, first principles must be introduced as rigid boundaries for the model. Hard constraints are embedded in the logic: The physical conservation equations (such as the plastic deformation energy conservation equation) and mechanical equilibrium equations (such as the tension-velocity compatibility equation describing the relationship between pull-out force, counter-tension force, and tube wall stress) in classical metal forming theory are treated as inviolable prior knowledge and directly encoded into the model's structure or solution process. During state prediction or parameter identification in the model, any calculation results must be verified and corrected by physical constraints. For example, if the deformation rate output by the model causes the calculated plastic deformation energy to exceed the material's limit, or if the calculated traction speed and tension relationship violates mechanical equilibrium, the output is considered invalid and forcibly adjusted to the feasible region using constraint satisfaction algorithms (such as the penalty function method or the projected gradient method). Let the fused attribute vector obtained after preprocessing be: Material intrinsic properties (chemical composition weighted influence factor = 0.8), billet initial state properties (initial hardness = 150HV), mold geometric properties (cavity sizing zone length / diameter ratio = 5.0), process medium properties (lubricant viscosity = 0.05Pa·s), and environmental conditions properties (ambient temperature = 25°C).

[0033] Cross-coupling mapping focuses on the interaction effects between different types of attributes. For example, it analyzes the combined effect of the initial hardness of the billet and the geometric parameters of the mold cavity. That is, the higher the hardness, the stronger the material's resistance to plastic deformation. The larger the length / diameter ratio of the sizing band, the stronger the mold's constraint on the tube and the more concentrated the local deformation. The coupling of the two may lead to a nonlinear jump in the deformation rate. Another example is the coupling between lubricant viscosity and ambient temperature. That is, an increase in temperature will reduce the viscosity of the lubricant, thereby reducing its load-bearing and friction-reducing capabilities, increasing the actual friction coefficient, and thus affecting dimensional accuracy.

[0034] Using existing production data, a Gaussian process regression nonlinear regression model is employed for pattern mining. The processing logic is as follows: the fused attribute vector is used as input features, and the corresponding process outputs (deformation rate V and dimensional accuracy deviation ΔD) are used as labels. This trains an implicit nonlinear function relationship f([0.8,150,5.0,0.05,25])→(V,ΔD). After integrating and reasonably simplifying all such cross-coupling relationships, a multi-attribute coupling model of material-mold-equipment is constructed. This model is essentially a high-dimensional nonlinear function F(X), with the fused attribute vector X as input and the predicted deformation rate and dimensional accuracy as output.

[0035] To ensure the model has physical consistency and extrapolation reliability, hard constraint embedding logic is introduced: The plastic deformation energy conservation equation E_p=∫σ:dε≤E_max (the maximum absorbable plastic work of the material at the current temperature) and the tension-velocity coordination equation F_draw=k·v (the mechanical balance relationship between traction force and velocity) are encoded into the model as prior knowledge. Assuming the model initially predicts a deformation rate V_pred=200mm / s, the required plastic deformation energy E_p=1.5×10⁻⁶ is estimated based on the billet cross-sectional area and flow stress. 6 J, and the maximum E_max of this type of copper material at 25°C is 1.2 × 10⁻⁶. 6 J. Obviously, E_p > E_max, violating the law of conservation of energy; or the calculated traction speed v and tension F_draw do not satisfy the balance relationship F_draw = k·v, then the output is deemed invalid. The penalty function method is then used for correction. Based on the original model output, construct a modified objective function L = Loss_data + λ·Penalty, where Penalty takes a larger value when constraints are violated (e.g., λ = 10).3 This can be achieved by optimizing the prediction to force it back into the feasible region, for example, by reducing V_pred to 160 mm / s, making E_p_new ≈ 1.15 × 10⁻⁶. 6 J≤E_max, satisfying the physical constraints. Alternatively, the projection gradient method can be used to directly project the parameters back into the constraints along the normal direction of the feasible region boundary.

[0036] In one embodiment disclosed in this application, after the physical model is constructed, it is further abstracted into a Markov process decision framework suitable for sequence decision optimization, laying a mathematical foundation for subsequent intelligent control. Its processing logic is as follows: Define system state (S): All observable and decision-critical variables during the drawing process are aggregated into a multidimensional state vector. According to the description, this vector should at least include the current real-time dimensions of each pass (e.g., outer diameter, wall thickness), internal tension, tube temperature, and lubrication status (e.g., quantifiable as the reciprocal of the friction coefficient or an indicator of oil film effectiveness).

[0037] Define decision-making and control actions (A): Define the specific set of control measures that the controller can take at each decision moment. That is, the adjustment amount of servo speed, the set value of tension compensation, the adjustment command of lubrication dosage or injection frequency, etc., together constitute a control action vector.

[0038] Define the state transition function (T): This function describes how the system state evolves after a certain regulatory action is performed. Its internal logic is governed by the multi-attribute coupled model built in the first two steps and physical constraints. Specifically, given a current state and a decision-making regulatory action, the state transition function calls the coupled model to predict deformation and size changes, and uses physical equations to verify and determine the values ​​of all state variables for the next time step, ensuring that the entire transition process conforms to physical laws.

[0039] Define the reward function (R): This function is the core criterion for evaluating the merits of a state or state-control action. According to its description, it is defined as the negative of the total loss. Its calculation logic is as follows: For each decision cycle, the system calculates the three main losses caused by the current control action: energy consumption (related to servo motor power and time), dimensional deviation (related to the degree of deviation between the measured size and the target size), and equipment wear (which can be modeled as a function related to load and speed). These losses are quantified into a unified cost value, and the summation yields the total loss. The reward value is the negative of this total loss. Therefore, the goal of the system agent is transformed into finding a strategy that maximizes the long-term cumulative reward, i.e., minimizes the long-term cumulative loss.

[0040] Within one cycle, the system state (S) is defined as a multi-dimensional vector, containing the current measured outer diameter of the pipe (20.1 mm), the measured wall thickness (1.05 mm), the internal tension (5000 N), the pipe temperature (80 °C), and quantified lubrication status indicators (oil film effectiveness of 0.8). The decision action (A) is the control command selected by the controller, including increasing the servo speed by 50 rpm, setting the tension compensation to -100 N (i.e., slightly easing), and increasing the lubrication injection frequency by 10%. The internal logic of the state transition function (T) is then triggered: The current state vector and the decision action vector are input into the previously constructed multi-attribute coupled model. The model predicts that after executing this action, the outer diameter of the pipe will shrink to 20.0 mm, the wall thickness will decrease to 1.00 mm, and the temperature will rise slightly to 82℃ due to deformation heat generation. Subsequently, the physical constraint module verifies and confirms that this deformation rate and force-energy relationship are reasonable, and finally outputs the determined state vector for the next moment. Next, the reward function (R) begins to calculate the cost of this action: The system calculates the energy cost increase due to speed increase as 0.5 units, the cost of dimensional deviation due to size error (target outer diameter 20.0 ± 0.02 mm) as 1.0 unit (the deviation is within tolerance but not zero), and the cost of equipment wear due to increased load and speed as 0.2 units. Adding these three together, the total loss is 1.7 units. According to the definition of the reward function, it is the negative of the total loss; therefore, the reward value for this decision is -1.7. The MDP framework transforms the complex control problem into a quantifiable objective: The core task of an intelligent agent (controller) is to learn an optimal policy by selecting a sequence of actions that maximizes long-term cumulative rewards (i.e. minimizes long-term cumulative losses) to achieve efficient, high-quality, and device-friendly automated pull-out.

[0041] The mapping module bridges the gap between the idealized dynamic response of the Theoretical Decision-Making (MDP) framework and the complexity of real physical systems. By introducing the concept of dynamic virtual impedance into the MDP framework, unknown or time-varying dynamic characteristics (such as nonlinear friction and lumped inertia) that are difficult to model accurately but significantly affect system performance are parameterized. Then, through a rigorous offline calibration process, the decision model can adaptively approximate the ideal dynamic behavior under various complex operating conditions.

[0042] In one embodiment disclosed in this application, the differences between the idealized assumptions implied in the Markov process decision-making framework and the actual physical system are analyzed. Ideal MDP models typically assume that state transitions are instantaneous and without inertia, or only consider simplified dynamics. However, a real drawing machine, consisting of a motor, reducer, traction wheel, belt, and the tube itself, constitutes a complex electromechanical system. Its behavior is significantly constrained by the frictional damping of the actual system (originating from Coulomb friction and viscous friction at the bearings, guide rails, and contact interfaces) and the inertia (manifested as kinetic energy changes during startup, acceleration, and deceleration) caused by the actual traction wheel assembly and the mass distribution of the tube.

[0043] These two dynamic effects, not explicitly included in the core MDP model, constitute the residuals of the system's dynamic response. To systematically address this problem, a processing logic introduced by virtual impedance is employed: We envision an adaptive system whose dynamic characteristics precisely offset the aforementioned residuals, resulting in a composite system (the actual system and the adaptive system connected in series) exhibiting the desired ideal dynamic response (e.g., fast, no overshoot, low oscillation). To this end, we define two key virtual parameters: The virtual damping coefficient, whose physical meaning is defined as the difference between the adaptive damping and the actual system friction damping, is used to simulate or compensate for friction effects in the digital domain. The virtual inertia coefficient, physically defined as the difference between the adapted inertia and the actual mass distribution of the traction wheel assembly and tubing, is used to simulate or compensate for inertial effects in the digital domain. By introducing these two virtual parameters into the model, an enhanced decision-making model that more closely resembles physical reality is constructed.

[0044] The ideal MDP model assumes that the system tension can instantly reach the target value after the servo speed command is issued. However, in the actual physical system, due to the existence of actual system frictional damping (such as the Coulomb friction of the bearing of about 200N and the viscous friction coefficient of the guide rail of 0.1N·s / m) and actual inertia (the equivalent concentrated mass of the traction wheel assembly and tubing of about 50kg), significant hysteresis and overshoot will occur. To compensate for this dynamic residual, an adaptive system is conceived, whose adaptive damping and adaptive inertia are designed to precisely counteract the above effects.

[0045] Accordingly, the virtual damping coefficient is defined as the difference between the adaptive damping and the actual system friction damping, used to compensate for friction in the digital domain; the virtual inertia coefficient is defined as the difference between the adaptive inertia and the actual inertia, used to compensate for inertia. For example, to ensure that the composite system exhibits an ideal first-order inertial response (time constant of 0.5 seconds) after a velocity step, it is deduced that the adaptive damping needs to be set to 250 N·s / m and the adaptive inertia to 25 kg. Then, the virtual damping coefficient = adaptive damping - actual system friction damping ≈ 250 N·s / m - 200 N = 50 N·s / m (Note: here, Coulomb friction is simplified to equivalent viscous damping), and the virtual inertia coefficient = adaptive inertia - actual inertia = 25 kg - 50 kg = -25 kg. By introducing these two virtual parameters (positive values ​​enhance damping, negative values ​​effectively reduce inertial perception) into the enhanced MDP decision model, the speed command issued by the controller is as if it were acting on a system with more ideal dynamic characteristics that has been digitally reshaped, thus making the model prediction more realistic.

[0046] In one embodiment disclosed in this application, the system's multi-objective output is defined as a set of ideal reference values, including an energy consumption reference value (representing high efficiency), a dimensional accuracy reference value (representing high quality), and an equivalent control speed modulus (representing smoothness and speed; a smaller value generally means smoother acceleration / deceleration and less impact). To achieve comprehensive optimization of potentially conflicting objectives, a processing logic constructed using mapping relationships is employed: For each given pair of virtual parameters (virtual damping coefficient, virtual inertia coefficient), a large number of simulations are performed on the enhanced MDP model, and its comprehensive performance under typical disturbances is recorded, and the corresponding multi-objective output values ​​are calculated.

[0047] By employing multivariate function approximation techniques (such as radial basis function networks), a complex nonlinear mapping function is learned from the virtual parameter space to the multi-objective output space. This mapping function allows for backward lookup or optimization of virtual parameter configurations most likely to achieve the objectives, using the multi-objective output as a guide. For example, if the goal is to minimize energy consumption and stability metrics while maintaining dimensional accuracy, this mapping function can help evaluate the potential of different combinations of virtual parameters to achieve this objective.

[0048] On the enhanced MDP model, intensive simulations were performed for a series of different virtual parameter pairs (virtual damping coefficient, virtual inertia coefficient). For example, multiple simulations were performed under parameter pairs (10,-5), (30,-15), and (50,-25) (units N·s / m and kg, respectively). Each simulation introduced typical disturbances such as fluctuations of ±5% in raw material hardness and ±10% in mold friction coefficient, and recorded the multi-objective output values ​​under steady state, including energy consumption reference value (unit energy consumption, kWh / t), dimensional accuracy reference value (mean absolute dimensional deviation, μm), and equivalent control speed modulus value (root mean square value of speed command, characterizing stability).

[0049] Assuming the simulation results are shown in Table 1, the Radial Basis Function Network (RBFN) multivariate function approximation technique is used for learning: Table 1:

[0050] Using virtual parameter pairs as network input and multi-objective output vectors as network output, RBFN learns a complex nonlinear mapping function F from a two-dimensional virtual parameter space to a three-dimensional multi-objective output space by weighting a set of Gaussian functions centered on the training samples. The power of this function F lies in its ability to perform inverse lookups, i.e., given a set of desired multi-objective output vectors, it evaluates the potential of different virtual parameter configurations. For example, if the process objective is to minimize energy consumption and equivalent control speed modulus while ensuring dimensional accuracy reference value ≤ 8μm, the target vector is [Min(energy consumption), ≤ 8μm, Min(speed modulus)]. The target vector can be compared with the mapping relationship learned by RBFN. The query revealed that the output of parameter pair (30,-15) is [95,7.5,120], which fully meets the dimensional accuracy requirements and has moderate energy consumption and speed modulus. While parameter pair (10,-5) has the lowest energy consumption (90), its dimensional accuracy exceeds the standard (9.0). Parameter pair (50,-25) has the best dimensional accuracy (6.0), but its speed modulus is too high (160), which can easily lead to mechanical impact.

[0051] In one embodiment disclosed in this application, since the constructed mapping function is usually implicit, the optimal virtual parameters cannot be directly solved analytically. Therefore, an efficient numerical optimization algorithm is needed. The Newton-Raphson joint iterative method is used to solve for the parameters. Starting from an initially guessed virtual parameter point, the parameter values ​​are continuously corrected iteratively until suitable parameters are found that allow the multi-objective outputs to approximate their respective reference values. Specifically: Define an ideal target vector that includes all multi-target output reference values. Initialize a set of dummy parameter values. Enter the iteration loop: The mapping function is invoked to calculate the actual multi-objective output vector corresponding to the current virtual parameters, and the error vector between it and the ideal target vector is calculated. It is then determined whether the magnitude of the error vector is less than a preset convergence threshold; if so, the iteration terminates, and the current parameters are the desired fitting parameters; otherwise, the Jacobian matrix of the mapping function at the current parameter point (i.e., the local rate of change of the error vector with respect to each virtual parameter) is calculated. Using the update rule of the Newton-Raphson method, the virtual parameters are incrementally corrected along the direction of the fastest error decrease (the inverse of the negative Jacobian matrix multiplied by the error vector). The above iterative process is repeated using the corrected new parameter values. Joint iteration refers to the process of simultaneously optimizing the virtual damping coefficient and the virtual inertia coefficient to ensure that the globally optimal joint parameter solution is found.

[0052] The ideal target vector for the process is set as [energy consumption ≤ 100, dimensional accuracy = 7.0 μm, velocity modulus ≤ 125], and simplified to a target point to be approximated. The upper limits of each parameter are used to construct a reference vector T = [100, 7.0, 125]. Starting from a set of initial virtual parameter values, let's say X0 = (20, -10) (i.e., virtual damping coefficient 20 N·s / m, virtual inertia coefficient -10 kg). Enter the iterative loop: The pre-trained RBFN mapping function F is invoked to calculate the actual multi-target output vector Y0=F(X0)=[92,8.5,115] corresponding to the current parameter X0. Next, the error vector E0=T-Y0=[8,-1.5,10] is calculated. Its magnitude is much larger than the preset convergence threshold (e.g., 0.5), so iterative iteration is required.

[0053] Calculate the Jacobian matrix J, which reflects the local rate of change of the error vector E with respect to each virtual parameter. Suppose that the Jacobian matrix at point X0, estimated using the finite difference method, is J = [[∂E1 / ∂D,∂E1 / ∂I],[∂E2 / ∂D,∂E2 / ∂I],[∂E3 / ∂D,∂E3 / ∂I]] = [[0.5,0.2],[-0.1,0.8],[0.4,-0.3]] (where D represents the virtual damping coefficient and I represents the virtual inertia coefficient). Then, according to the update rule of the Newton-Raphson method, calculate the parameter increment ΔX = -J. -1 •E0. Here, we simplify the solution by using a pseudo-inverse. After calculation (omitting the tedious steps of matrix inversion and multiplication), a feasible increment ΔX≈(-10,5) is obtained. Therefore, the new virtual parameter value X1=X0+ΔX=(20-10,-10+5)=(10,-5). The system uses X1 to repeat this process: Calculate Y1=F(X1), obtaining Y1=[90,9.0,110], and the error vector E1=[10,-2.0,15]. The magnitude is still relatively large, requiring further iteration. This process is repeated, continuously updating the virtual parameters to gradually reduce the magnitude of the error vector until a set of suitable parameters is found that makes the multi-target output closest to the ideal target vector T.

[0054] In one embodiment disclosed in this application, considering that conditions such as raw material hardness, mold wear, and cooling efficiency are constantly changing in actual production, a fixed set of virtual parameters cannot cope with all scenarios. To ensure the universality of the model, scenario-specific calibration is adopted: Based on the significance of the impact of key operating condition variables on the process, the entire operating domain is divided into sub-scenarios. For example, it can be divided into three sub-scenarios based on the hardness of the raw material: soft material, medium-hard material, and hard material; into three sub-scenarios based on the wear degree of the mold: new mold, semi-worn, and near-scrap mold; and into three sub-scenarios based on the cooling conditions: strong cooling, normal cooling, and weak cooling.

[0055] The combination of sub-scenes constitutes a calibration grid covering the vast majority of actual operating conditions. Subsequently, for each specific sub-scene combination in the calibration grid (e.g., hard material + semi-wearing mold + ambient cooling), a representative dynamic residual model of the actual system is constructed. Virtual impedance is introduced, its specific mapping function is built, and the optimal set of virtual parameters for that scenario is solved using the Newton-Raphson joint iterative method. Through this systematic offline batch processing, a virtual parameter database covering the entire operating condition range is obtained. Finally, this database, along with its corresponding operating condition identification information, is sent to the control output module.

[0056] The control output module is the intelligent central hub of the entire control system. It builds upon the theoretical framework and parameter models established by the preceding modules. Through the deep integration of deep learning and reinforcement learning technologies, it extracts process knowledge from massive amounts of historical data, formulates intelligent control strategies, and ultimately achieves online, real-time, forward-looking, and adaptive control of the complex drawing process. Without interacting with physical equipment, it utilizes existing historical multi-attribute process data to train and solidify two core neural network components, providing powerful initial capabilities and prior knowledge for online applications.

[0057] In one embodiment disclosed in this application, massive amounts of historical process data collected across multiple passes are serialized and reconstructed to form process trajectory samples with each pass as a time step. The sample data for each time step includes an attribute vector comprising the inlet / outlet dimensions, tension, temperature, lubrication status, mold parameters used, raw material batch information, and the final quality inspection result for that pass. An LSTM network with an attention mechanism (Att-LSTM) is constructed and trained. The training is as follows: The network reads a batch of serialized data sequentially by channel. LSTM units handle basic short-term dependencies, with their internal cell states and hidden states responsible for transmitting and compressing information from past channels. The attention mechanism layer then builds upon this foundation, performing importance assessment and weighted aggregation of the LSTM hidden states from all past channels in parallel. For example, when analyzing the instability risk of channel N, the attention mechanism might identify a small lubrication fluctuation in channel N-2 and a slight size deviation in channel N-1 with extremely high warning weights, thus highlighting them in the final context vector.

[0058] The goal of training is to enable the Att-LSTM network to accurately predict the key parameters of the next stage (i.e., the future state) based on the sequence data of the previous stages. These parameters include not only intuitive physical quantities, such as the trend of the inlet diameter of the next stage (whether it converges or diverges), but more importantly, abstract indicators that imply process risks, such as the potential instability risk value (a normalized scalar value; a higher value indicates a greater probability of defects such as bamboo-like joints or breakage). Through iterative training, the model's prediction error is reduced to an acceptable range, resulting in an Att-LSTM pre-trained model with strong long-term temporal feature understanding and state prediction capabilities. An example of a historical process trajectory sample containing 5 stages is used for illustration: In this sequence, the attribute vector for each pass (time step) includes features such as inlet diameter, outlet diameter, tension, and temperature. Let the features of interest be the outlet diameter (Do) and the lubrication status score (Ls). The network reads data sequentially by pass. When the LSTM unit processes the 5th pass, its internal cell state has sequentially compressed all the information from the first 4 passes. At this point, the attention mechanism layer begins to work, performing parallel importance evaluation on the LSTM hidden states (h1, h2, h3, h4) of the first 4 passes. A small neural network (typically a single-layer feedforward network) is used to calculate an attention score for each hidden state, which is then normalized to weights using a softmax function. The calculated weights are [0.1, 0.1, 0.6, 0.2], indicating that the network considers the information from the third track (h3) to be crucial for prediction. Weighted aggregation is then performed, summing the weights of each hidden state to generate a context vector C, calculated as C = Σ(αi, hi), where αi is the attention weight of the i-th track. Substituting the numerical values, we get: C ≈ 0.1h1 + 0.1h2 + 0.6h3 + 0.2h4.

[0059] In this example, the small fluctuations in lubrication status (low Ls3) in lane 3 are given high weight by the attention mechanism, thus becoming prominent in C. Finally, the context vector C, along with the current input features from lane 5, is fed into the final output layer to predict the key parameters of the next lane (lane 6). Let the output have two values: The predicted trend of the next pass's entrance diameter is -0.2 (negative values ​​indicate convergence), and the predicted potential instability risk is 0.85 (higher risk). The training objective is to iteratively adjust the network parameters to minimize the error (mean squared error) between this predicted value (trend = -0.2, risk = 0.85) and the actual result of the 6th pass (e.g., actual trend = -0.18, actual risk = 0.90), ultimately obtaining an Att-LSTM pre-trained model that can deeply understand long-term temporal characteristics and has high-precision state prediction capabilities.

[0060] In one embodiment disclosed in this application, while acquiring state prediction capability, it is necessary to train a decision-maker capable of mapping from any state to the optimal control action. This task is undertaken by the improved DDPG algorithm. Its processing logic is as follows: Initialize an ActorNetwork (A) and a CriticNetwork (B), and create a copy of the target network for each to stabilize the training process. The ActorNetwork takes a state vector as input (either the current state or a future state predicted by Att-LSTM) and outputs a specific control action vector (such as a servo speed setpoint or tension compensation). The CriticNetwork takes a state and the corresponding control action generated by the Actor as input and outputs a scalar value Q(s,a) to evaluate the long-term expected reward (i.e., the expected cumulative reward) of performing the control action in that state. The improved DDPG exploration and learning logic is as follows: The algorithm randomly samples or explores the state-regulation action space of the MDP defined by the framework's building module, generating a large amount of state-regulation action-new state-reward experience data, which is stored in an experience replay buffer. Data is randomly drawn in small batches from the buffer for learning. The Critic network continuously corrects the accuracy of its Q-value evaluation through learning, making its score more accurately reflect the long-term loss (negative reward) caused by the regulation action. The Actor network updates based on the gradient signal provided by the Critic network, and its update direction tends to select regulation actions rated higher by the Critic, thereby gradually approaching the optimal regulation policy distribution.

[0061] Throughout the training process, improvements are made to the network structure (such as introducing a noisy layer to enhance exploration), reward function design, or optimizer selection to enhance learning efficiency and policy quality. After training, the parameters of the best-performing Actor and Critic networks are saved to form a transferable pre-trained parameter set. At this point, the system possesses the ability to predict and make decisions about the initial state without relying on online interaction.

[0062] In a training iteration, the system state is s (consisting of current size, tension, temperature, etc.). Actor network A outputs a regulatory action a according to its current policy, namely [velocity increment = +30 rpm, tension compensation = -50 N]. After this action is applied to the environment, the system transitions to a new state s', and due to the energy consumption, size deviation, and other losses caused by the action, it receives a negative reward r = -2.0. This experience tuple (s, a, r, s') of state-action-new state-reward is stored in the experience replay buffer. Subsequently, the algorithm randomly selects a small batch (e.g., 64) of such experience data from the buffer for learning. For one of the experiences, the system first inputs s into the target Actor network A' (a stable copy of the Actor) to obtain the target action a' for the next state, and then inputs (s', a') into the target Critic network B' (a stable copy of the Critic) to calculate the target Q value y.

[0063] The calculation of y typically follows the Bellman equation, for example, y = r + γ * Q_B'(s',a') (where γ is a discount factor representing the importance placed on future rewards; let γ = 0.9 and Q_B'(s',a') = 5.0), then y = -2.0 + 0.9 * 5.0 = 2.5. Simultaneously, the original empirical value (s,a) is input into the online Critic network B to obtain its current Q-value estimate Q_B(s,a), set to 2.0. The training objective of the Critic network is to reduce the gap between the target y and the estimated Q_B(s,a), i.e., by minimizing the loss function (e.g., mean squared error Loss = (y - Q_B(s,a))). 2The Critic network updates its own parameters to make its evaluation more accurate. After updating, it provides a gradient signal to the online Actor network A, indicating how to fine-tune its parameters to increase the Q_B(s,a) value of its output action a. The Actor's update direction is precisely towards maximizing Q_B(s,A(s)), i.e., ∇_θJ≈E[∇aQ_B(s,a)|{a=A(s)}*∇_θA(s)], where θ is the Actor network parameter. This means that if a certain action can make the Critic score high, the Actor will learn to tend to produce such actions. Through this iterative game of the Critic teaching the Actor, the Actor's strategy is gradually optimized. Throughout the process, the exploration is enhanced by introducing parameterized noise into the Actor network output layer or adding OU process noise during action selection to ensure that better strategies are discovered. After training, the parameters of the best-performing Actor and Critic networks are saved as a transferable pre-trained parameter set.

[0064] In one embodiment disclosed in this application, when the system goes online, the Att-LSTM pre-trained model and the DDPG pre-trained parameter set obtained during the offline learning phase are first loaded from non-volatile memory. Simultaneously, a virtual parameter database from the mapping module is received. The online deep decision optimization system is initialized as a dynamically switchable enhanced MDP solver, whose internal structure incorporates process intuition gained through offline learning and physical adaptability obtained through parameter calibration.

[0065] During each control cycle, real-time status data for the current pass is collected, including but not limited to real-time outer diameter, wall thickness, traction tension, pipe temperature, and the status of each lubrication monitoring point. This real-time data is then used to construct a current state vector, which is directly fed into the decision-making process and also input into the loaded Att-LSTM prediction network. Based on the current state and historical state sequences, and according to the long-term time-series patterns it has learned, the Att-LSTM network predicts the process state for the next moment (i.e., the next control cycle or the next critical process node).

[0066] During initialization, the Att-LSTM pre-trained model and DDPG pre-trained parameter set were loaded from non-volatile memory, and a virtual parameter database from the mapping module was received, forming an enhanced MDP solver with process intuition and physical adaptability. Upon entering the current control cycle, the sensors acquire a series of real-time status data: The real-time outer diameter is 20.12 mm, wall thickness is 1.03 mm, traction tension is 5100 N, pipe temperature is 81 °C, and lubrication status score is 0.78 (a higher value indicates a more effective oil film). The data is organized into a current state vector s_t and simultaneously fed into the main decision-making process and the Att-LSTM prediction network. The Att-LSTM network, based on the long-term temporal patterns learned during its training, combines the state sequences from several previous passes (e.g., t-3, t-2, t-1 times), and the LSTM units pass and compress historical information into the hidden state sequence {h_{t-3}, h_{t-2}, h_{t-1}} pass by pass. The attention mechanism layer evaluates the importance of each historical hidden state to predicting the future state in parallel, calculates the attention score of each hidden state, and normalizes it to a weight α_i using softmax. The calculated weights are [0.05, 0.15, 0.80]. Therefore, the formula for calculating the context vector C is C = Σ(α_i × h_i), i.e., C ≈ 0.05·h_{t-3} + 0.15·h_{t-2} + 0.80·h_{t-1}. This formula highlights the information contribution of the recent state (especially the t-1 pass). The context vector C and the current state vector s_t are input together into the output layer of the Att-LSTM, and the predicted process state value s_{t+1} for the next time step is obtained through forward calculation. The output is assumed to be a predicted outer diameter of 20.08 mm, a wall thickness of 1.01 mm, a tension of 5050 N, a temperature of 82 °C, and a lubrication score of 0.75, along with an additional trend value of -0.15 for the inlet diameter change of the next pass (negative values ​​indicate convergence) and a potential instability risk value of 0.88 (higher risk). This predicted value s_{t+1} will be fed into the subsequent Actor network to generate regulatory actions, thereby achieving forward-looking control based on long-term time series features.

[0067] In one embodiment disclosed in this application, after obtaining the predicted process state value, the decision-making process begins: The process state predictions generated by the Att-LSTM are input into the target network A (the online-deployed Actor network) in the online system. The Actor network evaluates the future state based on its built-in strategy and outputs a preliminary control action suggestion and its value assessment (i.e., the expected benefit it anticipates from this control action). This control action suggestion is then passed to the target network C (the online-deployed Critic network). The Critic network does not directly execute the control action; instead, it first calculates an optimized target value, the expected Q-value of the control action, based on the future state and the control action. Next, it obtains actual feedback from the quality inspection process (e.g., the tolerance between the actual diameter and the target obtained by a laser diameter gauge, or the surface defect detection results determined by a machine vision system). The Critic network compares this actual quality feedback with the previously calculated target based on the predicted state and calculates the deviation between the two. This deviation is a direct indicator of the accuracy of the current strategy.

[0068] The Critic network uses this bias signal for backpropagation to update its network parameters, thereby improving its accuracy in evaluating the value of future states. It also provides gradient information to the Actor network, driving it to adjust its policy parameters so that it can output better control actions to reduce bias when encountering similar states in the future. The complete process from state prediction to policy update is illustrated using a single actual control cycle as an example: Suppose the process state predicted by Att-LSTM is sp = {outer diameter = 20.08 mm, wall thickness = 1.01 mm, tension = 5050 N, temperature = 82 °C, lubrication score = 0.75}, with a potential instability risk value of 0.88. This prediction is input into the online deployed target network A (Actor network), which evaluates the future state according to its pre-trained strategy and outputs a preliminary control action suggestion ap = {speed increment = +40 rpm, tension compensation = -80 N}, and provides a value assessment (i.e., the expected benefit of this action) Q. a =3.2 (This value represents the expected long-term cumulative reward during training). Subsequently, ap is passed to the target network C (Critic network). The Critic calculates the optimization target value (i.e., the expected Q-value) Q_c based on the state sp and action ap. The processing logic can be expressed as Q_c = f_C(sp,ap), where f_C is the mapping function of the Critic network; the calculated result is Q_c = 3.0. At this point, the system obtains actual feedback through the quality detection process: The laser diameter gauge measures the actual outer diameter at the next moment as 20.10 mm, with a tolerance of +0.02 mm from the target value of 20.08 mm. No defects are found during surface visual inspection, and the quality score is converted to an equivalent loss cost of 0.8 (a lower value indicates better quality). The Critic network compares this actual quality feedback with the predicted target calculated based on the predicted state. First, it calculates the loss cost corresponding to the predicted target (defined by the reward function, assuming the estimated loss cost corresponding to the predicted state is 1.0). Then, it calculates the deviation e = |actual loss cost - predicted loss cost| = |0.8 - 1.0| = 0.2. This deviation measures the prediction accuracy of the current strategy. Next, the Critic network uses the deviation signal for backpropagation: The network parameters θ_C are updated to make Q_c closer to the true Q value implied by the actual quality feedback. The loss function can be expressed as Loss_C=(Q_real-Q_c). 2 Here, Q_real is the equivalent Q-value derived from the actual feedback (since the reward has negative loss, the smaller the loss, the larger Q_real). Simultaneously, the Critic transmits the gradient information ∇_aQ_c to the Actor network, driving the Actor to update its policy parameter θ_A according to ∇_θ_AJ≈E[∇aQ_c|{a=ap}×∇_θ_AA(sp)], enabling future actions output in similar states sp to further reduce bias (e.g., reducing outer diameter tolerance). Through this closed-loop iteration based on actual quality feedback, the system continuously optimizes the prediction and decision-making accuracy of the Actor and Critic, achieving dynamic adaptive optimal control.

[0069] After several rapid internal iterations, the current Actor network outputs an optimal control action under the current information conditions. This control action is specifically manifested as a set of executable control instructions, such as a smooth speed curve (specifying the servo motor's speed change over a future period), a precise tension compensation value (for real-time fine-tuning of the counter-tension), or an adjusted lubrication frequency (controlling the opening interval of the fuel injection valve). This set of instructions is ultimately sent to the PLC (Programmable Logic Controller) or directly applied to actuators such as servo drives and proportional valves, thereby completing closed-loop control for one control cycle.

[0070] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0071] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to specific implementations. Clearly, many modifications and variations can be made based on the content of this specification. The embodiments selected and specifically described in this specification are intended to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A multi-stage drawing control system for processing thin copper tubes, characterized in that: This includes a framework building module, a mapping module, and an output control module: Framework establishment module: By weighted fusion and cross-coupling mapping of multi-source parameters, a multi-attribute coupling model is established to establish the relationship between the deformation rate and dimensional accuracy of each pass. After the physical conservation and mechanical equilibrium equations are embedded as constraints into the multi-attribute coupling model, the multi-stage drawing process is abstracted into a Markov process decision framework. Mapping module: Introducing dynamic virtual impedance into the Markov process decision framework, establishing the mapping relationship between multi-objective output and virtual parameters in dynamic virtual impedance, and using the Newton-Raphson joint iterative method to calibrate parameters for different sub-scenarios in the offline stage; Control output module: Utilizing an improved deep deterministic policy gradient algorithm and long short-term memory network, it performs offline learning on historical multi-attribute process data to generate a transferable pre-trained parameter set and state prediction capability. The pre-trained parameter set and virtual parameters are injected into the online deep decision optimization unit to perform real-time rolling solution of the Markov process decision framework and output the current optimal control action.

2. The multi-stage drawing control system for processing thin copper tubes according to claim 1, characterized in that: The control output module injects the pre-trained parameter set and virtual parameters into the online deep decision optimization unit to perform real-time rolling solution of the Markov process decision framework and output the current optimal control action: The real-time status of the current process is input into the long short-term memory network to obtain the predicted value of the process status at the next moment. Target network A receives the predicted value of the process state and outputs the value assessment of the control action, which is then passed to target network C to calculate the optimization objective; The target network C iteratively updates the parameters of the target network A and the target network C based on the deviation between the actual quality feedback and the predicted target, and outputs the current optimal control action.

3. The multi-stage drawing control system for processing thin copper tubes according to claim 1, characterized in that: The mapping module introduces dynamic virtual impedance into the Markov process decision framework. The dynamic virtual impedance includes a virtual damping coefficient and a virtual inertia coefficient. The virtual damping coefficient is the difference between the adaptive damping and the actual system friction damping, and the virtual inertia coefficient is the difference between the adaptive inertia and the actual traction wheel set and pipe mass distribution. The multi-objective output includes energy consumption reference value, dimensional accuracy reference value and equivalent control speed modulus value.

4. The multi-stage drawing control system for processing thin copper tubes according to claim 2, characterized in that: The control output module utilizes an improved deep deterministic policy gradient algorithm and a long short-term memory network to perform offline learning on historical multi-attribute process data, generating a transferable pre-trained parameter set and state prediction capability. The network reads a batch of serialized data in the order of channels. The cell states and hidden states inside the long and short temporal memory network transmit and compress information from past channels. The attention mechanism layer performs importance evaluation and weighted aggregation of the LSTM hidden states of all past channels in parallel. Initialize target network A and target network B, and create a copy of the target network for each of target network A and target network B to stabilize the training process; After the target network A takes the state vector as input, it outputs a control action vector. The target network B takes the state and the corresponding control action generated by the target network A as input, and outputs a scalar value to evaluate the long-term expected benefit of performing the control action under the state vector.

5. A multi-stage drawing control system for processing thin copper tubes according to claim 4, characterized in that: The control output module injects the pre-trained parameter set and virtual parameters into the online deep decision optimization unit to perform real-time rolling solution of the Markov process decision framework and output the current optimal control action: In each control cycle, real-time status data of the current pass is collected. After the real-time data is combined into the current status vector, it is sent into the decision process and input into the long short-term memory network. The long short-term memory network predicts the process status prediction value of the next moment. The predicted process status values ​​are input into the target network A in the online system. The target network A evaluates the predicted process status values ​​and outputs preliminary control action suggestions and their value assessment. The proposed control action is transmitted to the target network C, which calculates the target value, i.e., the expected Q value of the control action. The target network C obtains actual feedback from the quality inspection process and compares the actual quality feedback with the predicted target calculated based on the predicted state to calculate the deviation between the two. The target network C uses the bias to backpropagate, update its own network parameters, provide gradient information to the target network A, and drive the target network A to adjust its policy parameters. After several internal iterations and updates, the current target network A outputs the optimal control action under the current information conditions.

6. A multi-stage drawing control system for processing thin copper tubes according to claim 3, characterized in that: The mapping module introduces dynamic virtual impedance into the Markov process decision framework, establishing a mapping relationship between multi-objective outputs and virtual parameters in the dynamic virtual impedance: For each given pair of virtual parameters, simulations are performed on the Markov process decision framework, and the overall performance under disturbances is recorded to calculate the corresponding multi-objective output values. By employing multivariate function approximation techniques, we learn nonlinear mapping functions from the virtual parameter space to the multi-objective output space.

7. A multi-stage drawing control system for processing thin copper tubes according to claim 6, characterized in that: The mapping module uses the Newton-Raphson joint iterative method to calibrate parameters for different sub-scenes in the offline phase. Set an ideal target vector containing all multi-target output reference values, initialize virtual parameter values, and enter the iteration loop: Call the mapping function to calculate the actual multi-target output vector corresponding to the current virtual parameters, and calculate the error vector between the actual multi-target output vector and the ideal target vector; Determine whether the magnitude of the error vector is less than the preset convergence threshold. If it is, the iteration terminates and the current parameter is the desired fitting parameter. If not, calculate the Jacobian matrix of the mapping function at the current parameter point. Incrementally correct the virtual parameters along the direction of the fastest error reduction, and repeat the iterative process using the corrected new parameter values ​​to find the globally optimal joint parameter solution.

8. A multi-stage drawing control system for processing thin copper tubes according to claim 1, characterized in that: The framework building module establishes a multi-attribute coupled model for the relationship between the deformation rate and dimensional accuracy of multiple sources by weighted fusion and cross-coupling mapping of multi-source parameters: Each attribute and its sub-features are assigned a weight coefficient, and the multi-dimensional sub-features are aggregated into a comprehensive evaluation index for each attribute by weighted summation. By analyzing the interaction effects between different categories of attributes, using existing production data, and employing a nonlinear regression model in machine learning for pattern mining, the quantitative correlation between cross-coupling effects and the output variables of trace deformation rate and dimensional accuracy is extracted. Integrate all relationships to complete the construction of a multi-attribute coupling model.

9. A multi-stage drawing control system for processing thin copper tubes according to claim 1, characterized in that: The multi-source parameters include the chemical composition and grain structure of the copper material, the initial size and hardness of the billet, the geometric parameters of the mold cavity, the viscosity and distribution of the lubricant, and the ambient temperature and cooling conditions.

10. A multi-stage drawing control method for processing thin copper tubes, implemented by the control system described in any one of claims 1-9, characterized in that: The control method includes the following steps: By weighted fusion and cross-coupling mapping of multi-source parameters, a multi-attribute coupling model is established to establish the relationship between the deformation rate and dimensional accuracy of each pass. By introducing physical conservation and mechanical equilibrium equations as constraints and embedding them into a multi-attribute coupled model, the multi-stage drawing process is abstracted into a Markov process decision framework. In the Markov process decision framework, a dynamic virtual impedance is introduced to establish a mapping relationship between multi-objective output and virtual parameters in the dynamic virtual impedance. The Newton-Raphson joint iterative method was used to calibrate the parameters of different sub-scenes in the offline stage. By utilizing an improved deep deterministic policy gradient algorithm and a long short-term memory network with attention mechanism, we can learn offline from historical multi-attribute process data to generate a transferable pre-trained parameter set and state prediction capability. The pre-trained parameter set and virtual parameters are injected into the online deep decision optimization unit to perform real-time rolling solution of the Markov process decision framework and output the current optimal control action.