A multi-electron gun smelting thermal hysteresis cooperative control method based on residual compensation

By adopting a multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation, the problems of large hysteresis compensation, nonlinear high-precision prediction and multi-source decoupling in the multi-electron gun melting process are solved, realizing efficient and safe intelligent control and reducing trial and error risks and computational costs.

CN122386688APending Publication Date: 2026-07-14XIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAN UNIV OF TECH
Filing Date
2026-04-21
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies struggle to simultaneously address issues such as large hysteresis compensation, nonlinear high-precision prediction, multi-source decoupling, and low-cost online adaptive processes in multi-electron gun melting, leading to frequent control oscillations, prediction distortions, and high-risk failures.

Method used

A multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation is adopted. By establishing a nonlinear MDP model that integrates thermal inertia characteristics, a physical-data dual-driven residual compensation surrogate model is constructed. The process boundary barrier function and neighborhood thermal coupling suppression mechanism are introduced to achieve online calibration and safety clamping, thereby optimizing the spatiotemporal decoupling collaborative control of multi-electron guns.

Benefits of technology

It effectively solves the control oscillation and overshoot problems caused by large thermal hysteresis of the system, improves the prediction accuracy in the nonlinear region, realizes automatic spatiotemporal decoupling of multiple heat sources and low-cost long-term stable control, and reduces trial and error risk and computational cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122386688A_ABST
    Figure CN122386688A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on residual compensation's multi-electron gun smelting thermal hysteresis collaborative control method, it is related to industrial process intelligent control technical field, comprising: S1, establish the nonlinear MDP model of fusion thermal inertia characteristic;S2, construct physical-data double-driven residual compensation type agent model, linear superposition is obtained state evolution equation with residual compensation model to physical benchmark model;S3, optimization sampling function, constraint perception based on smelting process safety boundary is carried out active sampling;S4, train multi-electron gun space-time decoupling collaborative optimization strategy;S5, construct online calibration closed loop, complete online incremental correction control, output collaborative control instruction;The application provides a kind of based on residual compensation's multi-electron gun smelting thermal hysteresis collaborative control method, solves the practical demand problems of existing technology difficult to meet large hysteresis compensation, nonlinear high-precision prediction, multi-source decoupling and low-cost online self-adapting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control technology for industrial processes, and in particular to a collaborative control method for thermal hysteresis in multi-electron gun melting based on residual compensation. Background Technology

[0002] With the surge in demand for high-end titanium alloys in aerospace, marine engineering, and other fields, the multi-electron gun collaborative control of electron beam cold hearth furnaces (EBCHM) has become crucial for improving product metallurgical quality and production efficiency. However, this melting process is a complex thermal system characterized by high dimensionality, strong coupling, large hysteresis, and strong nonlinearity. It involves physical mechanisms such as solid-liquid phase transitions and turbulent heat transfer, and multiple heat sources are highly coupled within a confined space, necessitating stringent process safety limits. Current mainstream control technologies suffer from three major limitations: First, severe thermal hysteresis leads to temperature overshoot and control oscillations in PID or conventional reinforcement learning algorithms, lacking the ability to predict thermal inertia; second, pure data models exhibit prediction distortion in the nonlinear region of phase transitions, while high-fidelity mechanism simulation is extremely time-consuming, making it difficult to obtain sufficient training samples in industrial settings; third, the strong coupling of multiple heat sources easily triggers faults such as localized overheating and burn-through, and offline optimization strategies cannot adapt to long-term operating condition drift such as furnace lining erosion and raw material fluctuations, resulting in high online retraining costs and significant trial-and-error risks. Therefore, there is an urgent need to build an integrated intelligent control scheme that takes into account thermal hysteresis compensation, high-precision model prediction, and safe decoupling of multiple heat sources in order to break through the bottleneck of special metallurgical equipment. Summary of the Invention

[0003] The purpose of this invention is to provide a multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation, which solves the problem that existing technologies cannot simultaneously meet the practical needs of large hysteresis compensation, nonlinear high-precision prediction, multi-source decoupling, and low-cost online adaptive control.

[0004] To achieve the above objectives, this invention provides a multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation, comprising the following steps: S1. Establish a nonlinear MDP model that integrates thermal inertia characteristics, define the state space and action space, and construct the reward function; S2. Construct a physical-data dual-driven residual compensation surrogate model, including a physical baseline model and an error compensation model. Linearly superimpose the physical baseline model and the residual compensation model to obtain the state evolution equation, and use the state evolution equation as the residual compensation surrogate model. S3. Define the process boundary barrier function, optimize the sampling function, and perform active sampling based on the constraint perception of the smelting process safety boundary; S4. Train a multi-electron gun spatiotemporal decoupling collaborative optimization strategy through a domain thermal coupling suppression mechanism and a safety clamping mechanism. S5. Construct an online calibration closed loop, complete online incremental correction control, and output cooperative control commands.

[0005] Preferably, in S1: in the state space, construct The state vector of the melting system at any time The uniformity of the thermal field is characterized by the energy distribution gradient in multiple regions, the trend of thermal inertia is captured by the sliding window feature of historical time series, and the temperature trend in the future is predicted by the first derivative of the rate of change of the molten pool temperature. In the action space, construct The action vector output by the time-control policy network The relative adjustment increments of the power of each electron gun and the scanning angle are taken as actions, and the maximum rate of change constraint of the physical actuator is set. reward function It includes a trend prediction potential energy term, which includes a potential energy reward mechanism based on the rate of temperature change: if the current temperature does not meet the target but the rate of temperature change points to the target value, a positive incentive is given.

[0006] Preferably, in S2: S21, constructing a physical baseline model. Based on Stefan-Boltzmann thermal radiation law and energy conservation equation, a lumped parameter energy balance equation is established to generate a physical benchmark model for rapid calculation of the basic temperature field trend under ideal working conditions. S22. Constructing an error compensation model The Gaussian process regression algorithm is used to fit the highly nonlinear residuals caused by in-furnace melt turbulence, latent heat of phase change and measurement noise to obtain the error compensation model. S23. By linearly superimposing the physical baseline model and the error compensation model, a state evolution equation is generated, enabling high-precision, physical-distortion-free state prediction under small sample conditions.

[0007] Preferably, in S21, the expression for the lumped parameter energy balance equation is: ; In the formula, Indicates in The temperature of the molten pool at the next moment, as predicted by the physical baseline model; Indicates in The current actual physical temperature of the molten pool at any given moment; Indicates the discrete time step of the control system; This indicates the specific heat capacity at constant pressure of the melt within the molten pool; This indicates the density of the melt within the molten pool; Indicates the volume of the molten pool; This represents the effective heat dissipation surface area of ​​the molten pool. This indicates the effective absorption rate of the electron gun's energy; Indicates the input power of the electron gun; Represents the Stefan constant; Indicates the emissivity of the alloy; express; This indicates the absolute temperature of the internal environment of the vacuum furnace; This represents the equivalent combined heat transfer coefficient between the surface of the molten pool and its boundary; Indicates the cooling boundary temperature; The expression for the physical baseline model is: .

[0008] Preferably, in S3: S31, based on the physical limits of electron gun power density and scanning dwell time, a process boundary barrier function is constructed. If the sampling point is close to the fault boundary that causes the raw material to "cold shut" or the liquid surface to "overheat", the value of the process boundary barrier function increases exponentially, and a soft constraint is applied to the sampling action. S32. Construct a multi-objective acquisition function that includes the process boundary barrier function, and perform active sampling through a constraint-aware sampling strategy based on the process safety boundary.

[0009] Preferably, the expression for the process boundary barrier function in S31 is: ; In the formula, This represents the state vector of the current sampling point; Indicates the fault boundary index; Indicates the number of fault boundaries; Let represent the penalty decay coefficient of the j-th class boundary; Indicates the current sampling point Distance to Euclidean distance of the fault boundary; Indicates the first Fault boundary.

[0010] Preferably, the expression for the multi-target acquisition function in S32 is: ; In the formula, This represents the expected improvement in the predicted mean of the residual compensation surrogate model based on a physical-data dual-driven approach; Indicates the adaptive weighting coefficient; This represents the error compensation model.

[0011] Preferably, in S4: S41, Neighborhood thermal coupling suppression mechanism: a mutual exclusion loss function is introduced into the loss function of the policy training, defining the scanning overlap threshold and local energy density limit of adjacent electron guns. If the distance and total power of the two electron guns both exceed the set threshold, causing the local energy density to exceed the safety threshold, the neighborhood thermal coupling suppression mechanism takes effect, forcing the spatial separation of the scanning area or reducing the overlap rate of adjacent electron guns, thereby achieving automatic spatiotemporal decoupling. S42, Safety Clamping Mechanism: The future temperature rise rate is calculated in real time using a physical benchmark model to obtain the predicted value of thermal potential energy. If the predicted value of thermal potential energy exceeds the thermal shock stability threshold of the furnace material, the action clamping is triggered to forcibly cut off the current high-power output command. S43. By using deep reinforcement learning algorithms, a neighborhood thermal coupling suppression mechanism and a safety clamping mechanism are introduced into the residual compensation surrogate model for policy training, thereby obtaining a multi-electron gun spatiotemporal decoupling collaborative optimization strategy that meets thermal safety specifications.

[0012] Preferably, the expression for the neighborhood thermal coupling suppression mechanism in S41 is: ; ; In the formula, This represents the neighborhood thermal coupling suppression penalty term, used to quantify the degree of danger of the spatial thermal field superposition between adjacent electron guns exceeding the limit; Indicates the electron gun index; Indicates the total number of electron guns; Indicates the first The instantaneous power of the electron gun; Indicates the first The instantaneous power of the electron gun; Indicates the first Branch and the first The geometric distance between the center point of the electron gun scan; This indicates the safe threshold for local energy density in the molten pool; This represents the overall total loss function value during the training of the policy network; This represents the original loss function value of the basic SAC reinforcement learning algorithm; This represents the thermal coupling penalty weighting coefficient.

[0013] Preferably, in S5: S51, thermal field reconstruction sensing: mapping limited surface temperature measurement data into a complete state vector containing energy gradient; S52, Residual Calibration Learning: Through the sliding window mechanism, the deviation between the actual running data and the predicted value of the physical benchmark model is collected online, and the root mean square error or mean absolute error of the deviation within the sliding window is calculated in real time. When the statistical error exceeds the preset prediction accuracy decay threshold for N consecutive control cycles, it is determined that the model performance has decayed, and the covariance matrix of the error compensation model is locally incrementally updated. S53. Load the trained multi-electron gun spatiotemporal decoupling cooperative optimization strategy, and output the cooperative control commands for each electron gun according to the calibrated state vector.

[0014] Therefore, the present invention employs the above-mentioned multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation, and the specific beneficial effects are as follows: (1) By introducing a trend prediction potential energy term into the reward function, the control strategy can sense the temperature change inertia in advance, which effectively solves the control oscillation and overshoot problem caused by the large thermal lag of the system. (2) By using a physical-data dual-driven residual compensation surrogate model, the prediction accuracy of the model in the nonlinear region of solid-liquid phase transition is significantly improved, overcoming the prediction distortion defect of the pure data-driven model in key physical processes, and greatly reducing the dependence on the amount of training samples. (3) Through the neighborhood thermal coupling suppression mechanism and the constraint perception sampling strategy based on the process safety boundary, the spatiotemporal automatic decoupling of multi-heat source collaborative operation and the automatic avoidance of high-risk process areas are realized, which effectively prevents local overheating or cold insulation failure caused by thermal field superposition and significantly reduces the trial and error risk in industrial field. (4) The online learning mechanism based on residual incremental calibration enables the system to quickly adapt to working condition drifts such as furnace lining erosion or raw material fluctuations without retraining the entire model, thereby achieving long-term stable, safe and efficient intelligent closed-loop control with extremely low computational cost.

[0015] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0016] Figure 1 This is an overall flowchart of a multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation according to the present invention. Figure 2 This is a schematic diagram of the dual residual compensation surrogate model structure according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the constraint-aware sampling strategy according to an embodiment of the present invention; Figure 4 This is a training and execution logic diagram of the multi-electron gun spatiotemporal decoupling and collaborative optimization strategy according to an embodiment of the present invention. Figure 5This is a schematic diagram of the system temperature response under dual-drive residual compensation control according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the dynamic curve of the controller action in an embodiment of the present invention. Detailed Implementation

[0017] The following detailed description of embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0018] Please see Figures 1-4 A collaborative control method for thermal hysteresis in multi-electron gun melting based on residual compensation includes the following steps: S1. Establish a nonlinear Markov decision process (MDP) model that incorporates thermal inertia characteristics, define the state space and action space, and construct the reward function; In the state space, construct The state vector of the melting system at any time , Including the current actual physical temperature of the molten pool At the same time, an energy distribution gradient is introduced. derivative with the rate of change of temperature Energy distribution gradient The temperature field on the surface of the molten pool is acquired using an infrared thermal imager, and the temperature gradient at the boundary between the roughing and refining zones is calculated to characterize the uniformity of the thermal field; the derivative of the rate of temperature change is also used. Using the sliding window data from the past 5 seconds, the first derivative of the temperature at key measuring points in the molten pool is calculated as a basis for predicting future thermal potential energy. In the action space, construct The action vector output by the time-control policy network , Including the input power of the electron gun The relative adjustment increments of the power of each electron gun and the scanning angle are taken as actions, and the maximum rate of change constraint of the physical actuator is set. reward function This includes a trend prediction potential energy term, which includes a potential energy reward mechanism based on the rate of temperature change: if the current temperature does not meet the target but the rate of temperature change points to the target value, a positive incentive is given to suppress overshoot; if the rate of temperature change exceeds the set threshold, resulting in the risk of overshoot, a negative incentive is given; the potential energy reward mechanism based on the rate of temperature change is used to guide the control strategy in advance in long-lag feedback and eliminate oscillations. The expression for the reward function is: ; In the formula, express The state vector of the melting system at any given moment; express The action vector output by the time-control policy network; This represents the basic target reward, including penalties for energy consumption, uniformity, and quality. The trend prediction potential term is expressed as: ; In the formula, Indicates the target melting temperature; This indicates the current temperature of the molten pool; Indicates the potential energy reward weight; Indicates the sensitivity coefficient; This represents the first derivative of the rate of temperature change fitted using sliding window data. If the temperature deviation is consistent with the sign of the temperature change trend, i.e., it is moving towards the target in a positive direction, it is still a positive value even if the current temperature does not meet the target. This ensures the "correct trend" of the potential energy reward mechanism control strategy and overcomes thermal inertia lag.

[0019] S2. Construct a physical-data dual-driven residual compensation proxy model, including a physical baseline model. And error compensation model The state evolution equation is obtained by linearly superimposing the physical baseline model and the residual compensation model, and the state evolution equation is used as a residual compensation surrogate model. S21. Construct a physical baseline model: Based on the Stefan-Boltzmann law of thermal radiation and the energy conservation equation, establish a simplified lumped parameter energy balance equation to generate a physical baseline model for quickly calculating the trend of the basic temperature field under ideal working conditions. The expression for the lumped parameter energy balance equation is: ; In the formula, Indicates in The temperature of the molten pool at the next moment, as predicted by the physical baseline model; Indicates in The current actual physical temperature of the molten pool at any given moment; Indicates the discrete time step of the control system; This indicates the specific heat capacity at constant pressure of the melt within the molten pool; This indicates the density of the melt within the molten pool; Indicates the volume of the molten pool; This represents the effective heat dissipation surface area of ​​the molten pool. This indicates the effective absorption rate of the electron gun's energy; Indicates the input power of the electron gun; Represents the Stefan constant; Indicates the emissivity of the alloy; This indicates the absolute temperature of the internal environment of the vacuum furnace; This represents the equivalent combined heat transfer coefficient between the surface of the molten pool and its boundary; Indicates the cooling boundary temperature; The expression for the physical baseline model is: ; S22. Constructing an error compensation model: Using the Gaussian process regression (GPR) algorithm, the prediction residuals of the physical benchmark model are used as training targets. The highly nonlinear residuals caused by furnace melt turbulence, latent heat of phase change, and measurement noise are fitted by the Matérn 5 / 2 kernel function to obtain the error compensation model. S23. By linearly superimposing the physical baseline model and the error compensation model, a state evolution equation is generated, achieving high-precision, physically distortion-free state prediction under small sample conditions; the expression of the state evolution equation is: ; In the formula, express The system state vector at the next moment is predicted by the dual-drive model; express The current state vector of the melting system at any given moment; express The action vector output by the time-control strategy network.

[0020] S3. Define the process boundary barrier function, optimize the sampling function, and implement constraint-aware active sampling based on the safety boundary of the smelting process. S31. Based on the physical limits of electron gun power density and scanning dwell time, a process boundary barrier function is constructed. If the sampling point approaches the fault boundary that causes raw material "cold shut-off" or liquid surface "overheating," the value of the process boundary barrier function increases exponentially, applying a soft constraint to the sampling action and guiding the sampling point to gather in the high uncertainty region of the solid-liquid phase transition boundary; the expression is: ; In the formula, This represents the state vector of the current sampling point; Indicates the fault boundary index; Indicates the number of fault boundaries; Let represent the penalty decay coefficient of the j-th class boundary; Indicates the current sampling point Distance to Euclidean distance of the fault boundary; Indicates the first Class fault boundary; S32. Construct a multi-objective acquisition function that includes the process boundary barrier function, and perform active sampling through a process safety boundary constraint-based sensing sampling strategy. The expression for the multi-target acquisition function is: ; In the formula, This represents the expected improvement in the predicted mean of the residual compensation surrogate model based on a physical-data dual-driven approach; Indicates the adaptive weighting coefficient; This represents the error compensation model.

[0021] S4. Train a multi-electron gun spatiotemporal decoupling collaborative optimization strategy through a domain thermal coupling suppression mechanism and a safety clamping mechanism. S41. Neighborhood Thermal Coupling Suppression Mechanism: A mutual exclusion loss function is introduced into the loss function of policy training. A threshold for the scanning overlap of adjacent electron guns and a local energy density limit are defined. If the distance and total power of two electron guns both exceed the set thresholds, causing the local energy density to exceed the safety threshold, the neighborhood thermal coupling suppression mechanism takes effect, forcibly separating the scanning region spatially or reducing the overlap rate of adjacent electron guns, thus achieving automatic spatiotemporal decoupling. The expression is: ; ; In the formula, This represents the neighborhood thermal coupling suppression penalty term, used to quantify the degree of danger of the spatial thermal field superposition between adjacent electron guns exceeding the limit; Indicates the electron gun index; Indicates the total number of electron guns; Indicates the first The instantaneous power of the electron gun; Indicates the first The instantaneous power of the electron gun; Indicates the first Branch and the first The geometric distance between the center point of the electron gun scan; This indicates the safe threshold for local energy density in the molten pool; This represents the overall total loss function value during the training of the policy network; This represents the original loss function value of the basic SAC reinforcement learning algorithm; This represents the thermal coupling penalty weighting coefficient.

[0022] S42, Safety Clamping Mechanism: The future temperature rise rate is calculated in real time using a physical benchmark model to obtain the predicted value of thermal potential energy. If the predicted value of thermal potential energy exceeds the thermal shock stability threshold of the furnace material, the action clamping is triggered to forcibly cut off the current high-power output command. S43. By using deep reinforcement learning algorithms, a neighborhood thermal coupling suppression mechanism and a safety clamping mechanism are introduced into the residual compensation surrogate model for policy training, thereby obtaining a multi-electron gun spatiotemporal decoupling collaborative optimization strategy that meets thermal safety specifications.

[0023] S5. To address the operational drift caused by furnace lining erosion or raw material fluctuations, an online calibration closed loop is constructed to complete online incremental correction control and output collaborative control commands. S51, Thermal Field Reconstruction Sensing: Mapping limited surface temperature measurement data into a complete state vector containing energy gradients; S52, Residual Calibration Learning: Using a sliding window mechanism, the deviation between the actual operating data and the predicted values ​​of the physical baseline model over the past 30 minutes is collected online. The root mean square error or mean absolute error of this deviation within the sliding window is calculated in real time. When the statistical error exceeds the preset prediction accuracy decay threshold for N consecutive control cycles, it is determined that the model performance has degraded. The parameters of the physical baseline model remain unchanged, and the covariance matrix of the error compensation model is locally incrementally updated. The expression is: ; In the formula, This represents the hyperparameters of the covariance matrix of the error compensation model after incremental updates; This represents the hyperparameters of the covariance matrix of the error compensation model before the update; This indicates the learning rate or step size for incremental updates. Indicates the parameter The gradient operator; Represents the residual loss function of the error compensation model; S53. Load the trained multi-electron gun spatiotemporal decoupling cooperative optimization strategy, and output the cooperative control commands for each electron gun according to the calibrated state vector.

[0024] Experimental verification For simulation experiments of the method of the present invention, please refer to [link / reference]. Figure 5 Set the target melting temperature ( Figure 5 The black dashed line in the figure represents 1600℃; during the initial heating phase, the actual melting temperature (…) Figure 5 The blue solid line in the diagram rises rapidly. When the actual melting temperature approaches the target melting temperature, thanks to the introduction of a trend prediction potential energy term and an overshoot penalty mechanism in the reward function, the controller senses the system's huge thermal inertia in advance and performs a smooth "advance braking" adjustment. This allows the temperature curve to quickly fix at the 1600℃ baseline after a very small dynamic convergence, overcoming the common problems of catastrophic overshoot and long-period oscillation in traditional control and achieving a smooth transition.

[0025] Please see Figure 6This demonstrates the details of the coordinated control of multiple electron guns. In the first 15 steps of the control cycle (the initial rapid heating phase), the output power of both electron guns (red and pink lines) reaches its peak (full load operation), and the local total energy input far exceeds the process safety limit. At this point, the neighborhood thermal coupling suppression mechanism of this invention is instantly triggered, and the system automatically instructs the spatial scanning distance between the two electron guns (solid green line) to be significantly increased to a safe distance of nearly 50cm, perfectly realizing the spatiotemporal decoupling of multiple heat sources and fundamentally avoiding the risk of overheating and burn-through caused by the superposition of thermal fields at the interface.

[0026] Furthermore, after the temperature enters the steady-state range (after step 30), the output power of the two electron guns does not remain static, but exhibits alternating, high-frequency, non-periodic dynamic fine-tuning. This directly confirms that the physical-data dual-driven residual compensation proxy model of this invention is operating in real time in the background, continuously resisting nonlinear physical disturbances caused by latent heat absorption of solid-liquid phase transitions and external random noise through millisecond-level micro-operation commands, ensuring the system's extremely high steady-state control accuracy and disturbance resistance robustness under complex operating conditions.

[0027] Therefore, this invention employs a multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation. By introducing a trend prediction potential energy term into the reward function, the control strategy can sense the inertia of temperature changes in advance, effectively solving the control oscillation and overshoot problems caused by large thermal hysteresis of the system. Through a physical-data dual-driven residual compensation surrogate model, the prediction accuracy of the model in the nonlinear region of solid-liquid phase transition is significantly improved, overcoming the prediction distortion defect of pure data-driven models in key physical processes, and greatly reducing the dependence on the amount of training samples. Through the neighborhood thermal coupling suppression mechanism and the constraint-aware sampling strategy based on process safety boundaries, the spatiotemporal automatic decoupling of multi-heat source collaborative operation and the automatic avoidance of high-risk process areas are realized, effectively preventing local overheating or cold shut-off faults caused by thermal field superposition, and significantly reducing the trial-and-error risk in industrial settings. The online learning mechanism based on residual incremental calibration enables the system to quickly adapt to operating condition drifts such as furnace lining erosion or raw material fluctuations without retraining the entire model, thereby achieving long-term stable, safe, and efficient intelligent closed-loop control with extremely low computational cost.

[0028] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation, characterized in that, Includes the following steps: S1. Establish a nonlinear MDP model that integrates thermal inertia characteristics, define the state space and action space, and construct the reward function; S2. Construct a physical-data dual-driven residual compensation surrogate model, including a physical baseline model and an error compensation model. Linearly superimpose the physical baseline model and the residual compensation model to obtain the state evolution equation, and use the state evolution equation as the residual compensation surrogate model. S3. Define the process boundary barrier function, optimize the sampling function, and perform active sampling based on the constraint perception of the smelting process safety boundary; S4. Train a multi-electron gun spatiotemporal decoupling collaborative optimization strategy through a domain thermal coupling suppression mechanism and a safety clamping mechanism. S5. Construct an online calibration closed loop, complete online incremental correction control, and output cooperative control commands.

2. The multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation according to claim 1, characterized in that, In S1: In the state space, construct The state vector of the melting system at any time The uniformity of the thermal field is characterized by the energy distribution gradient in multiple regions, the trend of thermal inertia is captured by the sliding window feature of historical time series, and the temperature trend in the future is predicted by the first derivative of the rate of change of the molten pool temperature. In the action space, construct The action vector output by the time-control policy network The relative adjustment increments of the power of each electron gun and the scanning angle are taken as actions, and the maximum rate of change constraint of the physical actuator is set. reward function It includes a trend prediction potential energy term, which includes a potential energy reward mechanism based on the rate of temperature change: if the current temperature does not meet the target but the rate of temperature change points to the target value, a positive incentive is given.

3. The multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation according to claim 2, characterized in that, In S2: S21. Constructing a physical baseline model Based on Stefan-Boltzmann thermal radiation law and energy conservation equation, a lumped parameter energy balance equation is established to generate a physical benchmark model for rapid calculation of the basic temperature field trend under ideal working conditions. S22. Constructing an error compensation model The Gaussian process regression algorithm is used to fit the highly nonlinear residuals caused by in-furnace melt turbulence, latent heat of phase change and measurement noise to obtain the error compensation model. S23. By linearly superimposing the physical baseline model and the error compensation model, a state evolution equation is generated, enabling high-precision, physical-distortion-free state prediction under small sample conditions.

4. The multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation according to claim 3, characterized in that, In S21: The expression for the lumped parameter energy balance equation is: ; In the formula, Indicates in The temperature of the molten pool at the next moment, as predicted by the physical baseline model; Indicates in The current actual physical temperature of the molten pool at any given moment; Indicates the discrete time step of the control system; This indicates the specific heat capacity at constant pressure of the melt within the molten pool; This indicates the density of the melt within the molten pool; Indicates the volume of the molten pool; This represents the effective heat dissipation surface area of ​​the molten pool. This indicates the effective absorption rate of the electron gun's energy; Indicates the input power of the electron gun; Represents the Stefan constant; Indicates the emissivity of the alloy; This indicates the absolute temperature of the internal environment of the vacuum furnace; This represents the equivalent combined heat transfer coefficient between the surface of the molten pool and its boundary; Indicates the cooling boundary temperature; The expression for the physical baseline model is: 。 5. The multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation according to claim 4, characterized in that, In S3: S31. Based on the physical limits of electron gun power density and scanning dwell time, a process boundary barrier function is constructed. If the sampling point is close to the fault boundary that causes raw material "cold shut-off" or liquid surface "overheating", the value of the process boundary barrier function increases exponentially, thus applying soft constraints to the sampling action. S32. Construct a multi-objective acquisition function that includes the process boundary barrier function, and perform active sampling through a constraint-aware sampling strategy based on the process safety boundary.

6. The multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation according to claim 5, characterized in that, The expression for the process boundary barrier function in S31 is: ; In the formula, This represents the state vector of the current sampling point; Indicates the fault boundary index; Indicates the number of fault boundaries; Let represent the penalty decay coefficient of the j-th class boundary; Indicates the current sampling point Distance to Euclidean distance of the fault boundary; Indicates the first Fault boundary.

7. The multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation according to claim 6, characterized in that, The expression for the multi-target acquisition function in S32 is: ; In the formula, This represents the expected improvement in the predicted mean of the residual compensation surrogate model based on a physical-data dual-driven approach; Indicates the adaptive weighting coefficient; This represents the error compensation model.

8. The multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation according to claim 7, characterized in that, In S4: S41. Neighborhood thermal coupling suppression mechanism: A mutual exclusion loss function is introduced into the loss function of policy training. The scanning overlap threshold and local energy density limit of adjacent electron guns are defined. If the distance and total power of the two electron guns exceed the set threshold, resulting in the local energy density exceeding the safety threshold, the neighborhood thermal coupling suppression mechanism takes effect, forcing the spatial separation of the scanning area or reducing the overlap rate of adjacent electron guns to achieve automatic spatiotemporal decoupling. S42, Safety Clamping Mechanism: The future temperature rise rate is calculated in real time using a physical benchmark model to obtain the predicted value of thermal potential energy. If the predicted value of thermal potential energy exceeds the thermal shock stability threshold of the furnace material, the action clamping is triggered to forcibly cut off the current high-power output command. S43. By using deep reinforcement learning algorithms, a neighborhood thermal coupling suppression mechanism and a safety clamping mechanism are introduced into the residual compensation surrogate model for policy training, thereby obtaining a multi-electron gun spatiotemporal decoupling collaborative optimization strategy that meets thermal safety specifications.

9. The multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation according to claim 8, characterized in that, The expression for the neighborhood thermal coupling suppression mechanism in S41 is: ; ; In the formula, This represents the neighborhood thermal coupling suppression penalty term, used to quantify the degree of danger of the spatial thermal field superposition between adjacent electron guns exceeding the limit; Indicates the electron gun index; Indicates the total number of electron guns; Indicates the first The instantaneous power of the electron gun; Indicates the first The instantaneous power of the electron gun; Indicates the first Branch and the first The geometric distance between the center point of the electron gun scan; This indicates the safe threshold for local energy density in the molten pool; This represents the overall total loss function value during the training of the policy network; This represents the original loss function value of the basic SAC reinforcement learning algorithm; This represents the thermal coupling penalty weighting coefficient.

10. The multi-electron gun melting thermal hysteresis collaborative control method based on residual compensation according to claim 9, characterized in that, In S5: S51, Thermal Field Reconstruction Sensing: Mapping limited surface temperature measurement data into a complete state vector containing energy gradients; S52, Residual Calibration Learning: Through the sliding window mechanism, the deviation between the actual running data and the predicted value of the physical benchmark model is collected online, and the root mean square error or mean absolute error of the deviation within the sliding window is calculated in real time. When the statistical error exceeds the preset prediction accuracy decay threshold for N consecutive control cycles, it is determined that the model performance has decayed, and the covariance matrix of the error compensation model is locally incrementally updated. S53. Load the trained multi-electron gun spatiotemporal decoupling cooperative optimization strategy, and output the cooperative control commands for each electron gun according to the calibrated state vector.