Reinforced learning-sliding mode PID (Proportion Integration Differentiation) cooperative control method and system for temperature of laser cladding molten pool
Through the reinforcement learning-sliding mode PID collaborative control method, combining sliding mode control and reinforcement learning PID control, the response speed and stability problems of traditional laser cladding control methods in complex environments are solved, high-precision and robust molten pool temperature control is achieved, and the process stability and equipment life of laser cladding are improved.
Patent Information
- Application Number
- CN202511137797.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional laser cladding control methods have difficulty achieving rapid response, high stability, and high-precision temperature control when faced with complex and changing environments or system characteristics, resulting in defects such as molten pool collapse, low dimensional accuracy, and coarse grains, which limits the development of laser cladding.
The reinforcement learning-sliding mode PID collaborative control method is adopted, combining sliding mode control and reinforcement learning PID control. Through melt pool temperature monitoring, disturbance observation, sliding mode control and hybrid strategy modules, the PID parameters are optimized in real time to achieve high-precision and robust control of the melt pool temperature.
It improves the steady-state accuracy and anti-interference ability of the laser cladding process, reduces the cost of manual parameter debugging, extends the life of the equipment, and improves the process stability and intelligence level.
Smart Images

Figure CN120802599A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of selective laser melting process control, in particular to a reinforcement learning-sliding mode PID collaborative control method and system for laser melting pool temperature. BACKGROUND
[0002] With the development of computer-aided design and manufacturing technology, additive manufacturing (AM) has received great attention, which can cut complex three-dimensional (3D) structures into two-dimensional layers and further melt and stack materials layer by layer to directly drive the production of parts. Laser cladding as a metal additive manufacturing method has higher manufacturing efficiency compared to other methods and can provide a larger process parameter window. In addition, the parts manufactured in laser cladding have excellent structural density and metallurgical bonding. Laser cladding is also widely used in surface repair and coating preparation.
[0003] However, due to the increase of heat input, laser cladding is more prone to defects such as molten pool collapse, low dimensional accuracy and coarse grains, which greatly affects the strength of metal structural workpieces and further limits the development of laser cladding. Implementing process control is an efficient method to produce repeatable, reliable and quality-assured parts in laser cladding.
[0004] Traditional single control methods often cannot achieve fast response, high stability and high precision requirements when facing complex changes in additive manufacturing or changes in system characteristics. By using an infrared thermal imaging camera to monitor the molten pool temperature of the laser cladding process online and in real time, and controlling the molten pool temperature in real time, a closed-loop control system is developed, a control algorithm is designed, and the requirements of fast response, high stability and high precision are achieved, which improves the quality of the final product.
[0005] Compared with the prior art, the differences are as follows:
[0006] Comparison of control algorithm and adaptability with patent CN112517926B "Method for regulating molten pool temperature gradient in laser cladding process"
[0007] Patent CN112517926B uses a traditional PID algorithm to regulate the temperature. The PID parameters (such as KP, KI, KD) in this method are initial set values, and the essence is to use this fixed PID parameter for control after printing several passes at a constant power, which lacks the ability to adjust online and adapt to the nonlinear and time-varying characteristics in the cladding process.
[0008] The patent constructs a reinforcement learning-sliding mode PID collaborative control architecture. One of its cores is the reinforcement learning PID module, which can autonomously learn through Q-learning and other mechanisms, dynamically optimize PID control parameters (Kp, Ki, Kd) online according to real-time temperature errors and error rates. This design enables the controller to autonomously learn and adapt to complex changes in the process, eliminating the need for repeated manual debugging, and improving intelligence and process adaptability.
[0009] The patent CN112517926B uses a single PID control strategy, and its robustness depends entirely on the PID algorithm itself. When facing strong disturbances such as laser power fluctuations and material thermal physical parameter changes, a single PID controller may have difficulty balancing response speed and steady-state accuracy.
[0010] The patent introduces a multi-mode hybrid control strategy, organically integrating sliding mode control (SMC) and reinforcement learning PID control. The sliding mode control module is used to handle large temperature errors, and its strong robustness and fast response characteristics are used for coarse adjustment; the reinforcement learning PID module is used for fine adjustment when the error is small, ensuring steady-state accuracy. In addition, the patent also designs a disturbance observation module that can estimate external thermal disturbances in real time and perform feedforward compensation, further enhancing the system's active anti-interference ability.
[0011] Comparison with the technology of patent CN111753400B "Laser cladding forming pool temperature control method"
[0012] There are essential differences in control mechanisms and modeling.
[0013] The patent CN111753400B uses a predictive control idea based on a physical model. It collects laser scanning paths and builds a geometric and thermodynamic mathematical model composed of virtual cubes inside the program to simulate and predict the temperature field distribution of the workpiece. The control decision is based on the predicted value of the next processing point temperature, aiming to improve control accuracy through feedforward prediction.
[0014] The patent (reinforcement learning-sliding mode PID collaborative control method and system for laser cladding pool temperature) uses a data-driven adaptive control idea. It does not establish a physical model of the workpiece, but estimates the system's thermal disturbance in real time through the disturbance observation module, and uses the autonomous learning mechanism of the reinforcement learning PID module to optimize the control parameters online according to the actual temperature error. There are essential differences between the two in system modeling and prediction mechanisms. The former relies on a pre-built physical process simulation model for predictive control, while the patent relies on real-time data and online learning models for adaptive control.
[0015] There are essential differences in control strategies and adaptability.
[0016] The core control algorithm of patent CN111753400B is a PID control model. Although its PID parameters can be dynamically adjusted according to the incoming predicted temperature value, its control framework is relatively single, mainly dealing with temperature changes caused by predictable factors such as heat accumulation.
[0017] The present patent adopts a hybrid control strategy of multiple modes in cooperation. It combines two control methods of sliding mode control (SMC) and reinforcement learning PID: when the temperature error is large, the sliding mode control is mainly used to quickly correct the error by using its strong robustness and high response speed characteristics; when the error is small, the reinforcement learning PID control is mainly used to realize stable tracking and accurate adjustment of the signal. This hybrid strategy adaptively fuses the two control components by analyzing the error amplitude, solving the contradiction between response speed, steady-state accuracy and anti-interference of traditional single control method.
[0018] There are essential differences between the two in terms of control strategy and execution mode. The former is a single controller optimization based on prediction, while the present patent is an advanced control strategy based on real-time state and adaptive fusion of multiple controllers, which has stronger adaptability to complex and nonlinear disturbances. SUMMARY
[0019] To solve the above problems, the present application discloses a reinforcement learning-sliding mode PID collaborative control method and system for laser cladding molten pool temperature.
[0020] To achieve the above purpose, the technical solution adopted by the present application is:
[0021] The reinforcement learning-sliding mode PID collaborative control system for laser cladding molten pool temperature comprises a molten pool temperature monitoring module, a disturbance observation module, a sliding mode control module, a reinforcement learning PID control module and a hybrid strategy module.
[0022] The molten pool temperature monitoring module is used to collect the actual temperature signal of the molten pool in the laser cladding process in real time, and form a temperature error with the preset temperature value.
[0023] The disturbance observation module is used to estimate the system thermal disturbance in real time based on the laser power control signal and the molten pool temperature variation characteristics, and generate a disturbance compensation signal.
[0024] The sliding mode control module receives the temperature error and the disturbance compensation signal, and generates a first control component with strong robustness through an anti-disturbance control strategy.
[0025] The reinforcement learning PID module dynamically optimizes the proportional-integral-derivative parameters through an autonomous learning mechanism, and generates a second control component according to the temperature error.
[0026] The hybrid strategy module analyzes the amplitude characteristics of the temperature error, adaptively fuses the first control component and the second control component, and outputs a fused control component.
[0027] As a further improvement of the system of the application, the molten pool temperature monitoring module comprises an infrared thermal imaging camera and a temperature signal processing, the laser cladding monitoring platform based on the infrared thermal imaging camera adopts a side-axis monitoring, one infrared thermal imaging camera is erected outside the cavity of the laser cladding printer, and the collected molten pool image picture completely contains the region to be monitored; the temperature signal processing corrects the emissivity of the infrared camera for the molten pool of the printing material 316L stainless steel.
[0028] As a further improvement of the system of the application, the disturbance observation module comprises a control signal acquisition unit and a temperature change characteristic extraction unit, the control signal acquisition unit acquires the power value of the output signal of the laser power execution module; the temperature change characteristic extraction unit performs first-order difference calculation on the temperature signal of the molten pool temperature monitoring module to obtain the temperature change rate;
[0029]
[0030] The temperature gradient is combined to calculate the temperature difference between adjacent sampling points, and an observer is input to obtain an estimated thermal disturbance
[0031]
[0032] Wherein the state observer initializes the nominal temperature T_nom as T_actual(k), updates T_nom through a recursive formula, and other recursive formulas are defined as:
[0033]
[0034] As a further improvement of the system of the application, the sliding mode control module comprises a sliding mode surface unit and a chattering suppression unit, the sliding mode surface unit generates a sliding mode surface according to a temperature error and an integral term, and the sliding mode surface in the sliding mode control module is defined as follows:
[0035] s=e+λ∫edt;
[0036] Wherein, λ=0.25 is a sliding mode surface weight coefficient, used to balance the dynamic response of the error term and the integral term, and e is a temperature error composed of an actual temperature and a preset temperature; the chattering suppression unit replaces the sign function sgn(s) with a saturation function to suppress the chattering phenomenon of the sliding mode control, and the saturation function in the sliding mode control module is defined as follows:
[0037] sat(s / φ), φ=18;
[0038] Wherein φ = 18 is the boundary layer thickness, when s ≤ 18, sat (s / 18) = s / 18, that is, linear output; when s > 18, sat (s / 18) = sign (s), that is, saturated output, so as to suppress the chattering phenomenon of sliding mode control.
[0039] As a further improvement of the system of the application, the reinforcement learning PID module comprises a state definition unit, a Q table storage unit and a parameter updating unit, the state definition unit discretizes the temperature error e and the error change rate ec into 3 levels of states; the Q table storage unit initializes a 5-dimensional Q table 3x3x3x3x3; the parameter updating unit updates the parameters according to a reward function, and the reward function is defined as
[0040]
[0041] Wherein Δe is the error change, and e is the temperature error. The PID parameters are optimized by a Q learning update rule, and the Q learning update rule is defined as:
[0042] Q (s, a) ← (1-α) Q (s, a) + α [R + γ · max Q (s') ] ;
[0043] Wherein α = 0.1 · exp (-t / 20) is the learning rate, which decays exponentially with time t; and γ = 0.95 is the discount factor, which focuses on long-term rewards.
[0044] As a further improvement of the system of the application, the hybrid strategy module comprises an error amplitude analysis unit and a weight distribution unit, the error amplitude analysis unit acquires the absolute value of the molten pool temperature error |eT| in real time, and compares it with the preset fusion center value blend_center and the transition bandwidth blend_width, and outputs the classification result of the error amplitude; the weight distribution unit realizes the smooth distribution of the weight by a tanh function, and the definition is:
[0045]
[0046] blend ratio = 0.5 · (tanh (tanh arg )< blend_center) + 1) ;
[0047] Wherein, blend_ratio is the fusion weight of the SMC control component and the PID control component, 0 ≤ blend_ratio ≤ 1, when |eT| < 12.5℃, blend_ratio smoothly transitions from 0 (pure PID) to 0.5, and SMC and PID each account for 50%; when |eT| = 12.5°C, blend_ratio = 0.5; when |eT| > 12.5℃, blend_ratio smoothly transitions from 0.5 to 1, pure SMC, and the output fusion control quantity Pl of the hybrid strategy moduleaser is defined as:
[0048] Plaser = blend_ratio Psmc + (1-blend_ratio) Ppid
[0049] wherein P laser is the laser output power, P smc is the sliding mode control output power, and P pid is the PID control output power.
[0050] The method of the laser cladding molten pool temperature reinforcement learning-sliding mode PID collaborative control system of the application comprises the following steps:
[0051] S1: The molten pool temperature monitoring module acquires the actual temperature signal of the molten pool in the laser cladding process through an infrared thermal imaging camera, and calculates the temperature error with the preset temperature value;
[0052] S2: The control signal acquisition unit of the disturbance observation module acquires the output signal of the laser power execution module; the temperature change characteristic extraction unit performs first-order difference calculation on the actual temperature in combination with the temperature gradient (calculated by the temperature difference of adjacent sampling points), to obtain the temperature change rate, which is input into the observer to obtain the estimated thermal disturbance;
[0053] S3: The sliding mode control module receives the temperature error and disturbance compensation signal, and generates the first control component through the anti-interference control;
[0054] S4: The reinforcement learning PID module generates the second control component according to the temperature error through the Q learning and other reinforcement learning mechanisms to online optimize the proportional, integral, and differential parameters;
[0055] S5: The hybrid control module analyzes the amplitude characteristics of the temperature error, adaptively fuses the first control component and the second control component according to the preset rule, and outputs the fused control amount;
[0056] S6: The laser power actuator converts the fused control amount into a laser power adjustment instruction, which acts on the laser cladding equipment to realize closed-loop control of the molten pool temperature.
[0057] As a further improvement of the method of the application, in step S1, the infrared thermal imaging camera is installed in a side-shaft manner;
[0058] In step S2, the anti-interference control strategy of the sliding mode control module includes designing a sliding mode surface and suppressing chattering through a saturation function to ensure fast convergence to the target temperature in a large error state;
[0059] In the step S4, the autonomous learning mechanism of the reinforcement learning PID module includes: taking the temperature error e and the error change rate ec as states, taking the PID parameter (kp, ki, kd) adjustment as actions, taking the temperature tracking error square sum as a reward function, updating the PID parameter through Q learning, and optimizing the control accuracy.
[0060] In the step S5, the adaptive fusion rule of the mixed strategy module is that: the weight of the second control component and the weight of the first control component are calculated through a tanh function, and a fusion control quantity is output.
[0061] In the step S6, the laser power adjustment instruction contains a power increment constraint, ensuring the safety of equipment operation.
[0062] In the above technical solution, the laser cladding molten pool temperature reinforcement learning-sliding mode PID collaborative control system and method provided by the application has the following beneficial effects:
[0063] Multi-mode collaborative control improves robustness and accuracy: through the mixed strategy of the sliding mode control module and the reinforcement learning PID module, the strong anti-interference characteristic (lambda = 0.25 in the sliding mode surface design balances the error and the integral term, accelerates the convergence and suppresses the steady-state error) of the sliding mode control and the fine adjustment ability of the PID control are combined, and the control weight is adaptively distributed through the error threshold judgment unit, the sliding mode control is mainly used to ensure rapid correction when the molten pool temperature error is large, and the PID control is mainly used to improve the steady-state accuracy when the error is small, which is significantly better than a single control mode, and solves the problems of chattering of traditional sliding mode control and poor adaptability of PID control to nonlinear disturbance.
[0064] Reinforcement learning adaptive optimization reduces the cost of artificial parameter adjustment: the reinforcement learning PID module optimizes the PID parameters (Kp, Ki, Kd) online through the Q learning mechanism, and combines the gradient reward design of the reward function (the smaller the error reduction, the higher the reward, and the error stagnation is punished), to guide the system to adapt to complex working conditions such as laser power fluctuation and molten pool thermal inertia change, avoiding the lag and subjectivity of traditional PID requiring artificial repeated parameter adjustment, and improving the process stability and intelligent level.
[0065] Sliding mode chattering suppression technology guarantees control smoothness: the saturation function is used to replace the traditional sign function in the sliding mode control module, effectively solving the problem of frequent actuator action caused by the sign function in the sliding mode control, reducing the mutation amplitude of the laser power adjustment instruction, prolonging the service life of the equipment, and improving the stability of the molten pool temperature control.
[0066] Multi-module collaborative implementation of real-time closed-loop control: the disturbance observation module quickly estimates external thermal disturbance through the state observer and the temperature change rate calculation, and provides compensation input for the sliding mode control. BRIEF DESCRIPTION OF DRAWINGS
[0067] Figure 1 The structural block diagram of the laser cladding molten pool temperature reinforcement learning-sliding mode PID collaborative control system disclosed in the embodiment of the application is shown in the figure;
[0068] Figure 2 The control flowchart in the embodiment of the application is shown in the figure;
[0069] Figure 3 The structural schematic diagram of the additive manufacturing platform in the embodiment of the application is shown in the figure;
[0070] Figure 4 The installation schematic diagram of the infrared thermal imaging machine in the embodiment of the application is shown in the figure;
[0071] Figure 5 The simulation diagram in the embodiment of the application is shown in the figure. DETAILED DESCRIPTION
[0072] The application will be described in further detail below in conjunction with the accompanying drawings and specific embodiments:
[0073] Exemplarily, as shown in the figure, the laser cladding molten pool temperature reinforcement learning-sliding mode PID collaborative control system of the application Figure 1
[0074] comprises a molten pool temperature monitoring module, a disturbance observation module, a sliding mode control module, a reinforcement learning PID control module and a hybrid strategy module, wherein the sliding mode control module, the reinforcement learning PID control module and the hybrid strategy module jointly constitute a hybrid controller, the disturbance observation module and the reinforcement learning module are connected to the input end of the hybrid controller, and the output end of the hybrid controller outputs a control quantity to act on the molten pool controlled system and the disturbance observer;
[0075] The molten pool temperature monitoring module is configured to collect an actual temperature signal of the molten pool in the laser cladding process in real time, and form a temperature error with a preset temperature value.
[0076] The disturbance observation module is configured to obtain a power value of the output signal of the laser power execution module, and perform first-order difference calculation on the temperature signal of the molten pool temperature monitoring module to obtain a temperature change rate
[0077]
[0078] In combination with a temperature gradient (calculated by the temperature difference of adjacent sampling points), an estimated thermal disturbance is input into the observer to generate a disturbance compensation signal.
[0079]
[0080] The sliding mode control module receives the temperature error and disturbance compensation signal, generates a first control component with strong robustness through an anti-disturbance control strategy, and balances the dynamic response of the error term and the integral term by setting a sliding surface
[0081] s = e + l * integral e dt
[0082] The chattering phenomenon of the sliding mode control is suppressed by replacing the sign function sgn(s) with a chattering suppression unit saturation function.
[0083] The reinforcement learning PID module optimizes the Kp, Ki, and Kd parameters online through a Q learning mechanism.
[0084] Q(s, a) <- (1 - a) * Q(s, a) + a * [R + gamma * max Q(s')]
[0085] The second control component is generated according to eT
[0086] The hybrid strategy module analyzes the amplitude characteristics of the temperature error, adaptively fuses the first control component and the second control component, and outputs a fused control quantity.
[0087] Plaser = blend_ratio * Psmc + (1 - blend_ratio) * Ppid
[0088] Where P laser is the output power of the laser, P smc is the output power of the sliding mode control, and P pid is the output power of the PID control. The smooth allocation of weights is realized through a tanh function, which is defined as:
[0089]
[0090] blend ratio = 0.5 * (tanh(tanh arg ) + 1)
[0091] Where blend_ratio is the fusion weight of the SMC control component and the PID control component (0 <= blend_ratio <= 1). When |eT| < 12.5℃, blend_ratio smoothly transitions from 0 (pure PID) to 0.5 (SMC and PID each accounting for 50%); when |eT| = 12.5℃, blend_ratio = 0.5; when |eT| > 12.5℃, blend_ratio smoothly transitions from 0.5 to 1 (pure SMC)
[0092] Exemplarily, the control flow chart is as shown in Figure 2 , and the infrared thermal imaging machine installation schematic diagram is as shown in Figure 4A laser cladding molten pool temperature reinforcement learning-sliding mode PID collaborative control method is shown in the present application.
[0093] S1: The molten pool temperature monitoring module acquires the actual temperature signal of the molten pool in the laser cladding process through the infrared thermal imaging camera, and calculates the temperature error with the preset temperature value.
[0094] S2: The control signal acquisition unit of the disturbance observer module acquires the output signal of the laser power execution module; the temperature change feature extraction unit performs first-order difference calculation on the actual temperature combined with the temperature gradient (calculated by the temperature difference of adjacent sampling points), obtains the temperature change rate, and inputs it into the observer to obtain the estimated thermal disturbance.
[0095] S3: The sliding mode control module receives the temperature error and disturbance compensation signal, and generates the first control component through anti-interference control.
[0096] S4: The reinforcement learning PID module optimizes the proportional, integral, and differential parameters online through reinforcement learning mechanisms such as Q learning, and generates the second control component according to the temperature error.
[0097] S5: The hybrid control module analyzes the amplitude characteristics of the temperature error, adaptively fuses the first control component and the second control component according to the preset rules, and outputs the fused control quantity.
[0098] S6: The laser power actuator converts the fused control quantity into laser power adjustment instructions, which act on the laser cladding equipment to realize closed-loop control of the molten pool temperature.
[0099] Exemplarily, referring to Figure 3 As shown, the experimental device for laser cladding molten pool temperature reinforcement learning-sliding mode PID collaborative control of the present application includes a laser 1, a printed part 2, a substrate 3, a powder feeder 4, a printing cabin 5, a protective gas device 6, an upper computer 7, an infrared thermal imaging camera 8, a water cooler 9, and a control panel 10. The water cooler is connected to the laser 1 for water cooling, the powder feeder 4 is connected to the laser for powder feeding during printing, the printed part 2 is printed on the fixed substrate 3, and the infrared thermal imaging camera 4 is fixed on the front side of the printing cabin 5 with a support. The upper computer 7 is connected to the control panel 10 and the infrared thermal imaging camera 8, the wire feeder 8 is connected to the welding gun 7 for continuous supply of welding wire, and the protective gas device 6 delivers protective gas through the printing cabin 5 to prevent chemical reactions such as oxidation.
[0100] When the experiment begins, first load the printing program through the control panel 10 and the host computer 7. Among them, the printing parameters include: laser power 900w, scanning speed 9mm / s, powder feeding rate 8g / s, preset temperature 1950°. Fix the base plate 3 on the platform, set the coordinate zero, open the protection gas device 6. The gas of the protection gas device 9 is pure argon with a concentration of 100%, and the flow rate is set to 15L / min. The infrared thermal imaging camera 8 starts monitoring and capturing the real-time temperature of the molten pool during the printing process. After receiving the parameters, the laser 1 starts to generate laser and moves along the predetermined path to print. According to the real-time temperature of the molten pool image transmitted by the infrared thermal imaging camera 8, the controller program in the host computer 7 continuously controls and adjusts the laser power according to the actual temperature transmitted, and transmits it to the control panel 10 to control the laser 1. By continuously adjusting the parameters, the predetermined additive manufacturing path is completed, and the control of the molten pool temperature is realized. The simulation diagram in the specific embodiment is shown in Figure 5 .
[0101] It should be further pointed out that the above listed embodiments are only used to illustrate the technical principles and application modes of the present application, and do not constitute a limitation on the protection scope of the present application. Although the present application has been described in detail in combination with a plurality of preferred embodiments, those skilled in the art should understand that various forms of modification, deformation or equivalent replacement of related technical details can still be made without deviating from the core technical idea of the present application. For example, simplification of the graph construction method, replacement of the graph neural network structure, switching of the reinforcement learning algorithm, adjustment of the interface protocol between modules, etc. should be regarded as reasonable technical extensions of the present application. Therefore, any equivalent change or replacement within the scope of the spirit and technical principles of the present application should be covered within the protection scope of the present application.
Claims
1. Reinforcement learning-sliding mode PID collaborative control system for laser cladding molten pool temperature, characterized by: It includes molten pool temperature monitoring module, disturbance observation module, sliding mode control module, reinforcement learning PID control module and hybrid strategy module; The molten pool temperature monitoring module is used to collect the actual temperature signal of the molten pool during the laser cladding process in real time, and to form a temperature error with the preset temperature value; The disturbance observation module is used to estimate the system thermal disturbance in real time and generate the disturbance compensation signal based on the laser power control signal and the temperature variation characteristics of the molten pool; The sliding mode control module receives the temperature error and the interference compensation signal, and generates a first control component with strong robustness through an anti-disturbance control strategy; The reinforcement learning PID module dynamically optimizes the proportional-integral-derivative parameters through an autonomous learning mechanism and generates a second control component based on the temperature error; The hybrid strategy module analyzes the amplitude characteristics of the temperature error, adaptively fuses the first control component and the second control component, and outputs a fused control quantity.
2. The reinforcement learning-sliding mode PID collaborative control system for laser cladding molten pool temperature according to claim 1 is characterized in that: The molten pool temperature monitoring module includes an infrared thermal imaging camera and temperature signal processing. The laser cladding monitoring platform based on the infrared thermal imaging camera adopts side-axis monitoring. An infrared thermal imaging camera is set up outside the cavity of the laser cladding printer. The collected molten pool image completely includes the area to be monitored; the temperature signal processing performs infrared camera emissivity correction on the molten pool of the printing material 316L stainless steel.
3. The reinforcement learning-sliding mode PID collaborative control system for laser cladding molten pool temperature according to claim 1 is characterized in that: The disturbance observation module includes a control signal acquisition unit and a temperature change feature extraction unit. The control signal acquisition unit obtains the power value of the output signal of the laser power execution module; the temperature change feature extraction unit performs a first-order difference calculation on the temperature signal of the molten pool temperature monitoring module to obtain the temperature change rate; Combined with the temperature gradient, the temperature difference between adjacent sampling points is calculated and input into the observer to obtain the estimated thermal disturbance The state observer initializes the nominal temperature T_nom to T_actual(k) and updates T_nom through the recursive formula. The other recursive formulas are defined as:
4. The reinforcement learning-sliding mode PID collaborative control system for laser cladding molten pool temperature according to claim 1, characterized in that: The sliding mode control module includes a sliding surface unit and a chattering suppression unit. The sliding surface unit generates a sliding surface according to the temperature error and the integral term. The sliding surface in the sliding mode control module is defined as follows: s=e+λ∫e dt; Wherein, λ=0.25 is the sliding mode surface weight coefficient, which is used to balance the dynamic response of the error term and the integral term, and e is the temperature error between the actual temperature and the preset temperature. The chattering suppression unit uses a saturation function instead of the sign function sgn(s) to suppress the chattering phenomenon of the sliding mode control. The saturation function in the sliding mode control module is defined as follows: sat(s / φ), φ=18; Where φ = 18 is the boundary layer thickness. When s ≤ 18, sat(s / 18) = s / 18, which is a linear output. When s > 18, sat(s / 18) = sign(s), which is a saturated output, thereby suppressing the chattering phenomenon of the sliding mode control.
5. The reinforcement learning-sliding mode PID collaborative control system for laser cladding molten pool temperature according to claim 1, characterized in that: The reinforcement learning PID module includes a state definition unit, a Q table storage unit and a parameter update unit. The state definition unit discretizes the temperature error e and the error change rate ec into three levels of state; the Q table storage unit initializes a 5-dimensional Q table 3×3×3×3×3; the parameter update unit updates the parameters according to the reward function, and the reward function is defined as Where Δe is the error change, e is the temperature error, and the PID parameters are optimized using the Q-learning update rule, which is defined as: Q(s,a)←(1-α)Q(s,a)+α[R+γ·maxQ(s')]; Where α = 0.1 exp(-t / 20) is the learning rate, which decays exponentially with time t; γ = 0.95 is the discount factor, which focuses on long-term rewards.
6. The reinforcement learning-sliding mode PID collaborative control system for laser cladding molten pool temperature according to claim 1, characterized in that: The hybrid strategy module includes an error amplitude analysis unit and a weight allocation unit. The error amplitude analysis unit collects the absolute value of the molten pool temperature error |eT| in real time, compares it with the preset fusion center value blend_center and transition band width blend_width, and outputs the classification result of the error amplitude; the weight allocation unit realizes the smooth distribution of weights through the tanh function, which is defined as: Blend ratio=0.5·(tanh(tanh arg )+1); Among them, blend_ratio is the fusion weight of SMC control component and PID control component 0≤blend_ratio≤1. When |eT|<12.5℃, blend_ratio smoothly transitions from 0 (pure PID) to 0.5, and SMC and PID each account for 50%; when |eT|=12.5°C, blend_ratio=0.5; when |eT|>12.5℃, blend_ratio smoothly transitions from 0.5 to 1, pure SMC, and the output fusion control quantity P of the hybrid strategy module l aser Defined as: P l aser =blend_ratio·P smc +(1-blend_ratio)·Ppid Among them, P laser is the laser output power, P smc is the output power of sliding mode control, P pid Output power for PID control.
7. A method for using the reinforcement learning-sliding mode PID collaborative control system for laser cladding molten pool temperature according to any one of claims 1 to 6, characterized in that: The following steps are involved: S1: The molten pool temperature monitoring module collects the actual temperature signal of the molten pool during the laser cladding process through an infrared thermal imaging camera, and calculates the temperature error compared with the preset temperature value; S2: The control signal acquisition unit of the disturbance observation module obtains the output signal of the laser power execution module; the temperature change feature extraction unit performs a first-order difference calculation on the actual temperature and combines it with the temperature gradient (calculated by the temperature difference between adjacent sampling points) to obtain the temperature change rate, which is input into the observer to obtain the estimated thermal disturbance; S3: The sliding mode control module receives the temperature error and the interference compensation signal and generates a first control component through anti-interference control; S4: The reinforcement learning PID module uses reinforcement learning mechanisms such as Q learning to optimize the proportional, integral, and differential parameters online and generate the second control component based on the temperature error; S5: The hybrid control module analyzes the amplitude characteristics of the temperature error, adaptively fuses the first control component and the second control component according to a preset rule, and outputs a fused control variable; S6: The laser power actuator converts the fusion control quantity into a laser power adjustment instruction, which acts on the laser cladding equipment to achieve closed-loop control of the molten pool temperature.
8. The method of reinforcement learning-sliding mode PID collaborative control system for laser cladding molten pool temperature according to claim 7, characterized in that: In step S1, the infrared thermal imaging camera is mounted using a rangefinder; In step S2, the anti-disturbance control strategy of the sliding mode control module includes designing a sliding mode surface and suppressing chattering through a saturation function to ensure that the target temperature is quickly approached under a large error state; In step S4, the autonomous learning mechanism of the reinforcement learning PID module includes: taking the temperature error e and the error change rate ec as the state, adjusting the PID parameters (kp, ki, kd) as the action, and the sum of squares of the temperature tracking error as the reward function, updating the PID parameters through Q learning to optimize the control accuracy; In step S5, the adaptive fusion rule of the hybrid strategy module is: calculating the weight of the second control component and the weight of the first control component by using a tanh function, and outputting a fusion control amount; In step S6, the laser power adjustment instruction includes a power increment constraint to ensure safe operation of the equipment.
Citation Information
Patent Citations
A method for controlling the temperature gradient of the molten pool during laser cladding.
CN112517926B