Laser welding control system and method based on OCT time-frequency spectrum and reinforcement learning

By using a five-layer closed-loop system based on OCT time spectrum and reinforcement learning, the problems of low utilization of monitoring information and single control strategy in laser welding are solved, achieving high quality and stability in the welding process and automatic adjustment to adapt to complex working conditions.

CN120949561AActive Publication Date: 2025-11-14NAMO INTELLIGENT TECHNOLOGY (SHANGHAI) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511056663.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-14
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

In existing laser welding processes, the utilization rate of OCT monitoring information is low, the echo signal is easily interfered with, and the control strategy is simple and lacks adaptive ability, resulting in unstable welding quality.

Method used

A five-layer real-time closed-loop system based on OCT time-frequency spectrum and reinforcement learning was designed, including an OCT acquisition module, a sliding window buffer module, a feature extraction module, a state tensor construction module, a policy encoding module, a reinforcement learning policy network module, and an instruction execution module. The system adjusts the laser power and welding speed in real time through multi-channel processing and deep learning.

Benefits of technology

It achieves precise response and optimized adjustment of welding status, improves the consistency of welding quality and system stability, and can automatically adapt to changes in working conditions and material differences, reducing the need for manual parameter adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120949561A_ABST
    Figure CN120949561A_ABST
Patent Text Reader

Abstract

The invention discloses a laser welding control system based on OCT time-frequency spectrum and reinforcement learning, and relates to the technical field of quality control in the laser machining process. Comprising an OCT acquisition module, a sliding window cache module, a feature extraction module, a state tensor construction module, a strategy coding module, a reinforcement learning strategy network module, an instruction mapping module and an instruction execution module. The invention further discloses a laser welding control method based on the OCT time-frequency spectrum and reinforcement learning. The laser welding control method comprises the steps that S100, OCT signals are collected and cached in a sliding mode; s200, time-frequency spectrum feature extraction; s300, state coding and strategy reasoning are carried out; s400, mapping and issuing a control instruction; s500, executing an instruction; and S600, closed-loop control is carried out. According to the method, the keyhole stability and the dynamic behavior of the molten pool are comprehensively reflected, accurate response and optimal adjustment of the welding state are achieved, and the consistency of the welding quality and the stability of system operation are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of quality control technology for laser processing, and in particular to a laser welding control system and method based on OCT time spectrum and reinforcement learning. Background Technology

[0002] Laser welding, with its advantages of being non-contact, having high energy density, and a small heat-affected zone, has been widely used in fields such as automotive manufacturing, aerospace, power batteries, and medical devices. However, the molten pool and keyhole exhibit extremely high instability during the welding process; even slight disturbances can lead to weld defects such as incomplete penetration, burn-through, porosity, or spatter. To ensure welding quality, industry increasingly relies on online monitoring and closed-loop control technologies to achieve stable adjustment of the weld penetration and real-time correction of the welding state.

[0003] Optical coherence tomography (OCT) has become an important tool for monitoring weld penetration depth in recent years due to its micrometer-level axial resolution and high sampling rate. Currently, the widely adopted strategy is typically based on maximum intensity projection, which extracts the location of the maximum echo intensity from the continuous OCTA-line signal as an estimate of the weld interface. This method is simple to implement, computationally inexpensive, and suitable for penetration depth estimation within a certain range. However, because it only utilizes the intensity information of the OCT signal and ignores the rich phase, spectral, and temporal variation characteristics of the interferometric signal, the estimation stability and accuracy are often severely affected when the signal-to-noise ratio decreases or the signal is distorted. Furthermore, different material types, surface roughness, laser reflectivity, and shielding gas flow conditions can significantly interfere with the OCT signal, resulting in a lack of sufficient robustness in the traditional maximum intensity projection method.

[0004] On the other hand, existing laser welding control systems mostly employ fixed power or simple PID loops for process regulation. This approach is slow to respond to process changes and cannot dynamically optimize control strategies based on weld conditions. Especially when facing complex conditions such as material inhomogeneity, thermal fluctuations, and changes in processing paths, frequent manual parameter adjustments are often required. Furthermore, while some welding control systems have attempted to introduce machine learning or data-driven methods for prediction and control, most are limited to offline modeling or static data fitting, lacking real-time response capabilities and adaptive strategy update capabilities.

[0005] In summary, current welding closed-loop control systems suffer from two major technical gaps: first, a lack of in-depth utilization of the dynamic high-dimensional features in OCT signals, relying solely on shallow projection, which fails to accurately perceive the true state of the weld; second, a simplistic control strategy lacks real-time strategy optimization capabilities, making it unsuitable for complex and ever-changing industrial scenarios. These issues limit the intelligence level of laser welding systems and hinder their widespread application in high-quality, high-stability welding tasks.

[0006] Therefore, those skilled in the art are dedicated to developing a laser welding control system and method based on OCT time spectrum and reinforcement learning. Summary of the Invention

[0007] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by this application is how to overcome the shortcomings of low utilization rate of OCT monitoring information, easy interference of echo signal, single control strategy and lack of adaptive capability in the existing laser welding process.

[0008] The applicant analyzed the laser welding control process, from signal acquisition, feature extraction, policy reasoning to control execution, and designed a complete real-time closed-loop system with a five-layer structure: data acquisition layer, data processing layer, state modeling layer, control decision layer, and execution feedback layer. The data acquisition layer includes an OCT acquisition module, the data processing layer includes a sliding window buffer module and a feature extraction module, the state modeling layer includes a state tensor construction module and a policy encoding module, the control decision layer includes a reinforcement learning policy network module and an instruction mapping module, and the execution feedback layer includes an instruction execution module.

[0009] In one embodiment of this application, a laser welding control system based on OCT time-spectrum and reinforcement learning is provided, comprising: The OCT acquisition module acquires A-line interference signals in the welding area in real time. The sliding window buffer module uses a double buffering mechanism to store A-line interference signals, ensuring that newly acquired data can be written when data is read. The feature extraction module performs multi-channel processing on the A-line interference signal using a parallel computing architecture to extract feature vectors. The state tensor construction module stacks feature vectors into a two-dimensional state tensor along the time dimension. The strategy encoding module performs tensor dimensionality reduction, encoding the two-dimensional state tensor into a fixed-length state vector. The reinforcement learning policy network module responds to the input of a fixed-length state vector and outputs continuous control actions; The instruction mapping module maps continuous control actions into laser power and welding speed adjustment commands. The instruction execution module executes the laser power and welding speed equipment adjustment instructions, thereby adjusting the laser power of the laser and the moving speed of the welding head displacement device; The OCT acquisition module is connected to the sliding window buffer module via a high-speed data bus. The sliding window buffer module, feature extraction module, state tensor construction module, and policy encoding module are sequentially connected. The policy encoding module is connected to the reinforcement learning policy network module via a high-speed internal interface. The reinforcement learning policy network module, instruction mapping module, and instruction execution module are connected. The instruction mapping module sends laser power and welding speed equipment adjustment instructions to the instruction execution module via an industrial real-time communication protocol.

[0010] Optionally, in the OCT time-spectrum and reinforcement learning-based laser welding control system in the above embodiments, the OCT acquisition module includes a frequency domain OCT probe.

[0011] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, the frequency domain OCT probe includes a swept frequency OCT probe and a spectral domain OCT probe.

[0012] Alternatively, in the OCT time-spectrum and reinforcement learning-based laser welding control system in any of the above embodiments, the high-speed data bus uses Camera Link or PCIe.

[0013] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, multi-channel processing includes amplitude envelope, FFT spectrum, wavelet coefficients, and phase difference.

[0014] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, parallel computing architectures such as CUDA and FPGA acceleration are used.

[0015] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, the shape of the two-dimensional state tensor is T×d, where T is the number of frames corresponding to the window length and d is the feature dimension.

[0016] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, the state tensor construction module avoids redundant data copying through memory sharing or zero-copy transmission.

[0017] Alternatively, in the OCT time-spectrum and reinforcement learning-based laser welding control system in any of the above embodiments, tensor dimensionality reduction uses a convolutional neural network (CNN) or a Transformer.

[0018] Optionally, in the OCT time-spectrum and reinforcement learning-based laser welding control system in any of the above embodiments, the high-speed internal interface includes a TensorRT inference engine.

[0019] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, the continuous control actions include laser power increment ΔP and welding speed increment ΔV.

[0020] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, the laser power and welding speed equipment adjustment commands include power 0% to 100% and speed 0-50 mm / s.

[0021] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, the industrial real-time communication protocols include EtherCAT and PROFINET.

[0022] Based on any of the above embodiments, another embodiment of this application provides a laser welding control method based on OCT time spectrum and reinforcement learning, including the following steps: S100, OCT signal acquisition and sliding buffer: The OCT acquisition module acquires A-line interference signals directly above the weld and sends them to the sliding window buffer module for buffering, and updates them in a sliding window manner; S200, Time-Spectrum Feature Extraction: The feature extraction module performs multi-channel feature processing on each frame of the A-line interferometric signal, generating a multi-dimensional feature vector x. The state tensor construction module connects multiple frames to form a two-dimensional state tensor. ; S300, State Encoding and Policy Reasoning, based on a deep reinforcement learning structure, uses a reinforcement learning policy network module to encode the two-dimensional state tensor. S The encoding is a fixed-length state vector z, which is then input into the reinforcement learning policy network. π (·), output control action a ; S400, Control Command Mapping and Issuance: The command mapping module controls the actions... a Perform normalization inverse transformation and rate limiting processing, map to process control instructions, and send them to the instruction execution module; S500, instruction execution, responds to process control instructions. The instruction execution module executes laser power and welding speed equipment adjustment instructions, adjusting the laser power of the laser and the moving speed of the welding head displacement device; S600, closed-loop control, cyclically executes steps S100-S500 until the laser welding task is completed and the welding process ends.

[0023] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in the above embodiments, the OCT acquisition module includes a frequency domain OCT probe.

[0024] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in any of the above embodiments, the frequency domain OCT probe includes a swept frequency OCT probe and a spectral domain OCT probe.

[0025] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in the above embodiments, step S100 includes: S110, OCT signal acquisition: A coaxial frequency domain OCT probe is used to continuously image the area directly above the weld and acquire A-line interference signals. S120. Store the A-line interference signal, cache the A-line interference signal in a time series manner, and update it in a sliding window manner.

[0026] Furthermore, in the laser welding control method based on OCT time-spectrum and reinforcement learning in the above embodiments, the sliding window length is set to... T frame, T The OCT acquisition time corresponding to the frame is less than the response cycle of the laser welding control system. Each time it is updated, it slides forward one frame to maintain the complete state information of the most recent welding process.

[0027] Furthermore, in the laser welding control method based on OCT time-frequency spectrum and reinforcement learning in the above embodiments, the complete state information of the most recent welding process is centered on the current frame, including the states before and after it. T / 2 frame.

[0028] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in the above embodiments, step S200 includes: S210. Amplitude feature extraction: Extract the depth location of the maximum echo point, and calculate the local contrast, edge gradient, and average signal-to-noise ratio estimate within a fixed window around the maximum echo point. S220. Spectral feature extraction: Perform a fast Fourier transform on each frame of the A-line interference signal to extract the position of the main frequency component, the energy ratio of the first three main frequencies, the position of the spectral centroid, the bandwidth, and the normalized entropy in the spectrum, forming a set of frequency domain descriptors. S230. Wavelet feature extraction: The Daubechies wavelet is used to decompose the A-line interference signal into multiple scales, retaining the detail components at multiple scales. The energy, number of peaks and number of zero intersections of the detail component coefficients at each scale are calculated to represent the complexity of the local texture. S240, Phase dynamic feature extraction: Through the phase unwrapping algorithm, the phase data of the current frame and the previous frame of the A-line interference signal are compared, and the average phase drift rate, maximum rate of change and continuous rising / falling segment length of the whole frame are calculated to reflect the vibration and motion stability of the keyhole interface. S250. Calculate the two-dimensional feature tensor. In the feature extraction module, after steps S210-S240, each frame of A-line interference signal generates a multi-dimensional feature vector. x The state tensor construction module connects multiple frames to form a two-dimensional state tensor. , as the state input for reinforcement learning.

[0029] Furthermore, in the laser welding control method based on OCT time spectrum and reinforcement learning in any of the above embodiments, the fixed window refers to the range of ±10 pixel values ​​centered on the maximum echo point.

[0030] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in any of the above embodiments, step S300 includes: S310. Calculate the fixed-length state vector. Based on the deep reinforcement learning structure, the reinforcement learning policy network module converts the two-dimensional state tensor... S Encoded as a fixed-length state vector z The formula is as follows: z = f_encoder(S) ; S320. Calculate the control action and input the fixed-length state vector z into the reinforcement learning policy network. π (·), output control action a The formula is as follows: a = π(z) ; Among them, control actions a It is a continuous variable vector representing the adjustment amount of process parameters.

[0031] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in any of the above embodiments, the deep reinforcement learning structure includes DDPG, PPO or SAC framework.

[0032] Furthermore, in the laser welding control method based on OCT time-frequency spectrum and reinforcement learning in the above embodiments, the control action... a Including laser power adjustment ΔP and scan speed adjustment ΔV Laser power adjustment ΔP The unit is a relative proportion or increment.

[0033] Furthermore, in the laser welding control method based on OCT time-frequency spectrum and reinforcement learning in the above embodiments, the reinforcement learning policy network... π (·) Training is completed during the offline training phase, with the training objective being to maximize the reward function: in, D_actualBased on the current estimated melt depth (taking the maximum echo point depth), D_target For the target melting depth, σ_ depth The standard deviation of the echo depth within this window. Q_stability The interface stability index is derived from the spectrum and phase. α 1. α 2. α 3 represents the task-weighted parameter; Network employing reinforcement learning strategies π The strategy is iteratively optimized using (·) and the valuation network Q(z, a), and after training, the network is frozen and deployed to run in an online system.

[0034] Furthermore, in the laser welding control method based on OCT time-spectrum and reinforcement learning in the above embodiments, the interface stability index... Q_stability The value range is [-1, 1], where positive values ​​represent a stable state and negative values ​​represent an unstable state.

[0035] Furthermore, in the laser welding control method based on OCT time-spectrum and reinforcement learning in the above embodiments, the task weighting parameters... α 1. α 2. α 3. Dynamically configure according to welding task requirements.

[0036] Optionally, in the laser welding control method based on OCT time-spectrum and reinforcement learning in the above embodiments, α 1 ∈[0.5, 2.0] , α 2 ∈[0.1, 1.0] , α 3 ∈[0.2, 1.5] .

[0037] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in any of the above embodiments, the process control instructions include the analog voltage or digital power setting value of the laser and the speed setting instructions of the scanning platform.

[0038] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in the above embodiments, step S400 includes: S410, Control Action Processing, Instruction Mapping Module for Control Actions a After performing normalized inverse transform and rate limiting, the laser power adjustment amount is obtained. ΔP and scan speed adjustment ΔV ; S420, process control command mapping, the process control commands include the analog voltage or digital power setting value of the laser and the speed setting command of the scanning platform, and the laser power adjustment amount. ΔPMapped to the laser's analog voltage or digital power setting, the scanning speed adjustment amount... ΔV Mapped to the speed setting command of the scanning platform; S430, Process control instructions are issued. The process control instructions are issued to the instruction execution module through the industrial communication protocol.

[0039] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in the above embodiments, the industrial communication protocols include EtherCAT and CANopen.

[0040] This application utilizes dynamic features in OCT interference signals, excluding the location of maximum intensity, such as phase drift, frequency domain energy distribution, and wavelet coefficients, to comprehensively reflect keyhole stability and molten pool dynamic behavior. By using high-dimensional time-frequency features as input to the control system, it can adjust laser power or scanning speed in real time, achieving precise response and optimized adjustment of the welding state. The control strategy possesses excellent generalization and online adaptability, eliminating the need for manual parameter tuning and automatically responding to changes in different material types, surface roughness, and processing paths, ensuring consistent welding quality and stable system operation. This application operates stably throughout the entire welding process, without relying on manual parameter tuning. It can adapt in real time to disturbances such as changes in working conditions, surface state fluctuations, and differences in material reflectivity, achieving stable control of weld depth and interface state, significantly improving the quality consistency, adaptability, and intelligence level of the welding process.

[0041] The following will further explain the concept, specific structure and technical effects of this application in conjunction with the accompanying drawings, so as to fully understand the purpose, features and effects of this application. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of the structure of a laser welding control system based on OCT time-frequency spectrum and reinforcement learning, which is an exemplary embodiment. Figure 2 This is a flowchart of an exemplary embodiment of a laser welding control method based on OCT time-frequency spectrum and reinforcement learning; Figure 3 This is a diagram showing the effect of OCT penetration monitoring in laser welding without feedback control. Figure 4 This is a diagram illustrating the effect of melt depth monitoring in an exemplary embodiment. Detailed Implementation

[0043] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of this application to make its technical content clearer and easier to understand. This application can be embodied in many different forms, and the scope of protection of this application is not limited to the embodiments mentioned herein.

[0044] In the accompanying drawings, components with the same structure are designated by the same numerical designation, and components with similar structures or functions are designated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and this application does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of components is schematically exaggerated in some places in the drawings.

[0045] The applicant designed a laser welding control system based on OCT time-spectrum and reinforcement learning, such as... Figure 1 As shown, it includes: The OCT acquisition module includes a frequency domain OCT probe, which is used to acquire A-line interference signals in the welding area in real time. The sliding window buffer module uses a double buffering mechanism to store A-line interference signals, ensuring that newly acquired data can be written when data is read. The feature extraction module performs multi-channel processing on the A-line interference signal through a parallel computing architecture accelerated by FPGA. The multi-channel processing includes amplitude envelope, FFT spectrum, wavelet coefficients, phase difference, and feature vector extraction. The state tensor construction module stacks feature vectors along the time dimension into a two-dimensional state tensor. The shape of the two-dimensional state tensor is... T×d , T The number of frames corresponding to the window length. d As a feature dimension, zero-copy transmission avoids redundant data duplication; The policy encoding module uses a convolutional neural network (CNN) to perform tensor dimensionality reduction, encoding the two-dimensional state tensor into a fixed-length state vector; The reinforcement learning policy network module, responding to an input with a fixed-length state vector, outputs continuous control actions, including laser power increments. ΔP Welding speed increment ΔV ; The instruction mapping module maps continuous control actions into equipment adjustment instructions for laser power and welding speed; The instruction execution module executes laser power and welding speed equipment adjustment instructions, including power 0% to 100% and speed 0-50 mm / s, thereby adjusting the laser power of the laser and the moving speed of the welding head displacement device. The OCT acquisition module is connected to the sliding window buffer module via a high-speed data bus using PCIe. The sliding window buffer module, feature extraction module, state tensor construction module, and policy encoding module are sequentially connected. The policy encoding module is connected to the reinforcement learning policy network module via the TensorRT inference engine. The reinforcement learning policy network module, instruction mapping module, and instruction execution module are connected. The instruction mapping module sends laser power and welding speed equipment adjustment instructions to the instruction execution module via EtherCAT.

[0046] Based on the above embodiments, the applicant provides a laser welding control method based on OCT time spectrum and reinforcement learning, such as... Figure 2 As shown, it includes the following steps: S100, OCT signal acquisition and sliding buffer: A frequency-domain OCT probe acquires A-line interference signals directly above the weld seam, sends them to the sliding window buffer module for buffering, and updates them using a sliding window method; specifically including: S110, OCT signal acquisition: A coaxial frequency domain OCT probe is used to continuously image the area directly above the weld and acquire A-line interference signals. S120. Store the A-line interference signal, cache the A-line interference signal in a time-series manner, and update it using a sliding window method, with the sliding window length set to [value missing]. T frame, T The OCT acquisition time corresponding to each frame is less than the response cycle of the aforementioned laser welding control system. Each update slides forward one frame to maintain the complete state information of the most recent welding process. The complete state information of the most recent welding process is centered on the current frame, with frames before and after it. T / 2 frame.

[0047] S200, Time-Spectrum Feature Extraction: The feature extraction module performs multi-channel feature processing on each frame of the A-line interference signal to generate a multi-dimensional feature vector. x The state tensor construction module connects multiple frames to form a two-dimensional state tensor. Specifically, it includes: S210, Amplitude feature extraction: Extract the depth location of the maximum echo point, and calculate the local contrast, edge gradient and average signal-to-noise ratio estimate within a fixed window around the maximum echo point. The fixed window refers to the range of ±10 pixel values ​​centered on the maximum echo point. S220. Spectral feature extraction: Perform a fast Fourier transform on each frame of the A-line interference signal to extract the position of the main frequency component, the energy ratio of the first three main frequencies, the position of the spectral centroid, the bandwidth, and the normalized entropy in the spectrum, forming a set of frequency domain descriptors. S230. Wavelet feature extraction: The Daubechies wavelet is used to decompose the A-line interference signal into multiple scales, retaining the detail components at multiple scales. The energy, number of peaks and number of zero intersections of the detail component coefficients at each scale are calculated to represent the complexity of the local texture. S240, Phase dynamic feature extraction: Through the phase unwrapping algorithm, the phase data of the current frame and the previous frame of the A-line interference signal are compared, and the average phase drift rate, maximum rate of change and continuous rising / falling segment length of the whole frame are calculated to reflect the vibration and motion stability of the keyhole interface. S250. Calculate the two-dimensional feature tensor. In the feature extraction module, after steps S210-S240, each frame of A-line interference signal generates a multi-dimensional feature vector. x The state tensor construction module connects multiple frames to form a two-dimensional state tensor. , as the state input for reinforcement learning.

[0048] S300, State Encoding and Policy Reasoning, based on a deep reinforcement learning structure, uses a reinforcement learning policy network module to encode the two-dimensional state tensor. S Encoded as a fixed-length state vector z And input the reinforcement learning policy network π (·), output control action a Specifically, it includes: S310. Calculate the fixed-length state vector. Based on the SAC framework, the reinforcement learning policy network module converts the two-dimensional state tensor... S Encoded as a fixed-length state vector z The formula is as follows: z = f_encoder(S) ; S320, Calculate the control actions and convert the fixed-length state vector... z Input reinforcement learning policy network π (·), output control action a The formula is as follows: a = π(z) ; Among them, control actions a It is a continuous variable vector representing the adjustment amount of process parameters and control actions. a Including laser power adjustment ΔP and scan speed adjustment ΔV Laser power adjustment ΔP The unit is a relative proportion or increment; reinforcement learning strategy network π (·) Training is completed during the offline training phase, with the training objective being to maximize the reward function: in,D_actual This is an estimate of the current melt depth (taking the maximum echo point depth). D_target Target melting depth; σ_ depth This represents the standard deviation of the echo depth within that window. Q_stability Interface stability metrics are derived from the spectrum and phase. Q_stability The range of values ​​is [-1, 1] Positive values ​​indicate a stable state, while negative values ​​indicate an unstable state. α 1. α 2. α 3 represents the task-weighted parameter, which is dynamically configured according to the welding task requirements. α 1 ∈[0.5, 2.0],α 2 ∈[0.1, 1.0],α 3 ∈[0.2, 1.5] ; Network employing reinforcement learning strategies π (·) and valuation network Q(z, a) The strategy is iteratively optimized, and after training, the network is frozen and deployed to run in an online system.

[0049] S400, Control Command Mapping and Issuance: The command mapping module controls the actions... a Perform normalization inverse transformation and rate limiting processing, map to process control instructions, and issue them to the instruction execution module; specifically including: S410, Control Action Processing, Instruction Mapping Module for Control Actions a After performing normalized inverse transform and rate limiting, the laser power adjustment amount is obtained. ΔP and scan speed adjustment ΔV ; S420, process control command mapping, the process control commands include the analog voltage or digital power setpoint of the laser, the speed setting command of the scanning platform, and the laser power adjustment amount. ΔP Mapped to the laser's digital power setting, the scanning speed adjustment value is... ΔV Mapped to the speed setting command of the scanning platform; S430, Process control instructions are issued. The process control instructions are issued to the instruction execution module via EtherCAT.

[0050] S500, instruction execution, responds to process control instructions. The instruction execution module executes laser power and welding speed equipment adjustment instructions, adjusting the laser power of the laser and the moving speed of the welding head displacement device.

[0051] S600, closed-loop control, cyclically executes steps S100-S500 until the laser welding task is completed and the welding process ends.

[0052] To verify the technical effectiveness of the above embodiments, the applicant conducted experiments and effect comparisons using the existing laser welding OCT penetration depth monitoring method without feedback control and the laser welding control method based on OCT time spectrum and reinforcement learning of the above embodiments. The laser welding test material was 304 stainless steel with dimensions of 5 mm × 8 mm × 50 mm. The welding length was set to 20 mm. Two identical pieces of material were joined together, and linear laser welding was performed along the longest direction of the gap formed by the joining of the materials, i.e., 50 mm. The initial welding speed was set to 5 mm / s, the initial laser power was set to 1 kW, and the target welding penetration depth was set to 3 mm.

[0053] like Figure 3 As shown, without closed-loop control, the penetration depth will fluctuate slightly as welding progresses. Figure 4 As shown, with the addition of closed-loop feedback control, the welding power and welding speed are adjusted in real time, and the weld penetration approaches stability. It is evident that the above embodiment uses high-dimensional time-frequency characteristics as input to the control system, enabling real-time adjustment of laser power or scanning speed, achieving precise response and optimized adjustment of the welding state.

[0054] The preferred embodiments of this application have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of this application without inventive effort. Therefore, any technical solutions that can be obtained by those skilled in the art based on the concept of this application through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A laser welding control system based on OCT time-spectrum and reinforcement learning, characterized in that, include: The OCT acquisition module acquires A-line interference signals in the welding area in real time. The sliding window buffer module uses a double buffering mechanism to store the A-line interference signal, ensuring that newly acquired data can be written when data is read. The feature extraction module performs multi-channel processing on the A-line interference signal using a parallel computing architecture to extract feature vectors; The state tensor construction module stacks the feature vectors into a two-dimensional state tensor along the time dimension; The strategy encoding module performs tensor dimensionality reduction, encoding the two-dimensional state tensor into a fixed-length state vector. The reinforcement learning policy network module responds to the input of the fixed-length state vector and outputs continuous control actions; The instruction mapping module maps the continuous control actions into laser power and welding speed adjustment instructions; The instruction execution module executes the laser power and welding speed equipment adjustment instructions, thereby adjusting the laser power of the laser and the moving speed of the welding head displacement device; The OCT acquisition module is connected to the sliding window buffer module via a high-speed data bus. The sliding window buffer module, the feature extraction module, the state tensor construction module, and the policy encoding module are sequentially connected. The policy encoding module is connected to the reinforcement learning policy network module via a high-speed internal interface. The reinforcement learning policy network module, the instruction mapping module, and the instruction execution module are connected. The instruction mapping module sends the laser power and welding speed equipment adjustment instructions to the instruction execution module via an industrial real-time communication protocol.

2. The laser welding control system based on OCT time-spectrum and reinforcement learning as described in claim 1, characterized in that, The OCT acquisition module includes a frequency domain OCT probe.

3. The laser welding control system based on OCT time-spectrum and reinforcement learning as described in claim 1, characterized in that, The multi-channel processing includes amplitude envelope, FFT spectrum, wavelet coefficients, and phase difference.

4. The laser welding control system based on OCT time-spectrum and reinforcement learning as described in claim 1, characterized in that, The state tensor construction module avoids redundant data copying through memory sharing or zero-copy transmission.

5. The laser welding control system based on OCT time-spectrum and reinforcement learning as described in claim 1, characterized in that, The continuous control actions include laser power increments. ΔP Welding speed increment ΔV .

6. A laser welding control method based on OCT time-spectrum and reinforcement learning, using the laser welding control system based on OCT time-spectrum and reinforcement learning as described in any one of claims 1-5, characterized in that, Includes the following steps: S100, OCT signal acquisition and sliding buffer: The OCT acquisition module acquires A-line interference signals directly above the weld and sends them to the sliding window buffer module for buffering, and updates them in a sliding window manner; S200. Time-spectrum feature extraction: The feature extraction module performs multi-channel feature processing on each frame of the A-line interference signal to generate a multi-dimensional feature vector. x The state tensor construction module connects multiple frames to form a two-dimensional state tensor. ; S300, State Encoding and Policy Reasoning, based on a deep reinforcement learning structure, wherein the reinforcement learning policy network module encodes the two-dimensional state tensor. S Encoded as a fixed-length state vector z And input the reinforcement learning policy network π (·), output control action a ; S400, Control command mapping and issuance: the command mapping module maps and issues the control actions. a Perform normalization inverse transformation and rate limiting processing, map to process control instructions, and send them to the instruction execution module; S500, Instruction Execution: In response to the process control instruction, the instruction execution module executes the laser power and welding speed equipment adjustment instruction to adjust the laser power of the laser and the moving speed of the welding head displacement device; S600, Closed-loop control, cyclically execute steps S100-S500 until the laser welding task is completed and the welding process ends.

7. The laser welding control method based on OCT time-frequency spectrum and reinforcement learning as described in claim 6, characterized in that, Step S100 includes: S110, OCT signal acquisition: A coaxial frequency domain OCT probe is used to continuously image the area directly above the weld and acquire A-line interference signals. S120. Store the A-line interference signal, cache the A-line interference signal in a time series manner, and update it in a sliding window manner.

8. The laser welding control method based on OCT time-spectrum and reinforcement learning as described in claim 6 or 7, characterized in that, Step S200 includes: S210. Amplitude feature extraction: Extract the depth location of the maximum echo point, and calculate the local contrast, edge gradient, and average signal-to-noise ratio estimate within a fixed window around the maximum echo point. S220. Spectral feature extraction: Perform a fast Fourier transform on each frame of the A-line interference signal to extract the position of the main frequency component, the energy ratio of the first three main frequencies, the position of the spectral centroid, the bandwidth, and the normalized entropy in the spectrum, forming a set of frequency domain descriptors. S230. Wavelet feature extraction: The A-line interference signal is decomposed into multiple scales using Daubechies wavelets, retaining detail components at multiple scales. The energy, number of peaks, and number of zero intersections of the detail component coefficients at each scale are calculated to represent the complexity of the local texture. S240. Phase dynamic feature extraction: By using a phase unwrapping algorithm, the phase data of the current frame and the previous frame of the A-line interference signal are compared, and the average phase drift rate, maximum rate of change and continuous rising / falling segment length of the whole frame are calculated to reflect the vibration and motion stability of the keyhole interface. S250. Calculate the two-dimensional feature tensor. After steps S210-S240 in the feature extraction module, a multi-dimensional feature vector is generated for each frame of A-line interference signal. x The state tensor construction module connects multiple frames to form a two-dimensional state tensor. , as the state input for reinforcement learning.

9. The laser welding control method based on OCT time-frequency spectrum and reinforcement learning as described in claim 8, characterized in that, Step S300 includes: S310. Calculate the fixed-length state vector. Based on the deep reinforcement learning structure, the reinforcement learning policy network module converts the two-dimensional state tensor... S Encoded as a fixed-length state vector z The formula is as follows: z = f_encoder(S) ; S320. Calculate the control action and convert the fixed-length state vector... z Input the reinforcement learning policy network π (·), output control action a The formula is as follows: a = π(z) .

10. The laser welding control method based on OCT time-spectrum and reinforcement learning as described in claim 9, characterized in that, Step S400 includes: S410, Control action processing, the instruction mapping module processes the control action. a After performing normalized inverse transform and rate limiting, the laser power adjustment amount is obtained. ΔP and scan speed adjustment ΔV ; S420, Process control command mapping, the process control commands include the analog voltage or digital power setting value of the laser and the speed setting command of the scanning platform, and the laser power adjustment amount. ΔP Mapped to the analog voltage or digital power setting of the laser, the scanning speed adjustment amount ΔV Mapped to the speed setting command of the scanning platform; S430, Process control instructions are issued, and the process control instructions are issued to the instruction execution module through an industrial communication protocol.

Citation Information

Patent Citations

  • Deep reinforcement learning spectrum sharing method based on priority experience replay

    CN112383922A

  • Laser welding planning method, device and equipment based on automatic positioning

    CN118875491A

  • Laser welding spatter real-time identification method based on OCT keyhole depth measurement

    CN119426794A

  • Apparatus for cutting hairpin of hairpin stator

    KR1020230011132A

  • Add-on module for interposing between a control device and a laser machining head of a laser machining system

    WO2021130044A1