A laser welding control system and method based on OCT time-frequency spectrum and reinforcement learning
By using a laser welding control system based on OCT time-frequency spectrum and reinforcement learning, real-time control is achieved by utilizing high-dimensional time-frequency features. This solves the problems of low utilization rate of OCT monitoring information and single control strategy in existing technologies, and realizes the improvement of stability and intelligence in the welding process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2026-04-07
Smart Images

Figure CN120949561B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of quality control of laser processing, in particular to a laser welding control system and method based on OCT time-frequency spectrum and reinforcement learning. BACKGROUND
[0002] Laser welding has been widely used in the fields of automobile manufacturing, aerospace, power battery and medical devices due to its non-contact, high energy density and small heat-affected zone. However, the molten pool and keyhole have extremely high instability during the welding process, and slight disturbance may lead to welding defects such as incomplete penetration, burn-through, porosity or spatter. In order to ensure the welding quality, the industry increasingly relies on online monitoring and closed-loop control technology in the process to achieve stable adjustment of the penetration and real-time correction of the welding state.
[0003] Optical coherence tomography (OCT) has become an important means for monitoring welding penetration in recent years due to its micron-level axial resolution and high sampling rate. The strategy widely used at present is usually based on maximum intensity projection, that is, the position of the maximum echo intensity is extracted from the continuous OCT A-line signal as the estimation of the weld interface. This method is simple to implement, has low computational overhead, and is suitable for penetration estimation within a certain range. However, since only the intensity information of the OCT signal is used and the rich phase, spectral and timing variation characteristics in the interference signal are ignored, the estimation stability and accuracy are often severely affected when the signal-to-noise ratio is reduced or the signal is distorted. At the same time, different material types, surface roughness, laser reflectivity and protective gas flow state can significantly interfere with the OCT signal, resulting in a lack of sufficient robustness of the traditional maximum projection method.
[0004] On the other hand, existing laser welding control systems mostly use fixed power or simple PID loops for process adjustment. This approach is slow to react to process changes and cannot dynamically optimize the control strategy according to the weld state, especially when facing complex conditions such as uneven materials, thermal fluctuations, and changes in processing paths, often requiring frequent manual parameter tuning. In addition, some welding control systems have attempted to introduce machine learning or data-driven methods for prediction and control, but most are limited to offline modeling or static data fitting, lacking real-time response capability and adaptive strategy updating capability.
[0005] In summary, there are two major technical gaps in current welding closed-loop control systems: first, there is a lack of deep utilization of dynamic high-dimensional features in the OCT signal, relying only on shallow projection, which cannot accurately perceive the real state of the weld; second, the control strategy is single and does not have real-time strategy optimization capability, which cannot adapt to complex and variable industrial scenarios. These problems limit the intelligent level of laser welding systems and restrict their application in high-quality and high-stability welding tasks.
[0006] Therefore, the person skilled in the art is committed to developing a laser welding control system and method based on OCT time-frequency spectrum and reinforcement learning. SUMMARY
[0007] In view of the above defects of the prior art, the technical problem to be solved by the present application is how to overcome the defects of low utilization rate of OCT monitoring information, echo signal susceptible to interference, single control strategy and lack of adaptive ability in the existing laser welding process.
[0008] The applicant analyzes the process of laser welding control, from signal acquisition, feature extraction, strategy reasoning to control execution, and designs a complete real-time closed-loop system with five-layer structure of data acquisition layer, data processing layer, state modeling layer, control decision layer and execution feedback layer. The data acquisition layer includes an OCT acquisition module, the data processing layer includes a sliding window buffer module and a feature extraction module, the state modeling layer includes a state tensor construction module and a strategy encoding module, the control decision layer includes a reinforcement learning strategy network module and an instruction mapping module, and the execution feedback layer includes an instruction execution module.
[0009] In one embodiment of the present application, a laser welding control system based on OCT time-frequency spectrum and reinforcement learning is provided, comprising:
[0010] an OCT acquisition module, which acquires A-line interference signals of a welding area in real time;
[0011] a sliding window buffer module, which stores the A-line interference signals using a double buffering mechanism to ensure that new data can be written when data is read;
[0012] a feature extraction module, which extracts a feature vector by processing the A-line interference signals in multiple channels through a parallel computing architecture;
[0013] a state tensor construction module, which stacks the feature vector into a two-dimensional state tensor according to the time dimension;
[0014] a strategy encoding module, which encodes the two-dimensional state tensor into a fixed-length state vector by dimension reduction;
[0015] a reinforcement learning strategy network module, which outputs continuous control actions in response to the input of the fixed-length state vector;
[0016] an instruction mapping module, which maps the continuous control actions into laser power and welding speed adjustment instructions;
[0017] an instruction execution module, which executes the laser power and welding speed device adjustment instructions to adjust the laser power of the laser and the moving speed of the welding head displacement device;
[0018] The OCT acquisition module is connected with the sliding window buffer module through a high-speed data bus, the sliding window buffer module, the feature extraction module, the state tensor construction module and the strategy coding module are sequentially connected in communication, the strategy coding module is connected with the reinforcement learning strategy network module through a high-speed internal interface, the reinforcement learning strategy network module, the instruction mapping module and the instruction execution module are connected in communication, and the instruction mapping module sends the laser power and welding speed equipment adjustment instruction to the instruction execution module through an industrial real-time communication protocol.
[0019] Optionally, in the laser welding control system based on OCT time-frequency spectrum and reinforcement learning in the above embodiment, the OCT acquisition module comprises a frequency domain OCT probe.
[0020] Optionally, in the laser welding control system based on OCT time-frequency spectrum and reinforcement learning in any of the above embodiments, the frequency domain OCT probe comprises a swept source OCT probe and a spectral domain OCT probe.
[0021] Optionally, in the laser welding control system based on OCT time-frequency spectrum and reinforcement learning in any of the above embodiments, the high-speed data bus uses Camera Link or PCIe.
[0022] Optionally, in the laser welding control system based on OCT time-frequency spectrum and reinforcement learning in any of the above embodiments, the multi-channel processing includes amplitude envelope, FFT spectrum, wavelet coefficient and phase difference.
[0023] Optionally, in the laser welding control system based on OCT time-frequency spectrum and reinforcement learning in any of the above embodiments, the parallel computing architecture is CUDA or FPGA acceleration.
[0024] Optionally, in the laser welding control system based on OCT time-frequency spectrum and reinforcement learning in any of the above embodiments, the two-dimensional state tensor has a shape of Txd, T is the number of frames corresponding to the window length, and d is the feature dimension.
[0025] Optionally, in the laser welding control system based on OCT time-frequency spectrum and reinforcement learning in any of the above embodiments, the state tensor construction module avoids data redundancy copying through memory sharing or zero-copy transmission.
[0026] Optionally, in the laser welding control system based on OCT time-frequency spectrum and reinforcement learning in any of the above embodiments, the tensor dimension reduction uses convolutional neural network (CNN) or Transformer.
[0027] Optionally, in the laser welding control system based on OCT time-frequency spectrum and reinforcement learning in any of the above embodiments, the high-speed internal interface comprises a TensorRT inference engine.
[0028] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, the continuous control actions include laser power increment ΔP and welding speed increment ΔV.
[0029] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, the laser power and welding speed equipment adjustment commands include power 0% to 100% and speed 0-50 mm / s.
[0030] Optionally, in the laser welding control system based on OCT time spectrum and reinforcement learning in any of the above embodiments, the industrial real-time communication protocols include EtherCAT and PROFINET.
[0031] Based on any of the above embodiments, another embodiment of this application provides a laser welding control method based on OCT time spectrum and reinforcement learning, including the following steps:
[0032] S100, OCT signal acquisition and sliding buffer: The OCT acquisition module acquires A-line interference signals directly above the weld and sends them to the sliding window buffer module for buffering, and updates them in a sliding window manner;
[0033] S200, Time-Spectrum Feature Extraction: The feature extraction module performs multi-channel feature processing on each frame of the A-line interferometric signal, generating a multi-dimensional feature vector x. The state tensor construction module connects multiple frames to form a two-dimensional state tensor. S∈ ^ {T x d} ,in, T x d The shape of the two-dimensional state tensor. T The number of frames corresponding to the window length. d For feature dimensions;
[0034] S300, State Encoding and Policy Reasoning, based on a deep reinforcement learning structure, uses a reinforcement learning policy network module to encode the two-dimensional state tensor. S The encoding is a fixed-length state vector z, which is then input into the reinforcement learning policy network. π (·), output control action a ;
[0035] S400, Control Command Mapping and Issuance: The command mapping module controls the actions... a Perform normalization inverse transformation and rate limiting processing, map to process control instructions, and send them to the instruction execution module;
[0036] S500, instruction execution, responds to process control instructions. The instruction execution module executes laser power and welding speed equipment adjustment instructions, adjusting the laser power of the laser and the moving speed of the welding head displacement device;
[0037] S600, closed-loop control, cyclically executes steps S100-S500 until the laser welding task is completed and the welding process ends.
[0038] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in the above embodiments, the OCT acquisition module includes a frequency domain OCT probe.
[0039] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in any of the above embodiments, the frequency domain OCT probe includes a swept frequency OCT probe and a spectral domain OCT probe.
[0040] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in the above embodiments, step S100 includes:
[0041] S110, OCT signal acquisition: A coaxial frequency domain OCT probe is used to continuously image the area directly above the weld and acquire A-line interference signals.
[0042] S120. Store the A-line interference signal, cache the A-line interference signal in a time series manner, and update it in a sliding window manner.
[0043] Furthermore, in the laser welding control method based on OCT time-spectrum and reinforcement learning in the above embodiments, the sliding window length is set to... T frame, T The OCT acquisition time corresponding to the frame is less than the response cycle of the laser welding control system. Each time it is updated, it slides forward one frame to maintain the complete state information of the most recent welding process.
[0044] Furthermore, in the laser welding control method based on OCT time-frequency spectrum and reinforcement learning in the above embodiments, the complete state information of the most recent welding process is centered on the current frame, including the states before and after it. T / 2 frame.
[0045] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in the above embodiments, step S200 includes:
[0046] S210. Amplitude feature extraction: Extract the depth location of the maximum echo point, and calculate the local contrast, edge gradient, and average signal-to-noise ratio estimate within a fixed window around the maximum echo point.
[0047] S220. Spectral feature extraction: Perform a fast Fourier transform on each frame of the A-line interference signal to extract the position of the main frequency component, the energy ratio of the first three main frequencies, the position of the spectral centroid, the bandwidth, and the normalized entropy in the spectrum, forming a set of frequency domain descriptors.
[0048] S230. Wavelet feature extraction: The Daubechies wavelet is used to decompose the A-line interference signal into multiple scales, retaining the detail components at multiple scales. The energy, number of peaks and number of zero intersections of the detail component coefficients at each scale are calculated to represent the complexity of the local texture.
[0049] S240, Phase dynamic feature extraction: Through the phase unwrapping algorithm, the phase data of the current frame and the previous frame of the A-line interference signal are compared, and the average phase drift rate, maximum rate of change and continuous rising / falling segment length of the whole frame are calculated to reflect the vibration and motion stability of the keyhole interface.
[0050] S250. Calculate the two-dimensional feature tensor. In the feature extraction module, after steps S210-S240, each frame of A-line interference signal generates a multi-dimensional feature vector. x The state tensor construction module connects multiple frames to form a two-dimensional state tensor. S ∈ {T x d} , as the state input for reinforcement learning.
[0051] Furthermore, in the laser welding control method based on OCT time spectrum and reinforcement learning in any of the above embodiments, the fixed window refers to the range of ±10 pixel values centered on the maximum echo point.
[0052] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in any of the above embodiments, step S300 includes:
[0053] S310. Calculate the fixed-length state vector. Based on the deep reinforcement learning structure, the reinforcement learning policy network module converts the two-dimensional state tensor... S Encoded as a fixed-length state vector z The formula is as follows:
[0054] z = f encoder(S) ;
[0055] S320. Calculate the control action and input the fixed-length state vector z into the reinforcement learning policy network. π (·), output control action a The formula is as follows: a = π(z) ;
[0056] Among them, control actions aIt is a continuous variable vector representing the adjustment amount of process parameters.
[0057] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in any of the above embodiments, the deep reinforcement learning structure includes DDPG, PPO or SAC framework.
[0058] Furthermore, in the laser welding control method based on OCT time-frequency spectrum and reinforcement learning in the above embodiments, the control action... a Including laser power adjustment ΔP and scan speed adjustment ΔV Laser power adjustment ΔP The unit is a relative proportion or increment.
[0059] Furthermore, in the laser welding control method based on OCT time-frequency spectrum and reinforcement learning in the above embodiments, the reinforcement learning policy network... π (·) Training is completed during the offline training phase, with the training objective being to maximize the reward function:
[0060] R= –α 1 x |D actual - D target| - a 2 x s depth + a 3 x Q stability
[0061] in, D actual Based on the current estimated melt depth (taking the maximum echo point depth), D target For the target melting depth, s depth Q stability The standard deviation of the echo depth within this window. Q stability The interface stability index is derived from the spectrum and phase. α 1. α 2. α 3 represents the task-weighted parameter;
[0062] Network employing reinforcement learning strategies π The strategy is iteratively optimized using (·) and the valuation network Q(z, a), and after training, the network is frozen and deployed to run in an online system.
[0063] Furthermore, in the laser welding control method based on OCT time-spectrum and reinforcement learning in the above embodiments, the interface stability index... ΔP The value range is [-1, 1], where positive values represent a stable state and negative values represent an unstable state.
[0064] Furthermore, in the laser welding control method based on OCT time-spectrum and reinforcement learning in the above embodiments, the task weighting parameters... α 1. α 2. α 3. Dynamically configure according to welding task requirements.
[0065] Optionally, in the laser welding control method based on OCT time-spectrum and reinforcement learning in the above embodiments, α 1 ∈[0.5, 2.0] , α 2 ∈[0.1, 1.0] , α 3 ∈[0.2, 1.5] .
[0066] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in any of the above embodiments, the process control instructions include the analog voltage or digital power setting value of the laser and the speed setting instructions of the scanning platform.
[0067] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in the above embodiments, step S400 includes:
[0068] S410, Control Action Processing, Instruction Mapping Module for Control Actions a After performing normalized inverse transform and rate limiting, the laser power adjustment amount is obtained. ΔV and scan speed adjustment ΔP ;
[0069] S420, process control command mapping, the process control commands include the analog voltage or digital power setting value of the laser and the speed setting command of the scanning platform, and the laser power adjustment amount. ΔV Mapped to the laser's analog voltage or digital power setting, the scanning speed adjustment amount... Figure 1 Mapped to the speed setting command of the scanning platform;
[0070] S430, Process control instructions are issued. The process control instructions are issued to the instruction execution module through the industrial communication protocol.
[0071] Optionally, in the laser welding control method based on OCT time spectrum and reinforcement learning in the above embodiments, the industrial communication protocols include EtherCAT and CANopen.
[0072] This application utilizes dynamic features in OCT interference signals, excluding the location of maximum intensity, such as phase drift, frequency domain energy distribution, and wavelet coefficients, to comprehensively reflect keyhole stability and molten pool dynamic behavior. By using high-dimensional time-frequency features as input to the control system, it can adjust laser power or scanning speed in real time, achieving precise response and optimized adjustment of the welding state. The control strategy possesses excellent generalization and online adaptability, eliminating the need for manual parameter tuning and automatically responding to changes in different material types, surface roughness, and processing paths, ensuring consistent welding quality and stable system operation. This application operates stably throughout the entire welding process, without relying on manual parameter tuning. It can adapt in real time to disturbances such as changes in working conditions, surface state fluctuations, and differences in material reflectivity, achieving stable control of weld depth and interface state, significantly improving the quality consistency, adaptability, and intelligence level of the welding process.
[0073] The following will further explain the concept, specific structure and technical effects of this application in conjunction with the accompanying drawings, so as to fully understand the purpose, features and effects of this application. Attached Figure Description
[0074] Figure 2 This is a schematic diagram of the structure of a laser welding control system based on OCT time-frequency spectrum and reinforcement learning, which is an exemplary embodiment.
[0075] Figure 3 This is a flowchart of an exemplary embodiment of a laser welding control method based on OCT time-frequency spectrum and reinforcement learning;
[0076] Figure 4 This is a diagram showing the effect of OCT penetration monitoring in laser welding without feedback control.
[0077] Figure 1 This is a diagram illustrating the effect of melt depth monitoring in an exemplary embodiment. Detailed Implementation
[0078] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of this application to make its technical content clearer and easier to understand. This application can be embodied in many different forms, and the scope of protection of this application is not limited to the embodiments mentioned herein.
[0079] In the accompanying drawings, components with the same structure are designated by the same numerical designation, and components with similar structures or functions are designated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and this application does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of components is schematically exaggerated in some places in the drawings.
[0080] The applicant designed a laser welding control system based on OCT time-spectrum and reinforcement learning, such as... T x dAs shown, it includes:
[0081] The OCT acquisition module includes a frequency domain OCT probe, which is used to acquire A-line interference signals in the welding area in real time.
[0082] The sliding window buffer module uses a double buffering mechanism to store A-line interference signals, ensuring that newly acquired data can be written when data is read.
[0083] The feature extraction module performs multi-channel processing on the A-line interference signal through a parallel computing architecture accelerated by FPGA. The multi-channel processing includes amplitude envelope, FFT spectrum, wavelet coefficients, phase difference, and feature vector extraction.
[0084] The state tensor construction module stacks feature vectors along the time dimension into a two-dimensional state tensor. The shape of the two-dimensional state tensor is... ΔP , T The number of frames corresponding to the window length. d As a feature dimension, zero-copy transmission avoids redundant data duplication;
[0085] The policy encoding module uses a convolutional neural network (CNN) to perform tensor dimensionality reduction, encoding the two-dimensional state tensor into a fixed-length state vector;
[0086] The reinforcement learning policy network module, responding to an input with a fixed-length state vector, outputs continuous control actions, including laser power increments. ΔV Welding speed increment Figure 2 ;
[0087] The instruction mapping module maps continuous control actions into equipment adjustment instructions for laser power and welding speed;
[0088] The instruction execution module executes laser power and welding speed equipment adjustment instructions, including power 0% to 100% and speed 0-50 mm / s, thereby adjusting the laser power of the laser and the moving speed of the welding head displacement device.
[0089] The OCT acquisition module is connected to the sliding window buffer module via a high-speed data bus using PCIe. The sliding window buffer module, feature extraction module, state tensor construction module, and policy encoding module are sequentially connected. The policy encoding module is connected to the reinforcement learning policy network module via the TensorRT inference engine. The reinforcement learning policy network module, instruction mapping module, and instruction execution module are connected. The instruction mapping module sends laser power and welding speed equipment adjustment instructions to the instruction execution module via EtherCAT.
[0090] Based on the above embodiments, the applicant provides a laser welding control method based on OCT time spectrum and reinforcement learning, such as... {T x d} As shown, it includes the following steps:
[0091] S100, OCT signal acquisition and sliding buffer: A frequency-domain OCT probe acquires A-line interference signals directly above the weld seam, sends them to the sliding window buffer module for buffering, and updates them using a sliding window method; specifically including:
[0092] S110, OCT signal acquisition: A coaxial frequency domain OCT probe is used to continuously image the area directly above the weld and acquire A-line interference signals.
[0093] S120. Store the A-line interference signal, cache the A-line interference signal in a time-series manner, and update it using a sliding window method, with the sliding window length set to [value missing]. T frame, T The OCT acquisition time corresponding to each frame is less than the response cycle of the aforementioned laser welding control system. Each update slides forward one frame to maintain the complete state information of the most recent welding process. The complete state information of the most recent welding process is centered on the current frame, with frames before and after it. T / 2 frame.
[0094] S200, Time-Spectrum Feature Extraction: The feature extraction module performs multi-channel feature processing on each frame of the A-line interference signal to generate a multi-dimensional feature vector. x The state tensor construction module connects multiple frames to form a two-dimensional state tensor. S∈ ^ T x d ,in, {T x d} The shape of the two-dimensional state tensor. T The number of frames corresponding to the window length. d For feature dimensions; specifically including:
[0095] S210, Amplitude feature extraction: Extract the depth location of the maximum echo point, and calculate the local contrast, edge gradient and average signal-to-noise ratio estimate within a fixed window around the maximum echo point. The fixed window refers to the range of ±10 pixel values centered on the maximum echo point.
[0096] S220. Spectral feature extraction: Perform a fast Fourier transform on each frame of the A-line interference signal to extract the position of the main frequency component, the energy ratio of the first three main frequencies, the position of the spectral centroid, the bandwidth, and the normalized entropy in the spectrum, forming a set of frequency domain descriptors.
[0097] S230. Wavelet feature extraction: The Daubechies wavelet is used to decompose the A-line interference signal into multiple scales, retaining the detail components at multiple scales. The energy, number of peaks and number of zero intersections of the detail component coefficients at each scale are calculated to represent the complexity of the local texture.
[0098] S240, Phase dynamic feature extraction: Through the phase unwrapping algorithm, the phase data of the current frame and the previous frame of the A-line interference signal are compared, and the average phase drift rate, maximum rate of change and continuous rising / falling segment length of the whole frame are calculated to reflect the vibration and motion stability of the keyhole interface.
[0099] S250. Calculate the two-dimensional feature tensor. In the feature extraction module, after steps S210-S240, each frame of A-line interference signal generates a multi-dimensional feature vector. x The state tensor construction module connects multiple frames to form a two-dimensional state tensor. S ∈ z = f encoder(S) , as the state input for reinforcement learning.
[0100] S300, State Encoding and Policy Reasoning, based on a deep reinforcement learning structure, uses a reinforcement learning policy network module to encode the two-dimensional state tensor. S Encoded as a fixed-length state vector z And input the reinforcement learning policy network π (·), output control action a Specifically, this includes:
[0101] S310. Calculate the fixed-length state vector. Based on the SAC framework, the reinforcement learning policy network module converts the two-dimensional state tensor... S Encoded as a fixed-length state vector z The formula is as follows:
[0102] ΔP ;
[0103] S320, Calculate the control action and convert the fixed-length state vector. z Input reinforcement learning policy network π (·), output control action a The formula is as follows: a = π(z) ;
[0104] Among them, control actions a It is a continuous variable vector representing the adjustment amount of process parameters and control actions. a Including laser power adjustment ΔV and scan speed adjustment ΔP Laser power adjustment x |D actual - D target| - aThe unit is a relative proportion or increment; reinforcement learning strategy network π (·) Training is completed during the offline training phase, with the training objective being to maximize the reward function:
[0105] R= –α 1 x s depth + a 2 x Q stability 3 D actual ,
[0106] in, D target This is an estimate of the current melt depth (taking the maximum echo point depth). s depth Target melting depth; Q stability Q stability This represents the standard deviation of the echo depth within that window. Q(z, a) Interface stability metrics are derived from the spectrum and phase. ΔP The range of values is [-1, 1] Positive values indicate a stable state, while negative values indicate an unstable state. α 1. α 2. α 3 represents the task-weighted parameter, which is dynamically configured according to the welding task requirements. α 1 ∈[0.5, 2.0],α 2 ∈[0.1, 1.0],α 3 ∈[0.2, 1.5] ;
[0107] Network employing reinforcement learning strategies π (·) and valuation network ΔV The strategy is iteratively optimized, and after training, the network is frozen and deployed to run in an online system.
[0108] S400, Control Command Mapping and Issuance: The command mapping module controls the actions... a Perform normalization inverse transformation and rate limiting processing, map to process control instructions, and issue them to the instruction execution module; specifically including:
[0109] S410, Control Action Processing, Instruction Mapping Module for Control Actions a After performing normalized inverse transform and rate limiting, the laser power adjustment amount is obtained. ΔP and scan speed adjustment ΔV ;
[0110] S420, process control command mapping, the process control commands include the analog voltage or digital power setpoint of the laser, the speed setting command of the scanning platform, and the laser power adjustment amount. Figure 3 Mapped to the laser's digital power setting, the scanning speed adjustment value is... Figure 4 Mapped to the speed setting command of the scanning platform;
[0111] S430, Process control instructions are issued. The process control instructions are issued to the instruction execution module via EtherCAT.
[0112] S500, instruction execution, responds to process control instructions. The instruction execution module executes laser power and welding speed equipment adjustment instructions, adjusting the laser power of the laser and the moving speed of the welding head displacement device.
[0113] S600, closed-loop control, cyclically executes steps S100-S500 until the laser welding task is completed and the welding process ends.
[0114] To verify the technical effectiveness of the above embodiments, the applicant conducted experiments and effect comparisons using the existing laser welding OCT penetration depth monitoring method without feedback control and the laser welding control method based on OCT time spectrum and reinforcement learning of the above embodiments. The laser welding test material was 304 stainless steel with dimensions of 5 mm × 8 mm × 50 mm. The welding length was set to 20 mm. Two identical pieces of material were joined together, and linear laser welding was performed along the longest direction of the gap formed by the joining of the materials, i.e., 50 mm. The initial welding speed was set to 5 mm / s, the initial laser power was set to 1 kW, and the target welding penetration depth was set to 3 mm.
[0115] like As shown, without closed-loop control, the penetration depth will fluctuate slightly as welding progresses. As shown, with the addition of closed-loop feedback control, the welding power and welding speed are adjusted in real time, and the weld penetration approaches stability. It is evident that the above embodiment uses high-dimensional time-frequency characteristics as input to the control system, enabling real-time adjustment of laser power or scanning speed, achieving precise response and optimized adjustment of the welding state.
[0116] The preferred embodiments of this application have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of this application without inventive effort. Therefore, any technical solutions that can be obtained by those skilled in the art based on the concept of this application through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A laser welding control system based on OCT time-spectrum and reinforcement learning, characterized in that, include: The OCT acquisition module acquires A-line interference signals in the welding area in real time. The sliding window caching module uses a double buffering mechanism to store the A-line interference signal, ensuring that newly acquired data can be written when data is read. The feature extraction module performs multi-channel processing on the A-line interference signal using a parallel computing architecture to extract feature vectors; The state tensor construction module stacks the feature vectors into a two-dimensional state tensor along the time dimension; The strategy encoding module performs tensor dimensionality reduction, encoding the two-dimensional state tensor into a fixed-length state vector. The reinforcement learning policy network module responds to the input of the fixed-length state vector and outputs continuous control actions; The instruction mapping module maps the continuous control actions into laser power and welding speed adjustment instructions; The instruction execution module executes the laser power and welding speed equipment adjustment instructions, thereby adjusting the laser power of the laser and the moving speed of the welding head displacement device; The OCT acquisition module is connected to the sliding window cache module via a high-speed data bus. The sliding window cache module, the feature extraction module, the state tensor construction module, and the policy encoding module are sequentially connected. The policy encoding module is connected to the reinforcement learning policy network module via a high-speed internal interface. The reinforcement learning policy network module, the instruction mapping module, and the instruction execution module are connected. The instruction mapping module sends the laser power and welding speed equipment adjustment instructions to the instruction execution module via an industrial real-time communication protocol.
2. The laser welding control system based on OCT time-spectrum and reinforcement learning as described in claim 1, characterized in that, The OCT acquisition module includes a frequency domain OCT probe.
3. The laser welding control system based on OCT time-spectrum and reinforcement learning as described in claim 1, characterized in that, The multi-channel processing includes amplitude envelope, FFT spectrum, wavelet coefficients, and phase difference.
4. The laser welding control system based on OCT time-spectrum and reinforcement learning as described in claim 1, characterized in that, The state tensor construction module avoids redundant data copying through memory sharing or zero-copy transmission.
5. The laser welding control system based on OCT time-spectrum and reinforcement learning as described in claim 1, characterized in that, The continuous control actions include laser power increments. ΔP Welding speed increment ΔV .
6. A laser welding control method based on OCT time-spectrum and reinforcement learning, using the laser welding control system based on OCT time-spectrum and reinforcement learning as described in any one of claims 1-5, characterized in that, Includes the following steps: S100, OCT signal acquisition and sliding buffer: The OCT acquisition module acquires A-line interference signals directly above the weld and sends them to the sliding window buffer module for buffering, and updates them in a sliding window manner; S200. Time-spectrum feature extraction: The feature extraction module performs multi-channel feature processing on each frame of the A-line interference signal to generate a multi-dimensional feature vector. x The state tensor construction module connects multiple frames to form a two-dimensional state tensor. S∈ ^{T×d} ,in, T×d The shape of the two-dimensional state tensor. T The number of frames corresponding to the window length. d For feature dimensions; S300, State Encoding and Policy Reasoning, based on a deep reinforcement learning structure, wherein the reinforcement learning policy network module encodes the two-dimensional state tensor. S Encoded as a fixed-length state vector z And input the reinforcement learning policy network π (·), output control action a ; S400, Control command mapping and issuance: the command mapping module maps and issues the control actions. a Perform normalization inverse transformation and rate limiting processing, map to process control instructions, and send them to the instruction execution module; S500, Instruction Execution: In response to the process control instruction, the instruction execution module executes the laser power and welding speed equipment adjustment instruction to adjust the laser power of the laser and the moving speed of the welding head displacement device; S600, Closed-loop control, cyclically execute steps S100-S500 until the laser welding task is completed and the welding process ends.
7. The laser welding control method based on OCT time-frequency spectrum and reinforcement learning as described in claim 6, characterized in that, Step S100 includes: S110, OCT signal acquisition: A coaxial frequency domain OCT probe is used to continuously image the area directly above the weld and acquire A-line interference signals. S120. Store the A-line interference signal, cache the A-line interference signal in a time series manner, and update it in a sliding window manner.
8. The laser welding control method based on OCT time-spectrum and reinforcement learning as described in claim 6 or 7, characterized in that, Step S200 includes: S210. Amplitude feature extraction: Extract the depth location of the maximum echo point, and calculate the local contrast, edge gradient, and average signal-to-noise ratio estimate within a fixed window around the maximum echo point. S220. Spectral feature extraction: Perform a fast Fourier transform on each frame of the A-line interference signal to extract the position of the main frequency component, the energy ratio of the first three main frequencies, the position of the spectral centroid, the bandwidth, and the normalized entropy in the spectrum, forming a set of frequency domain descriptors. S230. Wavelet feature extraction: The A-line interference signal is decomposed into multiple scales using Daubechies wavelets, retaining detail components at multiple scales. The energy, number of peaks, and number of zero intersections are calculated for the detail component coefficients at each scale to represent the complexity of the local texture. S240. Phase dynamic feature extraction: By using a phase unwrapping algorithm, the phase data of the current frame and the previous frame of the A-line interference signal are compared, and the average phase drift rate, maximum rate of change and continuous rising / falling segment length of the whole frame are calculated to reflect the vibration and motion stability of the keyhole interface. S250. Calculate the two-dimensional feature tensor. After steps S210-S240 in the feature extraction module, a multi-dimensional feature vector is generated for each frame of A-line interference signal. x The state tensor construction module connects multiple frames to form a two-dimensional state tensor. S ∈ ^{T×d} , as the state input for reinforcement learning.
9. The laser welding control method based on OCT time-frequency spectrum and reinforcement learning as described in claim 8, characterized in that, Step S300 includes: S310. Calculate the fixed-length state vector. Based on the deep reinforcement learning structure, the reinforcement learning policy network module converts the two-dimensional state tensor... S Encoded as a fixed-length state vector z The formula is as follows: z = f_encoder(S) ; S320. Calculate the control action and convert the fixed-length state vector... z Input the reinforcement learning policy network π (·), output control action a The formula is as follows: a = π(z) .
10. The laser welding control method based on OCT time-spectrum and reinforcement learning as described in claim 9, characterized in that, Step S400 includes: S410, Control action processing, the instruction mapping module processes the control action. a After performing normalized inverse transform and rate limiting, the laser power adjustment amount is obtained. ΔP and scan speed adjustment ΔV ; S420, Process control command mapping, the process control commands include the analog voltage or digital power setting value of the laser and the speed setting command of the scanning platform, and the laser power adjustment amount. ΔP Mapped to the analog voltage or digital power setting of the laser, the scanning speed adjustment amount ΔV Mapped to the speed setting command of the scanning platform; S430, Process control instructions are issued, and the process control instructions are issued to the instruction execution module through an industrial communication protocol.
Citation Information
Patent Citations
Laser welding spatter real-time identification method based on OCT keyhole depth measurement
CN119426794A
KR20210091789A