Self-supervision diagnosis method for welding quality abnormity

By employing a self-supervised diagnostic method guided by multi-view high frame rate acquisition and dynamic semantic anchor points, the problems of feature drift and insufficient robustness during the welding process are solved, achieving efficient identification of welding quality anomalies and extending the model lifecycle.

CN121458686APending Publication Date: 2026-02-03JUXIN ELECTRONICS TECH MEIZHOU CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511624386.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing intelligent sensing and diagnosis technologies for welding processes have significant shortcomings in dealing with dynamic working conditions, adaptive semantic alignment, and online robust anomaly detection. The feature space is prone to distribution drift, and it relies on manual annotation or simulation samples. Traditional feature alignment and anomaly detection modules cannot achieve sensitive adaptation to the current welding feature fluctuation trend and process disturbance.

Method used

We employ multi-view high frame rate acquisition, synchronization, precise ROI region focusing, and spatial and illumination normalization preprocessing to construct a lightweight semantic anchor generator. We design a differentiable semantic constraint loss function and achieve dynamic alignment and semantic stabilization of the feature space through dynamic semantic anchor guidance and temporal consistency mechanisms. We use sliding statistics and adaptive threshold determination algorithms to periodically detect feature space drift and perform incremental updates.

Benefits of technology

It improves the model's adaptability and generalization in real and complex production environments, suppresses model misjudgments caused by process fluctuations or occasional interference, increases the speed of capturing welding quality change trends, extends the model's lifespan, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121458686A_ABST
    Figure CN121458686A_ABST
Patent Text Reader

Abstract

The invention relates to a self-supervision diagnosis method for welding quality abnormity, and aims to solve the problems of unstable welding seam state characterization, high abnormity diagnosis false alarm rate, insufficient adaptability and the like caused by multi-view image acquisition time sequence difference, complex illumination interference and characteristic drift in the welding process. According to the core scheme, the method comprises the steps of synchronously collecting multi-view-angle continuous images of a welding area through multiple cameras at a high frame rate, conducting time sequence synchronization, area focusing, denoising and brightness normalization processing on the original images, extracting normalized frame-level features through a pre-trained convolutional neural network, improving feature space semantic stability through a semantic anchor point generation and dynamic alignment mechanism, and improving the accuracy of image fusion. And abnormal detection and multi-stage dynamic loss adjustment are introduced, so that welding quality abnormal judgment and model self-adaptive optimization are realized. According to the method, the consistency and robustness of welding visual representation can be effectively enhanced, and the accuracy of welding quality anomaly detection and the system generalization ability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of "intelligent sensing and self-supervised feature alignment technology in welding process", and more particularly to a self-supervised diagnostic method for welding quality anomalies. Background Technology

[0002] With the advancement of intelligent and digital manufacturing, online quality monitoring and anomaly diagnosis technologies for welding processes have received significant attention from the industry. Welding, as a critical industrial process, involves various materials, variable process parameters, and complex on-site environments, resulting in a highly dynamic and uncertain process state. To ensure welding quality, traditional methods often rely on manual visual inspection and offline testing, making it difficult to achieve real-time and accurate identification of sudden anomalies and performance drift during the welding process.

[0003] Existing intelligent sensing and diagnostic technologies for welding processes still have significant shortcomings in dealing with dynamic working conditions, adaptive semantic alignment, and online robust anomaly detection: First, the feature space is prone to distribution drift when welding materials, environment, or process parameters change, and existing models are difficult to dynamically adjust and maintain stable performance; second, they rely excessively on manual annotation or simulation samples; third, traditional feature alignment and anomaly detection modules often use static thresholds or one-way memory mechanisms, which cannot achieve sensitive adaptation to the current welding feature fluctuation trend and process disturbance. Summary of the Invention

[0004] This application provides a self-supervised diagnostic method for welding quality abnormalities, aiming to solve one of the problems or issues of the prior art mentioned in the background section.

[0005] This application provides a self-monitoring diagnostic method for welding quality abnormalities, specifically including: S1: Acquire image streams from multiple perspectives during the welding process. The image streams include continuous video frames of the weld pool area, the heat-affected zone, and the surrounding environment area to obtain a multimodal visual representation of the welding process.

[0006] S2: Perform denoising and normalization preprocessing on the acquired image stream to eliminate illumination changes and background interference, thereby improving the stability and consistency of subsequent feature extraction.

[0007] S3: Input the preprocessed image frame into the pre-trained visual feature extraction model to obtain the initial frame-level feature vector, which serves as the basis for subsequent semantic anchor generation and feature alignment.

[0008] S4: Based on the evolution law of weld morphology, a lightweight semantic anchor generator is constructed to identify key time points of molten pool boundary changes and heat-affected zone expansion, and generate reference anchor features with semantic stability.

[0009] S5: Design a differentiable semantic constraint loss function to calculate the cross-frame semantic consistency error between the current frame-level features and the nearest semantic anchor point, so as to drive the dynamic alignment and semantic stabilization of the feature space.

[0010] S6: Based on the frame-level feature alignment results guided by semantic anchors, update the parameters of the feature learning module to improve the robustness of the model's feature representation under changes in material, environment, or process parameters.

[0011] S7: Input the aligned frame-level feature vector into the anomaly detection module to perform welding quality anomaly judgment based on dynamic threshold, so as to identify potential welding defects caused by feature drift.

[0012] S8: Based on the anomaly detection results and the semantic anchor update status, dynamically adjust the weight parameters of the semantic constraint loss function to achieve an adaptive balance between improving recognition accuracy and controlling training stability.

[0013] S9: Periodically evaluate the semantic consistency index of the feature space. If it is determined to be a scenario with significant feature drift, trigger the semantic anchor update mechanism to extend the model life cycle and enhance generalization ability.

[0014] The self-monitoring diagnostic method for welding quality abnormalities provided in this application has the following beneficial effects: This invention innovatively designs a collaborative mechanism of "dynamic semantic anchor point guidance + temporal consistency." Specifically, it employs multi-view high frame rate acquisition, synchronization, precise ROI region focusing, and spatial and illumination normalization preprocessing to ensure high signal-to-noise input across the entire welding process for critical areas. It utilizes adaptive local feature maps and first-order difference operations on morphological evolution to automatically capture key evolutionary events in the weld pool and heat-affected zone without manual annotation, dynamically generating a representative set of semantic anchor points. Compared to existing methods that rely on simulation samples, offline memory libraries, or GAN-generated demonstration samples, this significantly improves the model's adaptability and generalization in real-world, complex production environments.

[0015] This invention establishes a cross-frame semantic consistency metric between frame-level features and semantic anchor features, using it as a loss term to drive the update of the deep feature extraction network. Compared to traditional unsupervised feature alignment schemes that cannot detect temporal drift or rely solely on single-step local comparisons, this method introduces a dynamic sliding window, contrastive learning loss, and a multi-factor weight self-adjustment mechanism throughout the process. This effectively prevents uncontrolled feature space drift when material or process mutations or anomalies occur, achieving a semantic consistency improvement of up to 7-12%, and greatly suppressing model misjudgments caused by process fluctuations / occasional interference.

[0016] By employing hardware-level frame synchronization, multi-channel high-speed data transmission, and local high-speed caching technology, frame-level alignment with an error of less than 1µs across different viewpoints and efficient, high-quality feature acquisition are achieved. Compared with existing single-viewpoint or low-frame-rate acquisition methods, this technology can capture weld abnormal signs more comprehensively and accurately, improving the speed of capturing welding quality change trends by more than 40%.

[0017] This invention employs a sliding statistics and adaptive thresholding algorithm to periodically detect feature space drift. When the drift is significant, it automatically triggers anchor point library reconstruction and incremental updates. Compared with traditional methods of iterative freezing or manual intervention verification, this significantly reduces maintenance costs and downtime risks, achieving extended model lifecycles and significantly improved stability.

[0018] In summary, this invention, with dynamic semantic anchors as its core, combined with temporal consistency, differentiable alignment, and model adaptation techniques, solves common problems in the existing field of welding visual diagnosis, such as uncontrollable feature drift, insufficient robustness, short lifecycle, and inaccurate anomaly detection. It has significant beneficial effects in terms of improving model performance, increasing system efficiency, reducing maintenance costs, and expanding the scope of application, and possesses high inventiveness and industrial application promotion value. Attached Figure Description

[0019] Appendix Figure 1 This is the main flowchart of a self-supervised diagnostic method for welding quality abnormalities.

[0020] Appendix Figure 2 This is a sub-flowchart of a self-supervised diagnostic method for welding quality anomalies.

[0021] Appendix Figure 3 This is another sub-flowchart of a self-supervised diagnostic method for welding quality anomalies. Detailed Implementation

[0022] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0023] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0024] As attached Figure 1 As shown, this application provides a self-monitoring diagnostic method for welding quality abnormalities, specifically including: S1: Acquire image streams from multiple perspectives during the welding process. The image streams include continuous video frames of the weld pool area, the heat-affected zone, and the surrounding environment area to obtain a multimodal visual representation of the welding process.

[0025] S2: Perform denoising and normalization preprocessing on the acquired image stream to eliminate illumination changes and background interference, thereby improving the stability and consistency of subsequent feature extraction.

[0026] S3: Input the preprocessed image frame into the pre-trained visual feature extraction model to obtain the initial frame-level feature vector, which serves as the basis for subsequent semantic anchor generation and feature alignment.

[0027] S4: Based on the evolution law of weld morphology, a lightweight semantic anchor generator is constructed to identify key time points of molten pool boundary changes and heat-affected zone expansion, and generate reference anchor features with semantic stability.

[0028] S5: Design a differentiable semantic constraint loss function to calculate the cross-frame semantic consistency error between the current frame-level features and the nearest semantic anchor point, so as to drive the dynamic alignment and semantic stabilization of the feature space.

[0029] S6: Based on the frame-level feature alignment results guided by semantic anchors, update the parameters of the feature learning module to improve the robustness of the model's feature representation under changes in material, environment, or process parameters.

[0030] S7: Input the aligned frame-level feature vector into the anomaly detection module to perform welding quality anomaly judgment based on dynamic threshold, so as to identify potential welding defects caused by feature drift.

[0031] S8: Based on the anomaly detection results and the semantic anchor update status, dynamically adjust the weight parameters of the semantic constraint loss function to achieve an adaptive balance between improving recognition accuracy and controlling training stability.

[0032] S9: Periodically evaluate the semantic consistency index of the feature space. If it is determined to be a scenario with significant feature drift, trigger the semantic anchor update mechanism to extend the model life cycle and enhance generalization ability.

[0033] Step S1: Acquire image streams from multiple perspectives during the welding process. These image streams include continuous video frames of the weld pool area, the heat-affected zone, and the surrounding environment to obtain a multimodal visual representation of the welding process. Specifically, this includes: S1.1: Based on the industrial vision sensing system, multiple high frame rate industrial cameras are deployed simultaneously during the welding robot operation. The cameras are respectively arranged in front of the welding torch, above and to the side, and in the backlight area to obtain multi-view continuous video frames of the weld pool area, heat-affected zone, and surrounding environment area, thereby constructing a multi-modal visual input for the welding process.

[0034] Based on the industrial vision sensing system, multiple high frame rate industrial cameras (frame rate ≥ 120fps, resolution ≥ 1920×1080) are deployed during the welding robot operation as the basic acquisition device for multimodal vision input.

[0035] A spatial geometric calibration method (parameter: Zhang Zhengyou calibration method, calibration plate size 30mm×30mm) is adopted to accurately calibrate the intrinsic and extrinsic parameters of each camera, and determine their three-dimensional spatial relative position and attitude parameter matrix to ensure the spatial correspondence of imaging from different viewpoints.

[0036] Furthermore, by using an optical axis alignment optimization algorithm (parameter: minimize projection distortion error threshold ≤ 0.2 pixels), the camera mounting posture of three types of viewpoints—directly in front of the welding torch, above the side, and in the backlight area—is adjusted so that the molten pool area, heat-affected zone, and surrounding environment are within the effective imaging range from all viewpoints, and the effective pixel ratio of the overlapping field of view is maximized.

[0037] By employing a high dynamic range (HDR) exposure control strategy (parameters: short exposure time ≤ 1 / 5000 sec, long exposure time ≤ 1 / 1000 sec, composite intensity weight ratio 0.7:0.3), the imaging capability of the molten pool and heat-affected zone under strong arc light background is enhanced, and image frame streams with rich brightness levels and structural details are obtained.

[0038] Employing hardware-level buffering and a high-speed data bus protocol (Camera Link or 10GigE interface), low-latency image transmission from each industrial camera to the local data acquisition server is achieved, generating a multi-view continuous video frame sequence with high-precision hardware timestamps (accuracy ≤1µs).

[0039] Through the above chain processing, the physical layout of the camera and the results of image quality control are transformed into multimodal visual input data with high spatial alignment accuracy, strong temporal synchronization and full detail expression, so as to realize the comprehensive acquisition of multi-view information of the welding process and the high reliability of subsequent processing.

[0040] For example, on an automated welding production line, three high-speed industrial cameras are deployed in the working area of ​​a 6-axis welding robot. The front camera is mounted coaxially with the welding direction, with its lens at a 0° angle to the weld axis and a distance of 150mm from the weld. The side-above camera is mounted at a 45° angle to the left of the weld, with its lens at a 35° angle to the weld plane and a distance of 200mm from the weld. The backlight camera is positioned 180° directly behind the weld and a distance of 300mm from the weld. Each camera uses a fixed-focus industrial lens with a 25mm focal length, and the intrinsic parameter matrix is ​​obtained using the Zhang Zhengyou calibration method. and distortion coefficient The reprojection error from multiple viewpoints was controlled within 0.15 pixels. An HDR compositing strategy was employed, with short exposure times set to 1 / 6000 second and long exposure times set to 1 / 1500 second, effectively overcoming the saturation problem caused by arc flash during welding. In the data transmission stage, the three cameras transmitted image data to the local server via a 10GigE link, with an average bandwidth usage of 7Gbps, a single-frame transmission latency of less than 0.8ms, and a hardware timestamp accuracy of ±0.5µs. This resulted in a multi-view welding process video frame stream with a synchronization error of no more than 1µs, providing high-quality input for subsequent feature extraction and semantic anchor generation.

[0041] S1.2: Perform synchronous trigger control on multiple deployed industrial cameras, and use a timestamp alignment mechanism to perform frame-level synchronization processing on image frames from different perspectives to eliminate cross-view frame misalignment caused by differences in camera acquisition timing, and obtain time-consistent multi-view image stream data.

[0042] For the video frame sequences acquired by multiple high frame rate industrial cameras from the output of step S1.1, a hardware synchronous trigger control method (parameters: synchronization signal frequency ≥ 10kHz, trigger delay ≤ 100ns) is adopted to realize the image acquisition start control of each camera under a unified time reference.

[0043] Furthermore, a high-precision timestamp alignment mechanism (parameters: timestamp resolution 1µs, synchronization accuracy error ≤ ±0.5µs) is used to bidirectionally calibrate the embedded timestamps of each industrial camera with the system master clock, correcting time deviations caused by clock drift or link delay differences.

[0044] Furthermore, a cross-view frame-level matching algorithm (method: nearest neighbor matching strategy based on timestamp index mapping) is adopted to construct a one-to-one correspondence between the multi-view frame sequences output by each camera and generate a cross-view frame synchronization index table to achieve the unification of frame numbers between different views.

[0045] Furthermore, the inter-frame time difference calculation formula is used: in, For the first Timestamps of frames captured by the camera For the first Timestamps of frames captured by the camera The threshold condition is used to determine the acquisition time difference between corresponding frames. ( To set the upper limit of the synchronization tolerance (1µs), frame pairs that do not meet the synchronization accuracy requirements are removed to ensure that the data participating in subsequent processing has time consistency.

[0046] Furthermore, based on the multi-channel image stream buffer merging strategy of the synchronization index table (parameters: buffer depth = 5 frames, dequeue delay ≤ 2ms), multi-view frames that meet the synchronization conditions are output in parallel according to the timestamp order to generate a time-consistent multi-view welding process image stream.

[0047] By using synchronous trigger control, timestamp alignment, and cross-view frame matching processing, the multi-camera acquisition data from the previous step is transformed into synchronous video frames with a cross-view time error of ≤1µs, providing time-consistent and directly comparable multimodal visual input for ROI region extraction in S1.3.

[0048] For example, in an automated welding production line implementation scenario, three high-speed industrial cameras (model: Baslerboost series, 120fps frame rate, 1920×1080 resolution) are fixed at 0° directly in front of the welding torch, 45° to the side and above, and 180° behind the light, respectively, and connected to a hardware synchronization controller (model: National Instruments PXIe-6674T). The system is set to a synchronization pulse frequency of 10kHz and the GPS timing module (time synchronization accuracy 0.5µs) is activated to provide a unified UTC time reference for the controller and each camera. At the data receiving end, the video streams transmitted via the three 10GigE links are timestamped and compared through an FPGA card. When the calculated time difference between the three frames is... If the time deviation of any frame is less than 0.8µs, the frame group is marked as a valid synchronization frame group. If the time deviation of any frame is detected to be greater than the threshold, the frame is discarded and the previous valid synchronization frame group is output from the buffer queue. After this processing, the system achieved a frame-level synchronization rate of 99.98% within 2 hours of continuous operation, and the cross-view frame error rate was less than 0.02%, providing stable multi-view time series data for subsequent feature extraction and alignment.

[0049] S1.3: Perform ROI region segmentation processing on the synchronized multi-view image stream. Based on weld trajectory planning information and welding process parameters, identify and crop the regions of interest in the weld pool area, heat-affected zone and surrounding environment area to focus on key visual information and reduce redundant computation in subsequent processing.

[0050] S1.4: An FPGA-based image preprocessing unit is used to perform real-time filtering on image frames from various perspectives. A bilateral filtering algorithm is used to suppress high-frequency noise in the image, thereby improving image quality and enhancing the structural clarity of the molten pool boundary and heat-affected zone.

[0051] S1.5: Based on the illumination change characteristics of the welding process, dynamic histogram equalization is performed on the filtered image frame, and an adaptive contrast enhancement algorithm is used to normalize the brightness of the image to eliminate the problem of inconsistent image brightness caused by factors such as welding arc flash and metal reflection.

[0052] Step S2: The acquired image stream undergoes denoising and normalization preprocessing to eliminate illumination variations and background interference, thereby improving the stability and consistency of subsequent feature extraction. Specifically, this includes: S2.1: Perform spatial denoising processing based on Gaussian-Laplace hybrid filtering on the raw welding image stream acquired from multiple perspectives to suppress random noise introduced by the image sensor and obtain spatially smooth weld area image data.

[0053] For the raw welding image stream acquired synchronously from multiple perspectives, a Gaussian-Laplacian hybrid filtering algorithm (parameters: Gaussian kernel standard deviation σ=1.2, Laplacian eight-neighbor template coefficient center value -8, periphery value 1) is adopted to achieve the function of suppressing random noise and preserving detail edges in the image spatial domain.

[0054] Furthermore, by performing a Gaussian smoothing convolution operation (with a kernel size of 5×5 and normalization coefficients satisfying a kernel weight sum of 1), the amplitude of sensor thermal noise and high-frequency particle noise in the welding environment is reduced, and a preliminarily smoothed image matrix is ​​output. .

[0055] Furthermore, on Perform the Laplace operator enhancement operation using the formula: in Using Laplace template coefficients, the second derivative response for the contour details of the molten pool edge and heat-affected zone is calculated, yielding the detail enhancement matrix. .

[0056] Furthermore, the smoothing matrix With detail enhancement matrix Weighted fusion is performed, and the fusion formula is as follows: Among them, the fusion weight , This achieves a balance between noise reduction and detail preservation, and outputs a spatially denoised image matrix. .

[0057] Furthermore, the pixel intensity mean square error (MSE) detection algorithm is used to calculate the noise reduction rate of the images before and after processing. Output noise suppression evaluation metrics to ensure that the noise reduction effect meets the threshold. .

[0058] By using a Gaussian-Laplace hybrid filtering method, the multi-view raw image data from the previous step is transformed into spatially smooth, detailed, and low-noise weld area image data, achieving the expected technical effect of providing high signal-to-noise ratio input for subsequent illumination normalization and background suppression.

[0059] For example, on a fully automated welding production line using three high-speed industrial cameras, Gaussian-Laplacian hybrid filtering is performed on raw multi-view video frames with a resolution of 1920×1080 and a frame rate of 120fps. First, the Gaussian kernel σ=1.2 and the convolution template size is set to 5×5. The convolution sum is then calculated to obtain the smoothed image matrix. The sensor noise MSE_b in the original image is 25.4. Then, an eight-neighbor Laplacian template is applied. right Perform convolution to obtain the detail enhancement matrix. The mean second derivative of the sampled edge pixels is 18.6. and according to Fusion, output denoised image matrix And calculate MSE a =3.2, corresponding to a noise reduction rate The value exceeds the threshold of 0.85, indicating that noise suppression is sufficient and details are well preserved. In welding frames from different viewpoints, the continuity of the molten pool boundary and the heat-affected zone contour is maintained at more than 95%, providing a high-quality, low-noise input basis for the S2.2 illumination non-uniformity correction algorithm.

[0060] S2.2: Based on the spatially denoised weld area image data, an illumination non-uniformity correction algorithm based on Retinex theory is used for illumination normalization processing to eliminate brightness fluctuations caused by arc flicker or changes in ambient light during the welding process, thereby obtaining a weld image sequence with uniform illumination.

[0061] S2.3: Perform background suppression operation based on Otsu threshold segmentation on the weld image sequence after illumination equalization to extract the effective areas of the weld pool and heat-affected zone, and obtain a welding process image dataset with enhanced foreground.

[0062] S2.4: Based on the welding process image dataset with foreground enhancement, an adaptive histogram equalization algorithm is applied to enhance the image contrast, thereby improving the visibility of weld morphology details and the ability to express features, and obtaining welding image frames with optimized contrast.

[0063] S2.5: Perform a linear normalization operation based on the pixel value range on the contrast-optimized welding image frame to map its grayscale value to the [0,1] interval, so as to unify the image input scale and adapt to the input requirements of the subsequent pre-trained visual feature extraction model, and obtain a standardized visual input representation of the welding process.

[0064] Step S3: Input the preprocessed image frame into the pre-trained visual feature extraction model to obtain the initial frame-level feature vector, which serves as the basis for subsequent semantic anchor generation and feature alignment. Figure 2 As shown, it specifically includes: S3.1: Perform feature encoding processing based on the ResNet-18 architecture on the preprocessed multi-view welding image frames to obtain high-dimensional semantic feature maps; wherein, the preprocessed image frames are derived from the normalized video frame data output in step S2, and the ResNet-18 model is a lightweight visual feature extraction network pre-trained on the ImageNet dataset; perform this processing to generate frame-level visual embedding representations with spatial preservation properties.

[0065] S3.2: Perform a global average pooling operation on the high-dimensional semantic feature map output by ResNet-18 to compress the spatial dimension and preserve the channel feature response intensity; wherein, the high-dimensional semantic feature map is the 4D tensor data output by step S3.1, and the global average pooling operation compresses the spatial dimension of each channel into a single scalar value; perform this processing to generate an initial frame-level feature vector with a channel dimension of 512.

[0066] For sub-step S3.2, the input data is the high-dimensional semantic feature map output from step S3.1. This feature map is 4D tensor data from the ResNet-18 network, and the data dimension format is... ,in For batch size, For the number of channels, and Here are the spatial height and width of the feature map. This step aims to compress the two-dimensional spatial features of each channel into a single scalar through a global average pooling operation, thereby preserving the global response strength of that channel and removing spatial location dependencies.

[0067] Global Average Pooling (GAP) method is used (parameter: pooling kernel size). , Step length , This enables global statistical aggregation of two-dimensional features for each channel.

[0068] Furthermore, the input feature map is processed through a GAP operation. The Middle The characteristic values ​​of each channel The formula for calculating the mean is: in, This is the global average response value for this channel. and This refers to the spatial dimensions.

[0069] Furthermore, through parallel matrix operations in the batch processing dimension, The mean calculation is performed sequentially on each channel, generating a dimension of... Feature matrix Each row of this matrix corresponds to a 512-dimensional global feature description of a frame of image.

[0070] Furthermore, 32-bit floating-point precision storage is adopted. To avoid precision loss on small numerical channels, and to optimize memory read / write efficiency through cache alignment, the data is adapted to meet the real-time processing requirements of high frame rate welding video data.

[0071] Furthermore, for the abnormal channel responses caused by extreme changes in brightness, the response variance of each channel is calculated. Furthermore, weight reduction is applied to low-variance channels to improve the stability of subsequent feature normalization.

[0072] By using global average pooling, the high-dimensional semantic feature map from the previous step is transformed into a 512-dimensional initial frame-level feature vector, achieving the expected technical effect of spatial dimension compression and global semantic strength preservation.

[0073] For example, on a production line equipped with a dual-station welding robotic arm, the size of the input high-dimensional semantic feature map is... This corresponds to each batch containing 32 frames of welding image features encoded with ResNet-18. For each frame of data, the following methods are used: The global average pooling core configuration will allocate the power of each channel. Spatial features are averaged and calculated according to the above formula. The response values ​​of each channel. For example, the summation of the feature values ​​of the 100th channel in a certain frame is: The mean was calculated as follows After batch calculation is completed, the result is obtained. The feature matrix of the shape is represented using float32, reducing the data storage and transmission bandwidth requirements per frame to that of the original high-dimensional features. Regarding variance The channels were weighted with a reduction factor of 0.5 to minimize the impact of low-dynamic channels on subsequent normalization. The final output 512-dimensional initial frame-level feature vector was verified to maintain excellent separability and cross-frame stability during the subsequent S3.3 normalization and channel attention weighting processes.

[0074] S3.3: Perform L2 normalization on the initial frame-level feature vectors output by global average pooling to eliminate the impact of feature vector magnitude differences on subsequent semantic alignment; wherein, the initial frame-level feature vectors are derived from the 512-dimensional vectors output in step S3.2, and the L2 normalization is performed by dividing by the Euclidean norm of the vector to achieve normalization; this process is performed to generate normalized feature vectors with consistent magnitudes, thereby improving the comparability and stability of the feature space.

[0075] For step S3.3, the input data is the 512-dimensional initial frame-level feature vector output from step S3.2. This vector comes from the result of global average pooling of the high-dimensional semantic feature map output by ResNet-18. The modulus needs to be normalized to enhance the comparability of features across frames and across viewpoints and the stability of feature alignment.

[0076] The L2 normalization algorithm (parameters: precision floating-point type float32, normalization threshold ε=1e-12) is used to normalize the Euclidean norm of the feature vector to 1.

[0077] Furthermore, by calculating the input feature vector The Euclidean norm V2 is obtained by applying the following formula: in, For the first 3D channel response value, is the magnitude of the eigenvector.

[0078] Furthermore, normalization is performed on each channel component: in, These are the normalized component values. To prevent division by zero of small constants.

[0079] Furthermore, batch matrix operations are used to process the feature matrices along the batch dimension. The above operations are performed and parallelized using the SIMD instruction set to improve the real-time processing performance of high frame rate video streams.

[0080] Furthermore, the modulus and standard deviation are calculated from the normalized feature matrix. And compare thresholds ,when If the output is below this threshold, it will be marked as having acceptable modulus consistency to ensure that the feature scale does not introduce additional bias in subsequent semantic alignment.

[0081] By using L2 normalization, the 512-dimensional initial frame-level feature vector from the previous step is transformed into a normalized feature vector with consistent modulus length. This achieves unified feature scale, eliminates the interference of modulus length differences on semantic anchor similarity calculation, and improves the stability of the feature space in dynamic welding scenarios.

[0082] S3.4: Perform channel attention mechanism weighting on the normalized feature vector to enhance the feature channel response related to the welding state; wherein, the normalized feature vector is derived from the unitized vector output in step S3.3, and the channel attention mechanism is constructed based on the SENet structure, and adaptive weight allocation is performed on each channel through compression-excitation operation; this processing is performed to generate channel-weighted enhanced frame-level feature vectors, thereby improving the ability to represent key welding states.

[0083] For sub-step S3.4, the input data is the normalized frame-level feature vector with consistent magnitude output from step S3.3, and the dimension of this vector is... It is used to perform channel-level response selectivity enhancement within the feature space to highlight information channels related to the welding state and suppress interference from background and irrelevant feature channels.

[0084] Employing a channel attention mechanism based on the SENet architecture (parameter: compression ratio) The activation function is a combination of ReLU and Sigmoid, which enables adaptive calculation of channel weights for normalized feature vectors.

[0085] Furthermore, by using SENet's "Squeeze" operation, global information compression is performed, normalizing the input vector. Send into the fully connected layer This enables information fusion between channels and compresses its dimensions from 512 to [a smaller number]. This operation can be represented as: in, These are low-dimensional compressed feature vectors.

[0086] Furthermore, the ReLU activation function is used to... Perform a nonlinear transformation to introduce enhanced nonlinear correlations between channel features, resulting in... And send it to the fully connected layer. Perform dimensional restoration, mapping it back to the original channel number 512: Furthermore, by using the Sigmoid function... Perform a normalization mapping to confine the weights of each channel to the (0,1) interval, resulting in the channel attention weight vector. ,in: for The Dimensional components.

[0087] Furthermore, the channel attention weight vector With input normalized feature vector Perform element-wise multiplication to generate channel-weighted enhanced frame-level feature vectors. : in, For the weighted number of 3D channel characteristics.

[0088] By using a channel attention weighting processing method based on SENet, the normalized feature vector from the previous step is transformed into an enhanced frame-level feature vector with higher response on the feature channels of welding critical states, thereby achieving the expected technical effect of selective amplification and noise suppression of welding semantic features across frames and viewpoints.

[0089] For example, on a ship section welding production line using laser-MIG hybrid welding, the batch size... The shape of the input normalized feature matrix is Compression ratio For the data vector of the k-th frame ,through Mapping to 32 dimensions yields The maximum component value is 0.312, which remains non-negative after ReLU, and then... Mapping back to 512 dimensions yields The component range is [-0.85, 0.96]. After transformation using the Sigmoid function, we obtain... The mean was 0.532 and the standard deviation was 0.187, indicating a significant weighting distinction between channels. and After element-wise multiplication, the response coefficients of the 51 channels increased by more than 20%, mainly concentrated in feature channels related to the molten pool brightness gradient, texture directionality, and heat-affected zone edge contrast. During the welding process, the weights of low-relevance channels (such as background noise features) were suppressed to below 0.1, effectively filtering dynamic interferences such as welding fumes and spatter. In the subsequent S3.5 temporal consistency projection, this enhanced feature vector improved the cross-frame cosine similarity index by 12.6% compared to the unweighted vector, significantly improving the stability of semantic alignment in dynamic welding scenarios.

[0090] S3.5: Perform a temporally consistent projection transformation on the channel-weighted enhanced frame-level feature vector to align the feature space distribution between adjacent frames; wherein, the enhanced frame-level feature vector is derived from the weighted vector output in step S3.4, and the temporally consistent projection transformation uses a linear transformation matrix and is dynamically updated based on the feature mean of adjacent frames within the sliding window; this process is performed to generate a temporally aligned initial frame-level feature vector, which serves as the input basis for the subsequent semantic anchor generation and feature alignment modules.

[0091] For sub-step S3.5, the input data is the channel-weighted enhanced frame-level feature vector output from step S3.4, and the dimension of this vector is... It is necessary to align the feature space distribution of adjacent frames in the time dimension to reduce feature drift caused by changes in dynamic welding conditions and enhance cross-frame semantic consistency.

[0092] Employing a time-consistent projection transformation method (parameter: linear transformation matrix) Dimensions Sliding window length Frame, update step (Frame), to realize the function of aligning the distribution of features of adjacent frames in a unified reference space.

[0093] Furthermore, a sliding window mechanism is used in the current frame. and its adjacent Within a frame, calculate the enhanced feature vector for each frame. The element-wise mean is used to obtain the mean vector for that time window. : in, For the first in the window Enhanced feature vectors of frames.

[0094] Furthermore, a centered feature matrix is ​​constructed based on the mean vector. Subtract features from each frame. The operation is performed to obtain a sequence of relative change vectors within a time window, which is then used to estimate the subsequent linear projection matrix.

[0095] Step S4: Based on the evolution law of weld morphology, a lightweight semantic anchor generator is constructed to identify key time points of weld pool boundary changes and heat-affected zone expansion, generating semantically stable reference anchor features. For example... Figure 3 As shown, it specifically includes: S4.1: Local region features are extracted from the preprocessed multi-view image frame sequence. A sliding window mechanism is used to enhance the features of the weld pool area, heat-affected zone and surrounding environment frame by frame to obtain a local semantic feature map with spatial positioning capability.

[0096] For this sub-step, the input data are the time-aligned initial frame-level feature vector sequence output from step S3.5 and the preprocessed multi-view image frame sequence output from step S2. This data contains multimodal visual information of the weld pool area, heat-affected zone and surrounding environment area.

[0097] A local region feature extraction method is used (parameter: window size w). h =64 pixels, w w =64 pixels, step size s h =16 pixels, s w =16 pixels), to realize the function of spatial block sampling and regional feature encoding of each frame of multi-view image.

[0098] Furthermore, through a sliding window mechanism (parameter: time window length T) f =5 frames, stride Δt=1 frame), to achieve local spatiotemporal data extraction across time series, and to feed the spatial sub-block feature sequence within each local window as input into a lightweight convolutional encoder for local feature enhancement, resulting in a set of local semantic features with temporal context.

[0099] Furthermore, a combined weighted algorithm of channel attention and spatial attention is adopted (parameter: channel compression ratio r). c =8, spatial convolution kernel size k s =3), to achieve multi-scale weighted enhancement of local sub-block features, so as to amplify the local response channels that are highly correlated with weld morphology changes and suppress low-weight channels that are correlated with background noise.

[0100] Furthermore, through a feature map construction method, all local enhancement features are mapped according to their spatial location in the original image, generating a local semantic feature map with spatial localization accuracy up to sub-block resolution. And retain the characteristic change trajectory information across time windows in the graph.

[0101] By using a sliding window-based local feature enhancement processing method, the time-aligned frame-level feature vectors are combined with the spatiotemporal change patterns of key regions in multi-view images, and transformed into a local semantic feature map with spatial positioning capabilities. This enables precise characterization of the fine-grained evolution of the weld pool and heat-affected zone, and provides preliminary support for anchor point generation.

[0102] For example, on an aluminum alloy high-frequency pulse welding production line employing three-view monitoring, the resolution of the input multi-view weld seam image frame is... Pixels, sliding window space size set to The number of pixels is calculated with a step size of 16 pixels, a time window length of 5 frames, and a step size of 1 frame. For a given moment in the forward-looking camera frame sequence, a total of [number] pixels were extracted. Each sub-block contains features within a 5-frame time window. Features from each sub-block are input into a lightweight three-layer convolutional encoder with a kernel size of [size missing]. The output has 64 channels, which are then weighted by channel attention (compression ratio 8) to form enhanced features. All sub-block features are synthesized according to their original image positions. The feature map, with a resolution of 76×42 and 64 channels, records the feature change curves over time. Tests show that this feature map improves the response coefficient by an average of 18% at the edge of the molten pool and in the extended heat-affected zone, while reducing the background noise response to more than 0.27 times the original level. This provides a high signal-to-noise ratio regional feature input for the subsequent S4.2 morphology change rate calculation.

[0103] S4.2: Based on the evolution law of weld morphology, the change rate of the molten pool boundary and the expansion trend of the heat-affected zone in the local semantic feature map are dynamically modeled. The first-order difference method is used to calculate the feature change gradient to identify the key time points of morphological change.

[0104] For this sub-step, the input data is the local semantic feature map output from step S4.1. The map contains spatial location features across time windows and local response information of the weld pool boundary and heat-affected zone.

[0105] A method based on weld morphology time-series modeling is adopted (parameter: time window length). The frame (spatial resolution corresponding to the index position of the sub-block in the original image) enables a numerical characterization of the rate of change of the molten pool boundary and the expansion trend of the heat-affected zone.

[0106] Furthermore, using the first-order difference method (parameter: difference interval Δt = 1 frame), the feature response values ​​of the same spatial location in adjacent frames within the time window are obtained. Perform differential calculations to obtain the instantaneous gradient in the time dimension. : in, For the spatial sub-block index coordinates, For time frame index.

[0107] Furthermore, by analyzing each spatial location within the time window Take the average to obtain the average rate of change at that location. This serves as a quantitative indicator of the dynamic sensitivity of the local topography: Furthermore, the spatial location set corresponding to the weld boundary and heat-affected zone Perform regional statistics and calculate the boundary change rate separately. With the extended trend rate : in, For the set of boundary region indices, For the set of thermally affected zone indexes, and These represent the number of sub-blocks in the corresponding region.

[0108] Furthermore, based on a preset mutation determination threshold and ,right and Perform dual threshold detection when or At that time, the current time frame is marked as a candidate point for morphological change.

[0109] By using a processing method based on the evolution law of weld morphology and first-order difference modeling, the local semantic feature map is transformed into quantified boundary change rate and expansion trend index, realizing the automatic identification of key time points of morphological change in dynamic welding process, and providing a high-precision candidate event set for subsequent S4.3 time sequence consistency verification.

[0110] For example, on a stainless steel all-position pipe welding production line monitored by dual-view cameras, the spatial resolution of the input local semantic feature map is... Time window length Frame, differential interval Δt = 1 frame. Actual set threshold. , For a given welded section, the boundary area Contains 380 sub-blocks, heat-affected zone Contains 560 sub-blocks, calculated within this time window. mean ,get , In this case only Exceeding the mutation threshold According to the judgment rules, this time frame was marked as a candidate point for abrupt change in the weld pool boundary. In actual comparison, this frame precisely corresponds to the instant when the weld pool undergoes a sharp morphological change after process parameter adjustment. The slope of the feature change curve at the abrupt change point is 2.3 times higher than that in the stable phase, verifying the high sensitivity and low false alarm rate of this step. The identification result was successfully transmitted to S4.3 for temporal consistency verification, ensuring that the generated semantic anchor points have morphological representativeness and stability. S4.3: Perform temporal consistency verification on the identified morphological change time points, and combine the feature similarity calculation results between adjacent frames to select time points with semantic stability higher than the preset threshold as a candidate semantic anchor point set.

[0111] For sub-step S4.3, the input data is the set of candidate time points for morphological abrupt changes output by step S4.2 and the sequence of time-aligned frame-level feature vectors output by step S3.5. This data contains the feature space distribution features of the frame corresponding to each candidate time point and its adjacent frames.

[0112] Employ a timing consistency verification method (parameter: verification window length) The frame (with cosine similarity as the similarity metric) is used to verify the cross-frame semantic stability of candidate points for morphological abrupt changes.

[0113] Furthermore, each candidate time point is processed through a sliding time window. Before and after it Intra-frame feature vector Perform frame-by-frame similarity calculation.

[0114] Furthermore, by analyzing each candidate point Adjacent frame pairs within the corresponding window , Wait for... Calculate and take the average value to obtain the temporal consistency score of the candidate point. : in, The summation value covers all adjacent frame pairs within the verification window, representing the similarity score.

[0115] Furthermore, based on a preset semantic stability threshold (e.g., 0.85), filter out all Candidate time points are marked as semantically stable qualified nodes.

[0116] Furthermore, the set of semantically stable qualified nodes is organized into a candidate semantic anchor set in chronological order. Each node is assigned a similarity score and a corresponding morphological change rate index for subsequent anchor point feature aggregation.

[0117] By using a processing method based on temporal consistency verification and feature similarity screening, the morphological change candidate points identified in the previous step are transformed into a set of anchor point candidate points with high semantic stability, thereby achieving reliable filtering of anchor point generation and ensuring cross-frame semantic continuity in dynamic welding scenarios.

[0118] For example, on a high-strength steel frame robotic arc welding production line monitored by dual cameras, the number of input morphological change candidate points is 12, and the frame-level feature dimension is 512. Set to 3 frames, similarity threshold Regarding candidate time points The extracted 3-frame window is , , Calculated , Therefore, Points exceeding the threshold are identified as stable nodes. This calculation is performed on all candidate points, ultimately retaining 9 stable nodes, accounting for 75%. These nodes all appear in the stable molten pool state period after welding speed adjustment or sudden weld width change, providing a high-confidence time index for subsequent S4.4 anchor point feature aggregation, verifying the accuracy of anchor point selection under significant feature drift conditions.

[0119] S4.4: Based on the candidate semantic anchor set, construct the anchor feature representation, and use the max pooling and attention weighted fusion strategy to aggregate the frame-level features within the anchor time window to generate a representative and stable semantic anchor feature vector.

[0120] S4.5: Cache the generated semantic anchor feature vectors into the anchor feature pool and maintain the anchor update queue based on timestamp information to support dynamic anchor retrieval and semantic consistency constraint calculation in the subsequent feature alignment process.

[0121] Step S5: Design a differentiable semantic constraint loss function to calculate the cross-frame semantic consistency error between the current frame-level features and the nearest semantic anchor point, thereby driving dynamic alignment and semantic stabilization of the feature space. Specifically, this includes: S5.1: Based on the frame-level feature vector of the welding process and the feature vector of the semantic anchor point, construct a cross-frame semantic similarity matrix to quantify the semantic association strength between the current frame and the anchor point.

[0122] S5.2: The cosine similarity metric is used to calculate the similarity between the current frame feature vector and the feature vector of the nearest semantic anchor point to obtain a cross-frame semantic consistency metric.

[0123] For this sub-step, the input data includes two core inputs required for the cross-frame semantic similarity matrix output from step S5.1: the temporally aligned feature vector of the current frame. And the semantic anchor feature vector in the anchor feature pool that is closest to the current timestamp. Both are 512-dimensional vectors that have undergone L2 normalization and channel attention enhancement.

[0124] A semantic consistency measurement method based on cosine similarity is adopted (parameter: feature vector dimension). ), to achieve and Cross-frame semantic similarity calculation.

[0125] Furthermore, to enhance sensitivity to subtle semantic differences, the calculated... Perform a linear mapping transformation on the value, mapping it to the interval [0,1]: in, This is a normalized semantic consistency metric. The closer the value is to 1, the more semantically consistent the current frame is with the anchor point in the feature space.

[0126] Furthermore, for the current frame of The value is its time neighboring frame , of Values ​​are smoothed using a moving average (parameter: window length). (frames), reducing allergic reactions caused by transient abnormal fluctuations: Obtain a smooth cross-frame semantic consistency metric. This serves as the direct input for constructing semantic constraint terms in subsequent S5.3.

[0127] By using a multi-level calculation method based on cosine similarity combined with normalization and time smoothing, the semantic alignment degree between the current frame and the nearest anchor point is transformed into a stable and differentiable cross-frame semantic consistency index, thereby achieving accurate quantification of cross-view feature space semantic matching in dynamic welding scenarios.

[0128] For example, on a high-strength steel laser welding production line with three-view monitoring, the current frame feature vector and anchor point feature vector input in step S5.1 are both 512-dimensional and L2 normalized. The cosine similarity is calculated as follows: Through linear mapping, we obtain In the sliding window At frame rate, the normalized similarity between two consecutive frames is 0.925 and 0.941, respectively, and the smoothed calculation result is... This value, exceeding the preset high confidence threshold of 0.92 for semantic consistency on the production line, can be used as a positive sample alignment signal in the calculation of contrastive learning loss when entering S5.3. In tests under different operating conditions, the misclassification rate of unstable weld sections using this smoothed cosine similarity method decreased from 12.3% in the non-smooth scheme to 6.7%, verifying the effectiveness of this calculation strategy in anti-interference and feature drift suppression.

[0129] S5.3: Introduce a differentiable contrastive learning loss function and construct semantic constraint terms based on cross-frame semantic consistency metrics to enhance the semantic alignment capability of the feature space.

[0130] For sub-step S5.3, the input data is the smooth cross-frame semantic consistency metric output from step S5.2. In addition, the predefined set of positive and negative sample pair indices in the contrastive learning framework, positive sample pairs represent the current frame and anchor points with high semantic consistency, and negative sample pairs represent the current frame and non-matching anchor point combinations.

[0131] A loss function design method based on contrastive learning is adopted (parameter: temperature coefficient). This enables a differentiable constraint expression of the similarity distribution of positive and negative sample pairs in the feature space.

[0132] Furthermore, the InfoNCE loss formula is applied to each positive sample pair. The similarity is calculated using a normalized ratio: in, The total number of samples within the batch. To exclude the indicator function of the current sample itself.

[0133] Furthermore, the average loss of all positive sample pairs within a batch is taken to obtain the batch-comparison loss. : in, For the set of positive sample pairs, It is the number of its elements.

[0134] Furthermore, semantic constraint weighting coefficients are introduced. and will and Deviation from the benchmark threshold Combining penalty terms to construct semantic alignment constraint terms : in, This is the penalty weight parameter, used to increase the penalty intensity when the similarity is below a threshold.

[0135] Furthermore, the backpropagation mechanism ensures that each part of the loss function contributes gradients to the feature extraction network parameters, thereby simultaneously optimizing cross-frame alignment accuracy and semantic consistency robustness during model updates.

[0136] By constructing a loss function based on differentiable contrastive learning, the smooth semantic consistency metric is transformed into an optimized signal that indicates high similarity for positive samples and low similarity for negative samples, thereby enhancing the adaptive semantic alignment capability of the feature space in dynamic welding scenarios.

[0137] For example, on a stainless steel sheet TIG welding production line with dual-view monitoring, the batch size is set to... Frame, number of positive sample pairs Temperature coefficient Semantic constraint weighting coefficients Penalty weight Similarity benchmark threshold For a specific batch, the serial number Positive sample pairs, their smooth similarity The numerator term is calculated. The sum of the denominator terms is Then the pair Intra-batch average The similarity is 0.947. Since the similarity of this pair is higher than the threshold of 0.9, there is no penalty term contributing. Experiments show that, under this configuration, the cross-frame semantic consistency of the feature space remains above 0.92 after multiple iterations, compared to the configuration without the introduction of [specific feature space parameters]. Through comparative learning, the accuracy of anomaly detection improved by 7.8%, and the false alarm rate decreased to 0.65 times the original level.

[0138] S5.4: Weighted fusion of semantic constraints and the overall training loss function is performed to generate a semantically enhanced joint loss function to drive parameter updates in the feature learning module.

[0139] S5.5: Based on the semantically enhanced joint loss function, the backpropagation optimization algorithm is executed to update the weight parameters of the feature learning module with gradients, so as to improve the semantic consistency and feature alignment stability of the model in the dynamic welding scenario.

[0140] Step S6: Based on the frame-level feature alignment results guided by semantic anchors, update the parameters of the feature learning module to improve the robustness of the model's feature representation under changes in material, environmental, or process parameters. Specifically, this includes: S6.1: Perform gradient backpropagation calculation on the alignment error vector between the semantic anchor and the current frame to obtain the parameter gradient update amount of each network layer in the feature learning module.

[0141] S6.2: Based on the parameter gradient update amount, an adaptive optimization algorithm is used to update the parameters of the convolutional neural network in the feature learning module, so as to reduce the semantic deviation across the feature space and improve the consistency of feature representation.

[0142] The gradient update values ​​of each network layer parameter of the feature learning module obtained from the S6.1 sub-step An adaptive optimization algorithm is used (parameter: base learning rate). First-order moment estimation of attenuation coefficient Second-order moment estimation of attenuation coefficient Numerical stability constant This enables dynamic adjustment of the weight parameters of the convolutional neural network to reduce semantic bias across the feature space of frames.

[0143] Furthermore, by optimizing update rules through Adam, batch processing can be improved. Each layer of weight Cumulative first moment estimation With second-order moment estimation : in, The gradient is squared element by element to ensure that gradient variance information is captured.

[0144] Furthermore, regarding and Perform deviation correction to obtain and : The bias correction step compensates for the underestimation of the first and second moment estimates during the initial iteration.

[0145] S6.3: Apply the updated convolutional neural network parameters to the frame-level feature extraction process in the next time step to generate an optimized high-dimensional feature representation, which serves as the input basis for the subsequent anomaly detection module.

[0146] S6.4: Based on the record of process parameter changes during the welding process, the update range of the feature learning module is dynamically scaled to achieve an adaptive balance between feature drift compensation and model oscillation suppression.

[0147] S6.5: Periodically evaluate the similarity distribution between the updated feature space and semantic anchors. If it is determined that the semantic consistency has been significantly improved, save the current parameter snapshot to the model cache for subsequent model rollback or incremental training.

[0148] Step S7: The aligned frame-level feature vector is input into the anomaly detection module to perform welding quality anomaly judgment based on dynamic thresholds, in order to identify potential welding defects caused by feature drift. Specifically, this includes: S7.1: Perform temporal splicing on the aligned frame-level feature vectors to construct a fragment-level feature representation of the welding process, thereby enhancing the ability to capture the trend of welding quality changes.

[0149] S7.2: Based on the historical feature distribution of the welding process, the sliding window statistical method is used to calculate the Mahalanobis distance between the current segment features and the historical normal state in order to quantify the degree of deviation.

[0150] S7.3: Dynamically adjust the anomaly judgment threshold of the Marvin distance according to the current welding process parameters and material condition to adapt to the feature drift tolerance range under different working conditions.

[0151] The Mahalanobis distance between the current segment features and the historical normal state features output from step S7.2. An adaptive dynamic threshold adjustment method is adopted (parameter: initial threshold). Process parameter weighting coefficient Material state weighting coefficient This enables adaptive adjustment of the abnormal judgment threshold based on operating conditions.

[0152] By setting process parameter vectors (including welding current) Welding voltage Welding speed (etc.) and material state vector (including the grade of the base material) ,thickness Surface oxide film state The standardized representation of (etc.) is used to calculate the deviation of the current operating condition using the weighted Euclidean distance method. : in, and These are the first under normal historical conditions. Item process parameters and the first The average value of the material state.

[0153] Furthermore, a linear mapping algorithm is used to determine the deviation of the operating conditions. Convert to threshold correction factor : in, The proportional coefficient of the deviation to threshold correction factor is determined through optimization using historical validation sets.

[0154] Furthermore, combined with the initial threshold With correction factor Calculate the current dynamic threshold : This formula enables adaptive adjustment of the threshold, which is positively correlated with the degree of deviation from the welding conditions, so that the threshold approaches a certain value during the stable period of the welding conditions. During periods of significant changes in operating conditions, the threshold should be appropriately relaxed to reduce misjudgments caused by normal fluctuations.

[0155] Furthermore, for the dynamic threshold sequence Perform sliding window mid-value filtering (window length) This is to suppress the spike effects of sensing noise and transient disturbances on the threshold and ensure the stability of the decision boundary.

[0156] By using the above dynamic adjustment algorithm, the anomaly judgment threshold of Mahalanobis distance is linked with real-time welding condition parameters and material state, thereby realizing adaptive management of the feature drift tolerance range in a variable production environment.

[0157] For example, on a production line for stainless steel pressure vessel cylinders welded using dual-pulse MIG welding, the initial threshold... Set to 3.5, weighting coefficient for process parameters. The material condition weighting coefficient is set to 0.6. Take 0.4, the proportionality coefficient Set to 0.15. For a certain batch of welding tasks, real-time monitoring showed that the welding current deviated from the historical average by 8A, the welding speed deviated by 0.12 m / min, and the base material thickness was consistent but the surface oxide film state changed by 0.3 (normalized value). Substituting these values ​​into the formula, the deviation modulus of the process parameters was calculated to be... Material condition deviation value ,thus Correction factor Dynamic threshold The threshold becomes smoother after being filtered through a 5-frame window, effectively avoiding false alarms caused by local operating condition disturbances and ensuring the robustness of subsequent S7.4 judgments.

[0158] S7.4: Compare the calculated deviation with the dynamic threshold. If it exceeds the set threshold, generate a welding quality abnormality signal to indicate that there is a potential defect in the current welding process.

[0159] S7.5: Based on the persistence and intensity of abnormal signals, a multi-frame consistency verification mechanism is implemented to reduce the false alarm rate caused by occasional noise or transient disturbances and improve the reliability of diagnostic results.

[0160] S7.6: The final judgment result is encapsulated into a structured diagnostic report, including the anomaly type, occurrence timestamp, and confidence index, for use by the subsequent feedback optimization module.

[0161] Step S8: Based on the anomaly detection result and the semantic anchor update status, dynamically adjust the weight parameters of the semantic constraint loss function to achieve an adaptive balance between improving recognition accuracy and controlling training stability. Specifically, this includes: S8.1: Perform sliding window statistical processing on the current frame anomaly score and historical score sequence output by the anomaly detection module to obtain anomaly change trend indicators.

[0162] S8.2: Based on the anchor update frequency and anchor confidence information output by the semantic anchor update module, a semantic stability evaluation factor is generated as a reference for training stability.

[0163] S8.3: Normalize and weight the abnormal change trend indicators and semantic stability assessment factors to generate a comprehensive dynamic adjustment signal to drive the weight adjustment of the semantic constraint loss function.

[0164] Abnormal trend indicator sequence based on the output of step S8.1 Semantic stability evaluation factor output from step S8.2 As input data, the min-max normalization method is used (parameter: lower bound of normalization). Upper limit of normalization This achieves scale-uniform processing for the two types of input data. Furthermore, the normalized values ​​for the two types of data are calculated using the following formulas: in, These are the normalized outlier values. This is the normalized semantic stability value.

[0165] Furthermore, through the weighted fusion formula (parameter: anomaly weight coefficient) Semantic weight coefficient ,and A linear combination of two types of normalized data is achieved to obtain a comprehensive dynamic adjustment signal. : in, and The initial value is set based on the statistical analysis results of the historical model convergence speed and anomaly detection accuracy.

[0166] Furthermore, a sliding window mean smoothing method is employed (window length...). ), for comprehensive dynamic adjustment signals Smoothing is performed to suppress the interference of transient fluctuations on subsequent loss weight updates and to obtain a smoothed output. : in, This is the index for the current time step.

[0167] Furthermore, for the smoothed Standard deviation normalization is performed to ensure that the adjustment signal has a consistent dynamic range across different batches of welding tasks, thereby improving the usability of cross-task migration and ultimately generating a normalized smooth adjustment signal. .

[0168] By using the above-mentioned normalization-weighted fusion-smoothing-standardization algorithm, the abnormal change trend and semantic stability evaluation results are transformed into a unified dynamic adjustment signal that can directly drive the weight update of the semantic constraint loss function, thus achieving consistency and adaptability of the weight adjustment strategy under multiple tasks and multiple working conditions.

[0169] For example, on an automated welding production line, abnormal trend indicators were collected. The semantic stability evaluation factor ranges from [0.15, 0.42] over the most recent 50 frames. The range is [0.68, 0.92], which is obtained after normalization. , Set the weighting coefficient for abnormal indicators. Semantic stability weight coefficient Calculate the fused signal Select the smooth window length. Calculate the smoothing result Approximately 0.56192, after standard deviation normalization, we get (Assuming the baseline standard deviation for batch tasks is 8.02). In the... After inputting the weight ratio adjustment module in step S8.4, the system automatically reduced the weight of the semantic consistency term by 2.3% and increased the weight of the local difference retention term by 2.3%, effectively ensuring that the feature learning module converged smoothly and the recognition accuracy improved by 0.9% under the condition that the current abnormal trend was significantly enhanced but the stability of the semantic anchor point was still acceptable.

[0170] S8.4: Based on the comprehensive dynamic adjustment signal, the weight ratio of the semantic consistency term and the local difference preservation term in the semantic constraint loss function is adjusted to achieve an adaptive trade-off between improving recognition accuracy and controlling training stability.

[0171] S8.5: Apply the adjusted semantic constraint loss function to the gradient update process of the feature learning module to obtain the optimal feature alignment parameter configuration in the current welding state.

[0172] Step S9: Periodically evaluate the semantic consistency index of the feature space. If a scenario with significant feature drift is identified, trigger the semantic anchor update mechanism to extend the model's lifecycle and enhance its generalization ability. Specifically, this includes: S9.1: Perform sliding window statistics on the cross-frame semantic consistency error between semantic anchor features and current frame features to obtain a periodic semantic consistency index sequence, which is used to quantify the stability change trend of the feature space.

[0173] S9.2: Based on the semantic consistency index sequence within the sliding window, a moving average filtering algorithm is used for trend smoothing to extract the potential evolution law of feature drift and generate a drift intensity evaluation factor.

[0174] S9.3: Input the drift intensity evaluation factor into the preset dynamic threshold determination module, and perform adaptive threshold calculation based on the historical drift intensity distribution characteristics to determine whether the current scenario with significant feature drift has been entered.

[0175] S9.4: When a scene is determined to be characterized by significant feature drift, the semantic anchor update triggering mechanism is activated to generate a new set of semantic anchor candidates based on the feature vector of the current frame, so as to restore the semantic alignment capability of the feature space.

[0176] S9.5: Perform semantic stability evaluation based on the evolution law of weld morphology on the candidate set of semantic anchors, select representative new semantic anchor features, and complete the incremental update and version switching operation of the semantic anchor library.

[0177] For those skilled in the art, various other corresponding changes and modifications can be made based on the technical solutions and concepts described above, and all such changes and modifications should fall within the protection scope of the claims of this invention.

[0178] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms “first,” “second,” “third,” and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “comprising” or “including” and similar terms mean that the elements or objects preceding “comprising” or “including” encompass the elements or objects listed following “comprising” or “including” and their equivalents, and do not exclude other elements or objects. The “multiple” mentioned in the embodiments of this application refers to two or more. A and / or B indicate three possibilities: A; B; and A and B.

[0179] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A self-monitoring diagnostic method for welding quality abnormalities, specifically including: S1: Acquire image streams from multiple perspectives during the welding process to obtain a multimodal visual representation of the welding process; S2: Preprocess the acquired image stream; S3: Input the preprocessed image frame into the pre-trained visual feature extraction model to obtain the initial frame-level feature vector, which serves as the basis for subsequent semantic anchor generation and feature alignment. S4: Based on the evolution law of weld morphology, a lightweight semantic anchor generator is constructed to identify key time points of molten pool boundary changes and heat-affected zone expansion, and generate reference anchor features with semantic stability. S5: Design a differentiable semantic constraint loss function to calculate the cross-frame semantic consistency error between the current frame-level features and the nearest semantic anchor point; S6: Update the parameters of the feature learning module based on the frame-level feature alignment results guided by semantic anchors; S7: Input the aligned frame-level feature vector into the anomaly detection module to perform welding quality anomaly judgment based on dynamic threshold.

2. The self-monitoring diagnostic method for welding quality abnormalities according to claim 1, characterized in that, Step S7 is followed by: S8: Dynamically adjust the weight parameters of the semantic constraint loss function based on the anomaly detection result and the semantic anchor update status; S9: Periodically evaluate the semantic consistency index of the feature space. If it is determined to be a scenario with significant feature drift, trigger the semantic anchor update mechanism.

3. The self-monitoring diagnostic method for welding quality abnormalities according to claim 1, characterized in that, The image stream in step S1 includes consecutive video frames of the weld pool area, the heat-affected zone, and the surrounding environment area.

4. The self-monitoring diagnostic method for welding quality abnormalities according to claim 1, characterized in that, Step S2 involves preprocessing the acquired image stream, including denoising and normalization, and eliminating illumination variations and background interference.

5. The self-monitoring diagnostic method for welding quality abnormalities according to claim 4, characterized in that, Spatial denoising processing based on Gaussian-Laplace hybrid filtering is performed on the raw welding image stream acquired from multiple perspectives to suppress random noise introduced by the image sensor and obtain spatially smooth weld area image data.

6. The self-monitoring diagnostic method for welding quality abnormalities according to claim 4, characterized in that, Based on the illumination variation characteristics of the welding process, dynamic histogram equalization is performed on the filtered image frames, and an adaptive contrast enhancement algorithm is used to perform normalized brightness correction on the images.

7. The self-monitoring diagnostic method for welding quality abnormalities according to claim 1, characterized in that, In step S1, multiple industrial cameras are used for synchronous trigger control, and a timestamp alignment mechanism is used to perform frame-level synchronization processing on image frames from different perspectives to obtain time-consistent multi-view image stream data.

8. The self-monitoring diagnostic method for welding quality abnormalities according to claim 5, characterized in that, The synchronized multi-view image stream is divided into ROI regions. Based on weld trajectory planning information and welding process parameters, the regions of interest (ROIs) of the weld pool, heat-affected zone, and surrounding environment are identified and cropped.

9. The self-monitoring diagnostic method for welding quality abnormalities according to claim 1, characterized in that, Step S4 specifically includes: Local region features are extracted from the preprocessed multi-view image frame sequence. A sliding window mechanism is used to enhance the features of the weld pool area, heat-affected zone and surrounding environment frame by frame to obtain a local semantic feature map with spatial positioning capability. Based on the evolution law of weld morphology, dynamic modeling is performed on the change rate of the molten pool boundary and the expansion trend of the heat-affected zone in the local semantic feature map to identify the key time points of morphological abrupt change. The temporal consistency of the identified morphological change time points is checked, and the time points with semantic stability higher than the preset threshold are selected as the candidate semantic anchor point set by combining the feature similarity calculation results between adjacent frames. Anchor feature representations are constructed based on a candidate semantic anchor set. A max pooling and attention-weighted fusion strategy is used to aggregate frame-level features within the anchor time window to generate representative and stable semantic anchor feature vectors. The generated semantic anchor feature vectors are cached in the anchor feature pool, and the anchor update queue is maintained based on the timestamp information.

10. The self-monitoring diagnostic method for welding quality abnormalities according to claim 1, characterized in that, Step S5 specifically includes: Based on the frame-level feature vector of the welding process and the feature vector of the semantic anchor point, a cross-frame semantic similarity matrix is ​​constructed to quantify the semantic association strength between the current frame and the anchor point. The cosine similarity metric is used to calculate the similarity between the current frame feature vector and the feature vector of the nearest semantic anchor point in order to obtain a cross-frame semantic consistency metric. A differentiable contrastive learning loss function is introduced, and semantic constraint terms are constructed based on cross-frame semantic consistency metrics. The semantic constraint term and the overall training loss function are weighted and fused to generate a semantically enhanced joint loss function; The backpropagation optimization algorithm is executed based on the semantically enhanced joint loss function to update the weight parameters of the feature learning module by gradient, so as to improve the semantic consistency and feature alignment stability of the model in dynamic welding scenarios.

Citation Information

Cited By

  • Welding defect detection method based on multi-modal cross attention fusion module

    CN121999300A