ResNet-CGAN-based multi-layer multi-weld-seam feature identification method and system

Through the weld feature recognition method that combines ResNet and CGAN, combined with color attention and center of gravity loss function, the accuracy and real-time problem of weld information extraction in multi-layer multi-pass welding is solved, and high-precision and efficient welding quality improvement is achieved.

CN120563403APending Publication Date: 2025-08-29CHINA NUCLEAR IND HUAXING CONSTR

Patent Information

Application Number
CN202510501034.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Traditional machine vision methods are difficult to stably extract high-precision weld information in multi-layer multi-pass welding. They are disturbed by arc light, splashing and background noise, and are insufficient in real time, so they cannot meet the needs of high-speed welding.

Method used

A multi-layer multi-pass weld feature recognition method that combines the residual network (ResNet) and the conditional generation adversarial network (CGAN) is used, and combined with the color attention mechanism and auxiliary center of gravity loss function, high-quality weld feature maps are generated through the generator network to suppress background interference and improve positioning accuracy.

Benefits of technology

It realizes high-precision and real-time weld information extraction, reduces positioning error fluctuations, improves the intelligence level and welding quality of robot welding, and meets the needs of industrial high-speed welding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563403A_ABST
    Figure CN120563403A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of weld defect image recognition, and particularly relates to a ResNet-CGAN-based multi-layer and multi-pass weld feature recognition method and system, and the method comprises the steps: collecting the image data of multi-layer and multi-pass welding; inputting the image data into a pre-trained model fusing a residual network and a conditional generative adversarial network, wherein the model comprises a generator network and a discriminator network; the generator network is constructed based on an encoder-decoder structure of a residual network, and a color attention mechanism is fused to perform enhancement processing so as to generate a high-quality feature map; the discriminator network discriminates the high-quality feature map and the real feature map, and performs optimization training on the generator network; and extracting a welding line center line and welding line feature point information from a high-quality feature map output by the optimized generator network. According to the method, high-precision and high-real-time welding seam information extraction is realized, and reliable data support is provided for robot path planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of weld defect image recognition, and specifically relates to a multi-layer and multi-pass weld feature recognition method and system based on ResNet-CGAN. Background Art

[0002] Welding is a critical process in modern manufacturing, especially when joining medium-thick plates with a thickness exceeding 4mm. Multi-layer, multi-pass welding processes are indispensable. In multi-layer, multi-pass welding, multi-layer refers to the fusion of two or more weld layers to complete the entire weld, while multi-pass refers to the fusion of two or more weld passes to complete the entire weld. As the number of layers increases, the number of passes per layer also increases, and the weld centerline and weld characteristic points also change. However, traditional welding methods (such as manual or semi-automatic welding) are inefficient and labor-intensive, making them difficult to adapt to the needs of automated production. Robotic automated welding has emerged as a result, and accurate and real-time acquisition of weld information is the core of realizing intelligent welding.

[0003] The research and application of intelligent and efficient welding of medium and thick plates combining laser vision sensors and deep learning methods has become the key to improving the quality of multi-layer and multi-pass welding processes.

[0004] Aiming at the problem of blurred groove features caused by molten pool accumulation, arc interference and thermal deformation in multi-layer and multi-pass welding, the following technical difficulties are mainly solved: 1. Fuzzy groove information features and difficult positioning: As the number of weld layers increases, molten pool accumulation and thermal deformation cause the groove boundary to become blurred. Traditional machine vision methods have difficulty in stably extracting high-precision weld information. 2. Severe noise interference during welding: Arc light, spatter and background noise during welding significantly reduce the image signal-to-noise ratio, affecting the detection of weld feature points; 3. Decreased feature consistency: In the weld bead superposition scenario, the weld centerline and weld feature points change dynamically, resulting in large error fluctuations when extracting features using traditional models. 4. Insufficient real-time performance: Existing methods are difficult to meet the industrial demand for real-time feature extraction (>30 FPS) in high-speed welding scenarios.

[0005] The most important step in intelligent multi-layer, multi-pass welding is to detect and acquire basic weld bead information, including bead position, height, width, and initial point information. Traditional machine vision methods can effectively perform initial weld positioning, but the arc light during the welding process significantly affects their positioning accuracy, ultimately leading to larger errors and a serious decline in weld quality. With the rapid development of industrial automation, laser vision sensors have been widely used in the field of robotic automated welding. Compared with traditional detection methods, this sensor has the ability to quickly capture and process images, providing higher measurement accuracy and resolution. In addition, laser vision sensors can be well integrated with robotic automation systems to provide real-time monitoring and feedback. In multi-layer, multi-pass welding, laser vision sensors use lasers as the active light source. Combined with CCD vision sensors, they can extract weld depth information and accurately restore the weld bead shape. This feature makes them particularly effective in capturing images of groove welds.

[0006] Deep learning technology has made significant progress in computer vision and image processing, particularly in object detection and feature extraction, significantly improving the effectiveness of related research. For detecting feature points in multi-layer, multi-pass weld grooves, deep learning models can automatically extract the centerline of laser lines from weld images, effectively identifying and extracting weld feature points. This approach not only improves the accuracy and efficiency of weld feature point detection but also provides new insights into welding trajectory planning. By analyzing weld characteristics and providing real-time feedback, deep learning technology enables more precise welding path design and optimizes the welding process, effectively improving welding quality and production efficiency.

[0007] The traditional CNN model suffers from a decrease in feature map resolution due to gradient vanishing in deep networks, resulting in loss of deep network details and inability to restore fuzzy groove boundaries. Existing methods do not design an attention mechanism for laser color, resulting in insufficient color sensitivity and weak background noise suppression capabilities. The generated weld feature points are loosely distributed, resulting in large geometric matching errors with real welds and a lack of geometric consistency constraints. Traditional models have low frame rates due to parameter redundancy and cannot meet high-speed welding requirements. Summary of the Invention

[0008] Purpose of the invention: The purpose of the present invention is to solve the problem of restoring and identifying multi-layer and multi-pass welds and weld feature points in the above-mentioned complex situations. A multi-layer and multi-pass weld feature recognition method and system that integrates residual network (ResNet) and conditional generative adversarial network (CGAN) is proposed to solve the problem of fuzzy groove feature extraction; a color attention mechanism (COLOR-ATTENTION) is designed to enhance the model's sensitivity to the weld centerline (red / green / blue) and suppress background interference; an auxiliary center of gravity loss function is introduced to constrain the spatial distribution of weld feature points through connected domain analysis, reduce positioning error fluctuations, achieve high-precision and high-real-time weld information extraction, and provide reliable data support for robot path planning.

[0009] Technical solution: The multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN of the present invention includes the following steps: Collect image data of multi-layer and multi-pass welding; The image data is input into a pre-trained model that integrates a residual network and a conditional generative adversarial network, the model comprising a generator network and a discriminator network; wherein: the generator network is constructed based on the encoder-decoder structure of the residual network and incorporates a color attention mechanism to enhance the weld centerline and weld feature point area of ​​a preset color channel to generate a high-quality feature map; the discriminator network distinguishes between the high-quality feature map and the true feature map annotation, and optimizes the generator network using a combined loss function that combines the conditional generative adversarial network loss, the pixel-based content loss, and the geometric constraint loss based on the spatial distribution of the weld feature points; The weld centerline and weld feature point information are extracted from the high-quality feature map output by the optimized generator network.

[0010] The encoder progressively extracts multi-scale features of the image, and the decoder combines this information from the encoder (via skip connections) to restore spatial resolution and detail. Crucially, the generator network (G) incorporates a color attention mechanism. The generator network (G) learns to map the input low-quality weld image to a high-quality feature map. The output high-quality feature map significantly enhances and sharpens the weld centerline and weld feature points.

[0011] During the training phase, the discriminator network's task is to distinguish between the "fake" high-quality feature maps generated by the generator network and the real feature map annotations (typically ideal feature maps accurately annotated by humans or obtained through other high-precision methods). Trained within the framework of a conditional generative adversarial network (CGAN), the generator network (G) strives to generate increasingly realistic feature maps to "fool" the discriminator network (D), while the discriminator network continuously improves its discriminative capabilities. This adversarial process drives the generator network to learn to generate feature maps that are highly similar to the real annotations, both visually and in distribution.

[0012] In order to accurately guide the training of the generator network, the present invention adopts a combined loss function, which includes not only the standard CGAN loss (used to ensure the authenticity of the generated image), but also includes: pixel-based content loss: used to ensure that the generated image is similar to the real annotation at the pixel level and geometric constraint loss based on the spatial distribution of weld feature points, which aims to directly constrain the spatial position of the generated feature points to make it more consistent with the geometry of the actual weld.

[0013] After the model training is completed, the real-time weld image is input into the optimized generator network (G), and the required weld centerline and weld feature point information (for example, groove vertex and bottom point coordinates) are extracted from the high-quality feature map it outputs. The extracted information has high precision and high robustness.

[0014] To further improve the above technical solution, the generator network includes a residual block, an encoder, and a decoder. The residual block transmits information between the encoder and decoder paths through jump connections, and the encoder and decoder are both embedded with the color attention mechanism.

[0015] The ResNet structure utilizes residual block design and allows information to be directly transferred across layers through skip connections. This effectively alleviates the vanishing gradient problem in deep network training, allowing the network to be deeper. At the same time, it can effectively transfer shallow detail information in the encoder path (which is important for recovering blurred boundaries) to the decoder path, improving the detail recovery capability of feature maps.

[0016] Furthermore, the color attention mechanism enhances the response to preset colors by assigning learnable weights to different color channels and utilizing local maximum pooling operations.

[0017] Furthermore, the preset color corresponding to the weld centerline is red; the weld feature points include a weld feature point at the bottom of the weld bead and a weld feature point at the top of the weld bead, and the corresponding preset colors are green and blue respectively.

[0018] Furthermore, the assigning of learnable weights to different color channels includes independently generating a learnable scaling factor as a weight for each color channel; the local maximum pooling operation includes using sliding window maximum pooling to extract the highest response of the local area; and the color attention mechanism includes: multiplying the output of the local maximum pooling operation with the learnable weight of the corresponding channel, and obtaining the attention weight of each channel through a Sigmoid activation function, and multiplying the attention weight element-by-element by the original feature map to enhance the response of the preset color.

[0019] Furthermore, the geometric constraint loss based on the spatial distribution of the weld feature points is a centroid loss function, which is calculated by calculating the difference between the centroid of the weld feature point area in the feature map generated by the generator network and the centroid of the corresponding weld feature point area in the real feature map annotation, so as to improve the geometric consistency of the generated weld feature points.

[0020] Furthermore, the calculation of the centroid loss function includes: performing connected domain analysis on the weld feature point areas in the generated high-quality feature map and the real feature map annotation, identifying independent feature areas, calculating the centroid coordinates of the pixel distribution of each connected domain, and using the Huber loss function to calculate the difference in the centroid coordinates of the corresponding connected domains.

[0021] Furthermore, the pixel-based content loss is a Smooth L1 loss function, which combines the advantages of L1 loss being insensitive to outliers and L2 loss being smooth and easy to optimize when the error is small, helping to improve model robustness and training stability.

[0022] Furthermore, the extracted weld centerline and weld feature point information are used for real-time planning or adjustment of the robot welding path.

[0023] In addition, the present invention also provides a multi-layer and multi-pass weld feature recognition system based on ResNet-CGAN, which is intended to implement the above method and includes: a laser vision sensor configured to collect image data including a weld centerline and weld feature points; A processor; configured to: receive image data collected by the laser vision sensor; input the image data into a generator network, wherein the generator network is constructed based on an encoder-decoder structure of a residual network and incorporates a color attention mechanism, and the generator network is optimized and trained to generate a high-quality feature map containing an enhanced weld centerline and weld feature points; wherein the optimized training of the generator network utilizes a discriminator network to discriminate between the high-quality feature map and a true feature map annotation, and is performed through a combined loss function that combines a conditional generative adversarial network loss, a pixel-based content loss, and a geometric constraint loss based on the spatial distribution of the weld feature points; A feature extraction module is used to extract weld centerline and weld feature point information from the high-quality feature map output by the optimized and trained generator network; The output module is used to output the weld center line and weld feature point information for real-time planning or adjustment of the welding robot's path.

[0024] Beneficial effect: Compared with the existing technology, the advantages of the present invention are: the present invention integrates the deep feature extraction capability of ResNet and the generative distribution fitting capability of CGAN, combined with SmoothL1 content loss, and can accurately restore weld details from blurred and noisy images. Experiments have shown that error indicators such as MAE and RMSE are significantly reduced.

[0025] The color attention mechanism proposed in this paper effectively focuses on the color information of key laser streaks and weld feature points, suppressing interference from strong background noise such as arc light and spatter. The introduced center of gravity loss function directly constrains the generated weld feature points in terms of spatial distribution, making them more consistent with the geometry of the actual weld. This reduces positioning error fluctuations caused by factors such as weld bead accumulation and deformation, and avoids the structural bias that can result from relying solely on pixel-level loss. Experiments demonstrate stable performance (minimal fluctuations in mPA and mIoU) across different weld bead layers, demonstrating strong robustness and anti-interference capabilities.

[0026] Through reasonable network structure design, the present invention achieves a high processing speed (experimentally reaching 36.72FPS) while ensuring accuracy, meeting the real-time requirements of industrial high-speed welding and significantly improving the intelligence level and welding quality of robot multi-layer and multi-pass welding. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a schematic diagram of the overall architecture of the Resnet-CGAN network provided by the present invention.

[0028] Figure 2 Schematic diagram of the residual block structure used in the present invention.

[0029] Figure 3: is the error curve of the welding seam feature point extraction result in the present invention, where: (a) first pass; (b) second pass; (c) third pass; (d) fourth pass; (e) fifth pass; (f) sixth pass; Figure 4 Schematic diagram comparing the mean absolute errors of U-net, U-net++, and Resnet-CGAN of the present invention; Figure 5 : This is a schematic diagram comparing the root mean square error of U-net, U-net++ and Resnet-CGAN in the present invention; Figure 6 This is a schematic diagram of the visualization results of the Resnet-CGAN model training process of the present invention. DETAILED DESCRIPTION

[0030] The technical solution of the present invention is described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the embodiments.

[0031] Example 1: This invention provides a multi-layer, multi-pass weld feature recognition method based on ResNet-CGAN, which combines the residual structure of a conditional generative adversarial network (CGAN) with a residual network (ResNet). ResNet-CGAN consists of a generator network G, a color attention mechanism C, and a discriminator network D.

[0032] 1. Network model construction The generator network G is based on the ResNet framework, incorporating an encoder-decoder model. This architecture's key feature is its use of "residual blocks" to overcome the vanishing and exploding gradient problems associated with adversarial network training. By introducing skip connections, ResNet enables cross-layer information transfer, facilitating the training of deep networks and improving performance. Furthermore, a color attention mechanism is embedded in the encoder and decoder paths, dynamically focusing on important features across different color channels. The discriminator network D comprises multiple convolutional layers, progressively extracting spatial features from the image. Through layer-by-layer convolutional operations, it gradually learns from local features (such as edges and textures) to global structural features (such as shapes and patterns), ultimately mapping these features into a scalar for distinguishing real and fake images.

[0033] The overall architecture of the network model is as follows Figure 1 As shown, it includes input image, generator, and discriminator.

[0034] The input image (Capture image) is the original weld image captured by the laser vision sensor, with a size of 640*640 pixels and 3 color channels (RGB).

[0035] Generator Network ( Figure 1 The overall structure of the main path in the upper middle section adopts an encoder-decoder architecture and integrates the residual block and color attention (CA) mechanism.

[0036] The generator network's task in the model is to generate high-quality feature maps from low-quality weld images, where weld feature point information is amplified and enhanced. The generator uses a structure with downsampling layers, residual blocks, upsampling, and skip connections to ensure both spatial detail and global feature representation.

[0037] The input image is first down-sampled, with the spatial resolution gradually halved (from 1x to 1 / 8x). This process aims to extract multi-scale contextual features of the image. The color attention mechanism (CA) is applied in this path to enhance the feature response to specific colors. Residual Block: In the deepest layer of the network (where the spatial resolution is lowest), multiple residual blocks are used to perform deep feature transformation. The residual blocks utilize internal skip connections to help alleviate the gradient vanishing problem and enable the network to learn more complex feature representations. Up-sampling: The feature map processed by the residual block is passed through a series of upsampling modules, with the spatial resolution gradually doubled (from 1 / 8x to 1x). This process aims to reconstruct high-resolution details based on the extracted deep features and combined with possible skip connections from the encoder to the decoder. The color attention (CA) module is also applied in this path to maintain the focus on key color features during the reconstruction process.

[0038] The generator finally outputs a high-quality feature map (GenerateImage) with the same size as the input (or target size), in which the weld centerline (red) and weld feature points (green / blue) are clearly and accurately restored and enhanced.

[0039] Discriminator network ( Figure 1During the training phase, the discriminator network (the structure and associated losses in the lower middle section) distinguishes between "fake" images (Generated Images) generated by the generator and real "target" images (Labeled Images). The discriminator network's task is to classify the input image as real or fake. The discriminator enables the generator to continuously optimize itself through adversarial training. Each time the generator generates an image, it attempts to "trick" the discriminator into believing it is real. In this way, the generator and discriminator compete with each other, driving the model to continuously improve the quality of generated images.

[0040] Receives a pair of images as input, extracts features from the input feature map and finally judges its authenticity, outputs a discrimination result, and calculates the discriminator loss, thereby obtaining the total discriminator loss (L_discriminator = l_real + l_fake).

[0041] Discriminator loss (L_discriminator): Calculated based on the discriminator's discrimination results for real sample pairs and fake sample pairs, the goal is to maximize the ability to distinguish between true and false.

[0042] Generator loss (L_generator); is a combined loss used to guide the optimization of the generator; adversarial loss (g_loss: derived from the feedback of the discriminator, the goal is to minimize the probability that the generated image is judged as fake (i.e., "deceiving" the discriminator); content loss (SmoothL1_loss) calculates the pixel-level difference between the generated image and the real annotated image (such as Smooth L1 distance) to ensure that the generated content is consistent with the target, and λ is its weight coefficient; geometric loss (Centroid_loss) calculates the difference between the center of gravity of the weld feature point area in the generated image and the center of gravity of the corresponding area in the real annotation, which is used to constrain the spatial position accuracy of the generated weld feature points, and ω is its weight coefficient.

[0043] 2. Residual Module This invention introduces "Residual Connections" to the conditional adversarial network structure, which enables it to have better performance and stability when training deep networks. This is a simple but very effective design method that adds skip connections between different layers of the network, allowing information to be transmitted directly bypassing some intermediate layers. Figure 2As shown in Figure 1, the basic unit of ResNet is the residual block. In a residual block, the input data is not only processed by the normal convolution layer, but also passed directly to the output of the block through a "shortcut". The output of the residual block is calculated by adding the original input and the output after the convolution operation. This means that the network can directly utilize the original input information while learning more complex features.

[0044] 3. Color Attention Mechanism (COLOR-ATTENTION) In convolutional neural networks (CNNs), conventional convolution operations often treat different color channels equally. However, for many computer vision tasks, color information plays a crucial role in feature extraction. To extract weld centerlines and weld feature points, this paper uses red, green, and blue to annotate centerlines and weld feature points. Therefore, this paper designs a color attention mechanism that assigns different weights to specific color channels, enabling the model to focus on key color regions while suppressing less important ones.

[0045] The primary goal of this module is to enhance the model's performance through adaptive color weighting, simulating the human visual system's sensitivity to different colors. By calculating global color statistics and generating color weights, the module annotates the laser centerline and weld features with red, green, and blue. The color attention mechanism focuses on the centerline and weld features, achieving fine-grained control of input features.

[0046] The main design idea of ​​the color attention module is to generate a learnable scaling factor for each color channel (red, green, blue) to improve the color sensitivity of the model. The weight tensor shape of each color channel is , where C is the number of channels of the input features. These weight tensors serve as trainable parameters of the model and are updated according to the gradient of the loss function during backpropagation, allowing the model to gradually learn the importance of different colors. To better focus on local salient areas, this paper uses sliding window max pooling instead of traditional global average pooling. Local max pooling extracts the maximum value within each k×k window, allowing the model to capture the highest response within the region. Compared to average pooling, max pooling pays more attention to strong responses in local areas and can effectively filter out noise.

[0047] (1) Indicated in pixels The window area is centered on . b is the batch index, and c is the channel index. By sliding the window, the model generates a local maximum response for each pixel.

[0048] After completing the maximum pooling, the present invention multiplies the pooling result with the learnable weight of each channel and compresses the result to the range of [0, 1] using the Sigmoid activation function. After calculating the attention weight of each channel, the present invention multiplies it element-wise by the original feature map, so that the salient area of ​​each channel is enhanced: (2) in, is the cth color channel of the input feature map, The sigmoid function dynamically adjusts the weights of each channel to focus the model on the weld centerline (red) and weld feature points (blue / green), suppressing background interference.

[0049] 4. Loss Function The choice of loss function is crucial for optimizing model training. The L1 norm (absolute error) and L2 norm (mean squared error) are the most commonly used loss functions in adversarial networks. However, raw laser welding images are subject to significant noise, such as reflections. While the L1 loss function effectively reduces attention to abnormal locations, it also tends to focus on weld bends, resulting in large errors. While the L2 loss converges quickly, it focuses too much on noise, resulting in low weld feature extraction accuracy and a tendency to cause gradient explosion. The Smooth L1 loss function combines the advantages of both L1 and L2. When the difference between the predicted and actual values ​​is small (less than a threshold δ), Smooth L1 behaves like L1 for smoothing. When the difference is large (greater than the threshold δ), it transforms into L2, reducing sensitivity to outliers. This design allows Huber Loss to maintain the efficiency of L2 in most situations while maintaining the robustness of L1 in the presence of outliers.

[0050] (3) To improve the positional accuracy of weld feature points in images generated by a conditional generative adversarial network (CGAN), this paper designs a centroid loss function as an auxiliary loss term in the generator training. This loss function calculates the difference in the centroid positions of weld feature points in the generated image and the real image, guiding the generator to more accurately generate weld feature points that conform to the actual weld distribution. This loss function directly constrains the spatial distribution of weld feature points in the generated and real images, making the generated weld feature points more consistent with the actual weld shape and position. After the weld feature points are stably generated, the weld feature point regions of the green and blue channels are extracted from the generated and real images, respectively. Connected domains are then labeled for the extracted weld feature point regions, and independent weld feature regions are identified. For each connected domain, the centroid coordinates of its pixel distribution are calculated. The Huber loss function is then used to calculate the difference in the centroid coordinates of the corresponding connected domain in the generated image and the real image, using the following formula: (4) in, and are the centroid coordinates of the generated image and the real image of the i-th connected domain, respectively. Through connected domain analysis and centroid calculation, this method can effectively identify the complex distribution of weld feature points in multi-layer and multi-pass weld images, avoiding feature loss caused by noise or over-smoothing. The network combines GAN loss, L1 loss, and centroid loss to optimize G: (5) in, and is the loss weight.

[0051] Example 2: To improve network reliability, this example used a Yaskawa robot 3D scanning device to collect a real-world dataset. Six weld bead samples, ranging from 1 layer and 1 pass to 3 layers and 6 passes, were collected for experimentation. A total of 900 images with a resolution of 1280 x 960 pixels were collected as the training dataset. To avoid introducing excessive noise and useless information that could affect training efficiency, a region-of-interest (ROI) was used to extract images with a resolution of 640 x 640 pixels as the network input.

[0052] The model was developed using Python 3.11, Torch 2.0.0, and the cu118 library. Windows 11 was used as the development platform, equipped with an Intel(R) Core(TM) i7-14700KF CPU and an NVIDIA GeForce RTX-4090 GPU for training and testing. The initial learning rate of the neural network was 0.0005, and the Adam optimizer was used to adaptively optimize the learning rate.

[0053] To verify the effectiveness of laser line and weld feature point extraction, the proposed network was used to quantitatively evaluate each weld bead using mean pixel accuracy (mPA) and mean intersection pair union (mIOU). Unlike conventional pixel evaluation, which is meaningless due to the large amount of black background in the image, this method uses red, green, and blue as index categories to evaluate the weld centerline, weld bottom, and weld top features in the generated image.

[0054] (6) (7) At the same time, in order to verify the accuracy of the weld centerline and weld feature points, the mean absolute error (MAE) of the laser lines and weld feature points of 30 weld centerline images was evaluated.

[0055] (8) To further validate the advantages of the proposed network, we compared it with existing network architectures, U-net, and U-net++. The root mean square error (RMSE) and mean absolute error (MAE) were used to quantitatively evaluate the prediction results of the three networks for weld feature information.

[0056] Evaluation Results: As shown in Table 1, as the number of weld layers increases (Layer 1 to Layer 6), the mPA and mIoU fluctuate by only ±3%, demonstrating the model's robustness to multi-layer stacking scenarios. The decrease in accuracy (mPA = 89.8%) in Layer 4 is due to local feature blurring caused by weld slag accumulation in this layer. However, the centroid loss function constraint validates the model's ability to resist interference.

[0057] Table 1. Number of weld layers and corresponding mPA and mIoU evaluation table Weld pass mPA / % mIoU / % Layer 1 93.7 87.4 Layer 2 94.0 87.6 Layer 3 94.4 88.4 Layer 4 89.8 83.1 Layer 5 92.8 85.9 Layer 6 90.8 83.9 like Figure 3 As shown in the figure, the ResNet network has a good ability to extract the characteristic points of each weld layer. For images with gentle slope angles, the mean absolute error is relatively high, basically maintained within 2 pixels. The average value of each layer does not exceed 0.6 pixels.

[0058] like Figure 4 、 Figure 5As shown in the figure, the MAE and RMSE of the ResNet-CGAN (corresponding to ours in the figure) provided by the present invention are significantly lower than those of the U-net series (U-net++, U-net). The reasons are: 1) the residual connection alleviates the vanishing gradient of the deep network and retains more detailed features; 2) the adversarial training of CGAN enhances the geometric consistency between the generated image and the real distribution; 3) the color attention mechanism effectively suppresses arc interference (such as Figure 5 In addition, U-net++ introduces redundant parameters due to dense skip connections, resulting in insufficient real-time performance (FPS=28.5), while ResNet-CGAN achieves both accuracy and efficiency through a lightweight design.

[0059] Among the three network models, the Resnet-CGAN network has the lowest MAE and RMSE errors and the smallest curve fluctuation, which shows that the Resnet-CGAN network has the best feature extraction ability among the three and has strong robustness. The visualization effect is as follows Figure 6 shown.

[0060] Example 3: The present invention proposes a system for implementing a multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN, the system comprising: Laser vision sensor: installed on the robot's end effector to capture weld images.

[0061] Processor: an application program for implementing the method in Example 1 (including a pre-trained generator G model, a discriminator D model, a loss function calculation module, and a feature extraction module).

[0062] Robot controller: receives the weld feature information output by the processing unit and controls the movement of the welding robot.

[0063] Workflow: The laser vision sensor collects weld images and transmits them to the processor. The processor executes the application, calls the pre-trained generator G model to process the image, and generates a high-quality feature map. The feature extraction module extracts laser stripes and feature point information from the feature map. The processor sends this information to the robot controller, which adjusts the posture and trajectory of the welding robot's welding gun accordingly for precise welding.

[0064] The present invention has the following advantages: through the ResNet-CGAN fusion of deep features and adversarial training, the MAE is reduced by 30%, achieving high-precision extraction of weld information; the color attention mechanism suppresses arc and spatter noise, and the weld feature point positioning error fluctuation is less than 0.6 pixels, with strong anti-interference ability; the lightweight design enables a frame rate of 36.72 FPS, supporting high-speed welding scenarios and excellent real-time performance; the center of gravity loss function ensures that the generated weld feature points match the spatial distribution of the real weld, with geometric consistency constraints.

[0065] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as limiting the present invention itself. Various changes may be made to it in form and detail without departing from the spirit and scope of the present invention as defined in the appended claims.

Claims

1. A multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN, characterized by: The following steps are involved: Collect image data of multi-layer and multi-pass welding; The image data is input into a pre-trained model that integrates a residual network and a conditional generative adversarial network, the model comprising a generator network and a discriminator network; wherein: the generator network is constructed based on the encoder-decoder structure of the residual network and incorporates a color attention mechanism to enhance the weld centerline and weld feature point area of ​​a preset color channel to generate a high-quality feature map; the discriminator network distinguishes between the high-quality feature map and the true feature map annotation, and optimizes the generator network using a combined loss function that combines the conditional generative adversarial network loss, the pixel-based content loss, and the geometric constraint loss based on the spatial distribution of the weld feature points; The weld centerline and weld feature point information are extracted from the high-quality feature map output by the optimized generator network.

2. The multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN according to claim 1 is characterized in that: The generator network includes a residual block, an encoder, and a decoder. The residual block transfers information between the encoder and decoder paths through a jump connection. The encoder and decoder are both embedded with the color attention mechanism.

3. The multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN according to claim 1, characterized in that: The color attention mechanism enhances the response to preset colors by assigning learnable weights to different color channels and utilizing local maximum pooling operations.

4. The multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN according to claim 3 is characterized in that: The preset color corresponding to the weld centerline is red; the weld feature points include the weld bottom weld feature point and the weld top weld feature point, and the corresponding preset colors are green and blue respectively.

5. The multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN according to claim 3 or 4, characterized in that: The assigning of learnable weights to different color channels includes independently generating a learnable scaling factor as a weight for each color channel; the local maximum pooling operation includes using sliding window maximum pooling to extract the highest response of the local area; and the color attention mechanism includes: multiplying the output of the local maximum pooling operation with the learnable weight of the corresponding channel, and obtaining the attention weight of each channel through a Sigmoid activation function, and multiplying the attention weight element-by-element by the original feature map to enhance the response of the preset color.

6. The multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN according to claim 1, characterized in that: The geometric constraint loss based on the spatial distribution of the weld feature points is a centroid loss function, which is calculated by calculating the difference between the centroid of the weld feature point area in the feature map generated by the generator network and the centroid of the corresponding weld feature point area in the real feature map annotation, so as to improve the geometric consistency of the generated weld feature points.

7. The multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN according to claim 6, characterized in that: The calculation of the centroid loss function includes: performing connected domain analysis on the weld feature point areas in the generated high-quality feature map and the real feature map annotation, identifying independent feature areas, calculating the centroid coordinates of the pixel distribution of each connected domain, and using the Huber loss function to calculate the difference in the centroid coordinates of the corresponding connected domains.

8. The multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN according to claim 6, characterized in that: The pixel-based content loss is a Smooth L1 loss function.

9. The multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN according to claim 1, characterized in that: The extracted weld centerline and weld feature point information are used for real-time planning or adjustment of the robot welding path.

10. A system for implementing the multi-layer and multi-pass weld feature recognition method based on ResNet-CGAN according to claim 1, characterized in that: include: a laser vision sensor configured to collect image data including a weld centerline and weld feature points; processor; The system is configured to: receive image data collected by the laser vision sensor; input the image data into a generator network, wherein the generator network is constructed based on an encoder-decoder structure of a residual network and incorporates a color attention mechanism, and the generator network is optimized and trained to generate a high-quality feature map containing enhanced weld centerlines and weld feature points; wherein the optimized training of the generator network utilizes a discriminator network to discriminate between the high-quality feature map and true feature map annotations, and is performed through a combined loss function that combines a conditional generative adversarial network loss, a pixel-based content loss, and a geometric constraint loss based on the spatial distribution of the weld feature points; A feature extraction module is used to extract weld centerline and weld feature point information from the high-quality feature map output by the optimized and trained generator network; The output module is used to output the weld center line and weld feature point information for real-time planning or adjustment of the welding robot's path.

Citation Information

Patent Citations

  • Original generative adversarial network model-based residual error network method

    CN107944546A

  • Weld seam feature point extraction method and device, electronic equipment and storage medium

    CN113674218A

  • Multi-vision task filling container defect detection method and system and medium

    CN117173461A

  • X-ray weld defect image reconstruction method based on improved SRGAN

    CN118314018A

  • Lightweight image style migration method based on color and texture dual channels

    CN119295296A

Cited By

  • A robot welding guidance method and system based on weld geometric feature points

    CN122574032A