Welding penetration state recognition method based on improved YOLO X model

By improving the YOLO X model and combining it with welding current and speed parameters, extracting the molten pool image features and performing multimodal fusion, the problems of low efficiency and low accuracy in welding penetration state recognition in traditional methods are solved, and high-precision welding quality monitoring is achieved.

CN120611210APending Publication Date: 2025-09-09GUANGXI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510604410.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Traditional welding penetration state recognition methods are inefficient and have low recognition accuracy when processing images in complex environments. They also fail to effectively integrate welding parameters and image features, making it difficult to ensure welding quality.

Method used

The improved YOLO X model is used, combined with welding current and welding speed parameters, to extract the molten pool image features through multi-level convolution processing, and the self-attention mechanism is used to perform multimodal feature fusion to achieve accurate identification of the penetration state.

Benefits of technology

The recognition accuracy of deep penetration argon arc welding penetration status has been significantly improved, ensuring effective monitoring and quality control of the welding process and improving the stability and reliability of welding production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611210A_ABST
    Figure CN120611210A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of welding penetration state recognition, in particular to a welding penetration state recognition method based on an improved YOLO X model, which comprises the following steps of: acquiring a molten pool image sequence and welding parameter data of deep penetration argon arc welding in a welding process in real time to form an original multi-mode welding data set; preprocessing the original multi-modal welding data set to obtain standardized multi-modal welding data; the standardized multi-mode welding data are input into a penetration state recognition model for welding penetration state recognition, and a welding penetration state recognition result is obtained; the penetration state recognition model comprises an image feature extraction module, a welding parameter feature embedding module, a multi-modal feature fusion module and a penetration state classification module which are connected in sequence. According to the method, efficient and accurate recognition of the welding penetration state is achieved by fusing the welding parameter data and the multi-modal data of the molten pool image and combining the multi-scale weighted fusion mechanism of the improved YOLO X model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of welding penetration state recognition, and in particular to a welding penetration state recognition method based on an improved YOLO X model. Background Art

[0002] In the field of industrial manufacturing, welding is a key process for metal connection, and its quality directly determines the structural strength and performance of the product. Keyhole Tungsten Inert Gas (K-TIG), as an efficient welding technology, has important applications in the connection of medium and thick metal plates (such as pressure vessels). K-TIG welding has evolved from traditional tungsten inert gas shielded welding. By increasing the welding current and supplementing it with water cooling technology, it effectively avoids the problem of tungsten needle ablation. K-TIG welding can achieve single-sided welding and double-sided forming in the welding of medium and thick plates, eliminating tedious processes such as groove and root cleaning, and significantly improving welding efficiency.

[0003] With the acceleration of the process of industrial automation, the application of automatic welding equipment in welding production is becoming more and more extensive. In the K-TIG welding process, welding current and welding speed are the core parameters that affect the welding quality. Current that is too large or too small, welding speed that is too fast or too slow, can easily lead to welding defects such as incomplete penetration or over-penetration. Among them, incomplete penetration is manifested as a small hole closed by the molten pool, while over-penetration is manifested as the molten pool width being smaller than the small hole width. In manual welding, experienced welders can dynamically adjust parameters by observing the molten pool to ensure quality, but in automated welding, it is necessary to rely on computers to identify the penetration state in real time and adaptively adjust parameters, which puts higher requirements on the accurate detection of the penetration state. At present, the recognition of the penetration state of K-TIG welding mainly relies on machine vision methods and deep learning methods. In machine vision methods, welding monitoring systems usually build image acquisition systems based on cameras, which can be divided into active vision and passive vision according to the type of light source. There are two types of passive vision. Passive vision directly uses welding arc as light source. The system structure is simple, but it is seriously interfered by strong arc light and the image signal-to-noise ratio is low. Active vision improves imaging quality through auxiliary light sources (such as laser or filtering system), but increases the complexity of the equipment. After acquiring the welding image, traditional machine vision methods extract the molten pool morphology or weld deviation information through image processing algorithms (such as region of interest extraction, noise reduction, segmentation and feature extraction, etc.), and then combine classification or regression models to establish the association between image features and penetration state. However, this type of method has significant limitations. Strong arc light, spatter and smoke interference lead to insufficient robustness of image segmentation and feature extraction, and low recognition accuracy. In addition, most methods rely only on image information and do not consider the influence of key process parameters such as welding current and speed. In actual welding, the penetration state is the result of the joint action of multiple factors. Single modal data is difficult to fully reflect the welding state.

[0004] In summary, traditional welding penetration state recognition methods have problems such as low efficiency and low recognition accuracy when processing images in complex environments. Therefore, how to effectively integrate welding parameters and image features to improve the accuracy and stability of welding penetration state recognition has become a key issue that needs to be urgently solved in the current welding field. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a method for identifying welding penetration status based on an improved YOLO X model, the method comprising the following steps:

[0006] Real-time acquisition of a molten pool image sequence and welding parameter data during deep penetration argon arc welding to form an original multimodal welding data set; the welding parameter data includes at least welding current data and welding speed data;

[0007] Preprocessing the original multimodal welding data set to obtain standardized multimodal welding data;

[0008] Constructing a penetration state recognition model based on the improved YOLO X model; wherein the penetration state recognition model includes an image feature extraction module, a welding parameter feature embedding module, a multimodal feature fusion module, and a penetration state classification module connected in sequence;

[0009] The standardized multimodal welding data is input into the penetration state recognition model to perform welding penetration state recognition to obtain a welding penetration state recognition result.

[0010] In a further embodiment, the step of preprocessing the original multimodal welding data set to obtain standardized multimodal welding data comprises:

[0011] performing denoising processing on the melt pool image sequence to generate a standardized melt pool image;

[0012] Normalizing the welding current data and the welding speed data to generate a welding parameter vector;

[0013] The standardized molten pool image and the welding parameter vector are integrated with corresponding time stamps to obtain standardized multimodal welding data.

[0014] In a further embodiment, the step of inputting the standardized multimodal welding data into the penetration state recognition model to perform welding penetration state recognition and obtain a welding penetration state recognition result comprises:

[0015] The standardized molten pool image is input into the image feature extraction module for multi-level convolution processing to extract a multi-scale feature image of the molten pool containing geometric features of the molten pool edge, local morphological features of the small holes, and thermal distribution information;

[0016] The welding parameter vector is mapped to a high-dimensional feature space by the welding parameter feature embedding module, and the welding parameter vector is spatially aligned with the molten pool multi-scale feature image to obtain a welding parameter feature; the number of output channels of the welding parameter feature is the same as the number of channels of the molten pool multi-scale feature image;

[0017] The multi-modal feature fusion module splices the molten pool multi-scale feature image and the welding parameter feature in the channel dimension to obtain a multi-modal fusion feature map;

[0018] Calculating a feature weight coefficient for each pixel position in the multimodal fusion feature map using a self-attention mechanism, and performing pixel-level weighting on the multimodal fusion feature using the feature weight coefficient to obtain an attention-optimized fusion feature;

[0019] The attention optimization fusion feature is input into the penetration state classification module to perform penetration state classification and identification to obtain a welding penetration state identification result; the welding penetration state identification result includes an incomplete penetration state, a normal penetration state and an over-penetration state.

[0020] In a further embodiment, the image feature extraction module includes an input layer, a focusing layer, a convolution-batch normalization-activation function combination layer, a multi-level local residual block body and a global-local residual block body connected in sequence.

[0021] In a further embodiment, the multi-level local residual block body adopts a multi-level cascade architecture, each level of the local residual block body includes a convolution-batch normalization-activation function combination layer and a cross-stage local connection layer connected in sequence, and the number of output channels between the convolution-batch normalization-activation function combination layer and the cross-stage local connection layer of the local residual block body at different levels increases with the increasing level;

[0022] The global-local residual block includes a convolution-batch normalization-activation function combination layer, a spatial pyramid pooling layer, and a cross-stage local connection layer connected in series, wherein the spatial pyramid pooling layer is composed of a multi-scale maximum pooling layer and a convolution-batch normalization-activation function combination layer connected in series; the multi-scale maximum pooling layer includes multiple parallel pooling branches with different pooling sizes.

[0023] In a further embodiment, the step of inputting the standardized molten pool image into the image feature extraction module for multi-level convolution processing to extract a multi-scale feature image of the molten pool containing geometric features of the molten pool edge, local morphological features of the small holes, and thermal distribution information comprises:

[0024] Using the focus layer to perform a slicing and reorganization operation on the standardized melt pool image to obtain spatial detail information of the melt pool area;

[0025] The multi-level local residual block body gradually extracts the geometric features of the edge of the molten pool and the local morphological features of the small holes from the spatial detail information of the molten pool area, thereby obtaining the shallow residual features of the molten pool at different levels;

[0026] Inputting the shallow residual features of the melt pool into the main body of the global-local residual block to perform pooling operations at different scales to obtain multi-scale pooling features;

[0027] The multi-scale pooling feature is combined with the shallow residual feature of the melt pool to obtain the deep feature of the melt pool;

[0028] The deep features of the melt pool are sequentially upsampled, and the deep features of the melt pool after each upsampling are fused with the shallow residual features of the melt pool at different levels through cross-level splicing operations to generate a multi-scale feature image of the melt pool.

[0029] In a further embodiment, the welding parameter feature embedding module comprises a parameter embedding layer and a dimension expansion layer connected in series, wherein the parameter embedding layer is composed of a fully connected layer;

[0030] The fully connected layer is used to map the welding parameter vector to a high-dimensional feature space with the same number of channels as the molten pool multi-scale feature image, to obtain the mapped welding parameter features;

[0031] The dimension expansion layer is used to spatially expand the mapped welding parameter features so that the spatial size of the mapped welding parameter features is consistent with the multi-scale feature image of the molten pool, thereby obtaining the welding parameter features.

[0032] In a further embodiment, the multimodal feature fusion module includes a channel splicing layer, a self-attention mechanism, a cross-stage local connection and splicing layer, a downsampling layer and a dynamic weight adjustment layer.

[0033] In a further embodiment, the self-attention mechanism is composed of a convolutional layer and a Sigmoid activation function.

[0034] In a further embodiment, the melt penetration state classification module includes a YOLO detection head, which includes a convolutional layer and a classifier.

[0035] The present invention provides a method for identifying weld penetration status based on an improved YOLO X-model. The method collects a sequence of weld pool images and welding parameter data during deep penetration argon arc welding in real time to form an original multimodal welding dataset. The original multimodal welding dataset is preprocessed to obtain standardized multimodal welding data. A weld penetration status identification model based on the improved YOLO X-model is constructed. The weld penetration status identification model comprises an image feature extraction module, a welding parameter feature embedding module, a multimodal feature fusion module, and a weld penetration status classification module, which are connected in sequence. The standardized multimodal welding data is input into the weld penetration status identification model to perform weld penetration status identification and obtain a weld penetration status identification result. Compared with the prior art, this method significantly improves the recognition accuracy of deep penetration argon arc welding weld penetration status by fusing multimodal data of welding current parameters, welding speed parameters, and weld pool images, combined with the multi-scale dynamic weighted fusion mechanism of the improved YOLO X-model. This method effectively monitors and controls the welding process, ensuring the stability and reliability of welding production. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 1 is a flow chart of a method for identifying welding penetration status based on an improved YOLO X model provided by an embodiment of the present invention;

[0037] Figure 2 This is a schematic diagram of the welding experiment platform architecture provided by an embodiment of the present invention;

[0038] Figure 3 1 is a schematic diagram of the structure of a melt-through state identification model provided by an embodiment of the present invention;

[0039] Figure 4 Schematic diagram of the welding parameter and image feature fusion process provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0040] The following describes the embodiments of the present invention in detail with reference to the accompanying drawings. The embodiments are provided for illustrative purposes only and are not to be construed as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of protection of the present invention. Many changes may be made to the present invention without departing from the spirit and scope of the present invention.

[0041] refer to Figure 1 , the embodiment of the present invention provides a welding penetration state recognition method based on the improved YOLO X model, such as Figure 1 As shown, the method includes the following steps:

[0042] S1. Real-time acquisition of a molten pool image sequence and welding parameter data during deep penetration argon arc welding to form an original multimodal welding data set; the welding parameter data includes at least welding current data and welding speed data.

[0043] S2. Preprocessing the original multimodal welding data set to obtain standardized multimodal welding data.

[0044] In some embodiments, the step of preprocessing the original multimodal welding data set to obtain standardized multimodal welding data includes:

[0045] performing denoising processing on the melt pool image sequence to generate a standardized melt pool image;

[0046] Normalizing the welding current data and the welding speed data to generate a welding parameter vector;

[0047] The standardized molten pool image and the welding parameter vector are integrated with corresponding time stamps to obtain standardized multimodal welding data.

[0048] In the K-TIG welding process, welding current and welding speed are key factors affecting welding quality. Excessive current or too fast speed may lead to defects such as incomplete or excessive penetration, and the change in the molten pool state directly reflects the welding quality. Traditional machine vision methods are inefficient and inaccurate when processing welding images under strong arc interference, and most methods do not fully consider the influence of welding parameters. Therefore, this embodiment proposes a welding penetration state recognition method based on an improved YOLO X model, combining welding parameters and image features to achieve accurate recognition of the penetration state. Specifically, in the K-TIG welding operation scenario, this embodiment uses a pre-built K-TIG welding experimental platform to collect the molten pool image sequence in the K-TIG welding process in real time through a high-dynamic industrial camera to ensure that the image can clearly reflect the key information such as the process molten pool morphology and keyhole state in the welding process, and synchronously reads key welding parameter data such as welding current data and welding speed data through an industrial computer to ensure that the molten pool image sequence and the welding parameter data correspond one-to-one in the time series to form an original multimodal welding data set. The welding experimental platform architecture is as follows: Figure 2As shown, the molten pool image sequence is then subjected to dynamic range compression and noise suppression to eliminate noise interference in the image, generating a noise-suppressed molten pool image sequence. The noise-suppressed molten pool image sequence is then normalized, and the image size is adjusted to meet the input requirements of the YOLO X model to generate a standardized molten pool image. The welding current data and welding speed data are normalized and scaled to a specified range (e.g., between 0 and 1) through linear transformation to eliminate the influence of different dimensions on model training, generating a numerical welding parameter vector. In this embodiment, the standardized molten pool image and the welding parameter vector are integrated with corresponding timestamps to form standardized multimodal welding data containing image information and welding parameter information, ensuring that each set of data corresponds to the welding state at the same moment.

[0049] S3. Build a melt penetration state recognition model based on the improved YOLO X model.

[0050] In this embodiment, a pre-trained penetration state recognition model is obtained, and standardized multimodal welding data is input into the pre-trained penetration state recognition model to perform penetration state recognition, and a welding penetration state recognition result is obtained. The welding penetration state recognition result includes states such as incomplete penetration, normal penetration, and excessive penetration. Figure 3 As shown, the penetration state recognition model is constructed based on the improved YOLO X model. The penetration state recognition model includes an image feature extraction module, a welding parameter feature embedding module, a multimodal feature fusion module and a penetration state classification module connected in sequence, wherein the image feature extraction module is used to extract image features from the molten pool image; the welding parameter feature embedding module is used to map the welding current and welding speed parameters to a high-dimensional space that matches the image features, and adjust its spatial size through an expansion operation to align with the image feature map; the multimodal feature fusion module is used to splice the image features with the welding parameter features and perform weighted fusion through a self-attention mechanism; the state recognition module is used to classify and identify the penetration state according to the fused features. In this embodiment, the step of inputting the standardized multimodal welding data into the penetration state recognition model to perform welding penetration state recognition and obtain the welding penetration state recognition result includes:

[0051] The standardized molten pool image is input into the image feature extraction module for multi-level convolution processing to extract a multi-scale feature image of the molten pool containing geometric features of the molten pool edge, local morphological features of the small holes, and thermal distribution information;

[0052] The welding parameter vector is mapped to a high-dimensional feature space by the welding parameter feature embedding module, and the welding parameter vector is spatially aligned with the molten pool multi-scale feature image to obtain a welding parameter feature; the number of output channels of the welding parameter feature is the same as the number of channels of the molten pool multi-scale feature image;

[0053] The multi-modal feature fusion module splices the molten pool multi-scale feature image and the welding parameter feature in the channel dimension to obtain a multi-modal fusion feature map;

[0054] Calculating a feature weight coefficient for each pixel position in the multimodal fusion feature map using a self-attention mechanism, and performing pixel-level weighting on the multimodal fusion feature using the feature weight coefficient to obtain an attention-optimized fusion feature;

[0055] The attention optimization fusion feature is input into the penetration state classification module to perform penetration state classification and identification to obtain a welding penetration state identification result; the welding penetration state identification result includes an incomplete penetration state, a normal penetration state and an over-penetration state.

[0056] Specifically, this embodiment uses the cross-stage local connection network (CSPDarknet) in the image feature extraction module to perform multi-level convolution processing on the standardized molten pool image to extract the molten pool multi-scale feature image containing the geometric features of the molten pool edge, the dynamic morphology of the keyhole and the thermal distribution information; at the same time, the welding parameter feature embedding module maps the welding parameter vector to the high-dimensional feature space through a two-layer fully connected network, and the number of its output channels is consistent with the number of channels of the molten pool multi-scale feature image; the welding parameter features are adjusted to the same spatial size as the molten pool multi-scale feature image through the dimensional expansion operation, so as to achieve spatial alignment of the welding parameters and the visual image features; then, this embodiment splices the molten pool multi-scale feature image and the welding parameter features in the channel dimension to generate a multi-dimensional molten pool. modal fusion feature map; the weight coefficient of each pixel position in the multimodal fusion feature map is calculated by the self-attention mechanism, and the self-attention mechanism is composed of a 1×1 convolutional layer and a Sigmoid activation function to generate an attention map of the same size as the input, perform pixel-level weighting on the fusion features, and output the attention optimized fusion features, so that the penetration state classification module performs global average pooling on the attention optimized fusion feature map, and outputs three categories of probability distribution through the Softmax classifier, specifically including the incomplete penetration state (the molten pool completely closes the small hole), the normal penetration state (the molten pool matches the small hole size) and the over-penetration state (the small hole diameter is significantly larger than the molten pool width). The final recognition result is the highest probability category as the output, and the confidence of each category is also output.

[0057] It should be noted that, in this embodiment, the welding parameter vector is mapped to a high-dimensional feature space with the same dimension as the multi-scale feature image of the molten pool through the welding parameter embedding process, so that the features of the two modalities can be fused in the same space. Specifically, the welding speed and welding current are two numerical parameters in the welding process, and their initial shape is [batch_size, 2], where batch_size is the batch size. After being processed by the welding parameter feature embedding module, the welding parameter vector is mapped to the high-dimensional feature space, and the output dimension is base_channels (that is, the number of channels of the image feature), thereby generating a welding parameter feature with a shape of [batch_size, base_channels], thereby achieving effective fusion with the high-dimensional features of the image.

[0058] Then, in order to align the welding parameter features with the molten pool multi-scale feature image in the spatial dimension, this embodiment needs to resize the welding parameter features. Considering that the molten pool multi-scale feature image is a four-dimensional tensor with a shape of [batch_size, base_channels, H, W], H is the height and W is the width, while the current shape of the welding parameter features is only [batch_size, base_channels], which only contains channel information but no spatial information. For this reason, this embodiment uses the unsqueeze method to insert dimensions into the last two dimensions (i.e., spatial dimensions) of the welding parameter features and converts its shape to [batch_size, base_channels, 1, 1]. Next, this embodiment expands this low-dimensional feature vector to the same spatial size [batch_size, base_channels, H, W] as the image feature map through an expand operation, so that each pixel position contains the same welding parameter features. After this operation, this embodiment converts the original low-dimensional welding parameter features into a feature map with the same size as the image feature map.

[0059] In some embodiments, the image feature extraction module includes an input layer Inputs, a focusing layer Focus, a first convolution-batch normalization-activation function combination layer Conv2D_BN_SILU (320, 320, 64) connected in sequence, a multi-level local residual block body, and a global-local residual block body. In this embodiment, the input layer (Inputs) is used to receive a standardized melt pool image of size 640×640×3, where the image size is 640×640 and the number of channels is 3; the focusing layer (Focus) is used to downsample the input image resolution to a resolution of 320×320×12 through a slice reorganization operation, that is, the output feature map size is 320×320 and the number of channels is 12, retaining the spatial detail information of the melt pool area; the convolution-batch normalization-activation function combination layer (Conv2D_BN_SILU) is used to perform convolution (Conv2D), batch normalization (Batch Normalization) and SILU activation function in sequence to extract local texture features.

[0060] The multi-level local residual block body adopts a multi-level serial architecture, and each level of the local residual block body includes a convolution-batch normalization-activation function combination layer (Conv-BN-SiLU layer) and a cross-stage local connection layer (CspLayer) connected in sequence, and the number of channels between the convolution-batch normalization-activation function combination layer and the cross-stage local connection layer of different levels is configured differently. The number of output channels between the convolution-batch normalization-activation function combination layer and the cross-stage local connection layer of the local residual block body at different levels increases with the increasing level, so as to realize the progressive feature extraction of the feature map from shallow to deep and from local to global. In this embodiment, the output of the cross-stage local connection layer of the previous level is directly used as the input of the convolution-batch normalization-activation function combination layer of the next level to form a linear transmission path of the data stream, wherein the cross-stage local connection layer (CspLayer) is composed of a cross-stage local connection network (Cross Stage Partial Connection Network). The residual block body (Resblock_body) contains a residual connection structure, which enhances the deep feature expression capability through residual connections.

[0061] Specifically, the multi-level local residual block body includes at least three levels of residual block bodies, and the at least three levels of residual block bodies are respectively a first-level residual block body, a second-level residual block body and a third-level residual block body. The first-level residual block body includes a second convolution-batch normalization-activation function combination layer and a first cross-stage local connection layer in series, and the number of channels of the second convolution-batch normalization-activation function combination layer and the first cross-stage local connection layer is 128; the second-level residual block body includes a third convolution-batch normalization-activation function combination layer and a second cross-stage local connection layer in series, and the number of channels of the third convolution-batch normalization-activation function combination layer and the second cross-stage local connection layer is 256; The main body of the third-level residual block includes a fourth convolution-batch normalization-activation function combination layer and a third cross-stage local connection layer connected in series. The number of channels of the fourth convolution-batch normalization-activation function combination layer and the third cross-stage local connection layer is 512. Based on this structure, this embodiment gradually extracts the geometric features of the melt pool edge (such as melt pool width, small hole diameter), thermal distribution features (such as high temperature area morphology) and dynamic morphological features (such as small hole closure trend) according to the standardized melt pool image through multi-level convolution and cross-stage local connection, and outputs a multi-scale feature map group containing spatial semantic information. The feature map sizes are 320×320, 160×160, 80×80, 40×40 and 20×20 respectively.

[0062] The global-local residual block includes a fifth convolution-batch normalization-activation function combination layer, a spatial pyramid pooling bottleneck (SPPBottleneck), and a fourth cross-stage local connection layer connected in series. The spatial pyramid pooling layer is composed of a multi-scale maximum pooling layer and a convolution-batch normalization-activation function combination layer connected in series to achieve multi-scale feature aggregation of the input feature map and extract the thermodynamic distribution characteristics and global deformation characteristics of the melt pool. In this embodiment, the spatial pyramid pooling layer is used to fuse different receptive field features through multi-scale pooling operations, capture the global context information of the melt pool, and enhance the scale invariance of the model. The multi-scale maximum pooling layer includes multiple parallel pooling branches with different pooling sizes (such as pooling kernel sizes of 5×5, 9×9, and 13×13, respectively). Based on the structure of the image feature extraction module, in this embodiment, the step of inputting the standardized melt pool image into the image feature extraction module for multi-level convolution processing to extract the melt pool multi-scale feature image containing the geometric features of the melt pool edge, the local morphological features of the small holes, and the thermal distribution information includes:

[0063] Using the focus layer to perform a slicing and reorganization operation on the standardized melt pool image to obtain spatial detail information of the melt pool area;

[0064] The multi-level local residual block body gradually extracts the geometric features of the edge of the molten pool and the local morphological features of the small holes from the spatial detail information of the molten pool area, thereby obtaining the shallow residual features of the molten pool at different levels;

[0065] Inputting the shallow residual features of the melt pool into the main body of the global-local residual block to perform pooling operations at different scales to obtain multi-scale pooling features;

[0066] The multi-scale pooling feature is combined with the shallow residual feature of the melt pool to obtain the deep feature of the melt pool;

[0067] The deep features of the melt pool are sequentially upsampled (upsampling2D), and the deep features of the melt pool after each upsampling are fused with the shallow residual features of the melt pool at different levels through cross-level splicing operations to generate a multi-scale feature image of the melt pool.

[0068] Specifically, such as Figure 4 As shown, this embodiment uses a focusing layer to slice and reorganize the standardized melt pool image according to the input standardized melt pool image, concentrates the width and height information of the image into the channel dimension, reduces the size of the image data by half, increases the number of channels, and obtains the spatial detail information of the melt pool area. The spatial detail information of the melt pool area is input into the first convolution-batch normalization-activation function combination layer for preliminary convolution, batch normalization and activation operations to obtain the shallow features of the image; then, through the first level residual block body in the multi-level local residual block body, the second convolution-batch normalization-activation function combination layer is used to perform convolution, normalization and activation operations on the shallow features of the image, and then the input features are fused with the output features through the first cross-stage local connection layer to obtain the first level residual features of 160×160×128; the subsequent second level residual block body and the first The main body of the three-level residual block continues to perform further feature extraction and fusion on the residual features of the first level to enhance the feature expression ability. In the main body of the second-level residual block, the residual features of the first level are convolved, normalized and activated through the third convolution-batch normalization-activation function combination layer to enhance the local feature expression ability. Then, the geometric features of the edge of the melt pool (such as the ratio of the melt pool width to the hole diameter) are extracted through the second cross-stage local connection layer, and the second-level residual features of 80×80×256 are output; in the main body of the third-level residual block, the residual features of the second level are convolved, normalized and activated through the fourth convolution-batch normalization-activation function combination layer, and then the feature redundancy is optimized through the third cross-stage local connection layer to obtain the third-level residual features of 40×40×512, that is, the shallow residual features of the melt pool.

[0069] In this embodiment, the 40×40×512 shallow residual features of the melt pool output by the multi-level local residual block body are input into the global-local residual block body, and the fifth convolution-batch normalization-activation function combination layer in the global-local residual block body is used to perform convolution, normalization and activation operations on the shallow residual features of the melt pool. Then, the multi-scale maximum pooling layer in the spatial pyramid pooling layer is used to perform pooling operations of different scales on the feature map to extract thermodynamic distribution features and global deformation features of different scales, and then further fused through subsequent convolution-batch normalization-activation function combination layers to obtain multi-scale pooling features, and the multi-scale pooling features are combined with the shallow residual features of the melt pool. The features are fused through the fourth cross-stage local connection layer to obtain 20×20×1024 deep features of the melt pool. The shallow residual features and deep features of the melt pool at different levels are upsampled and spliced ​​across levels. Specifically, the 20×20×1024 deep features of the melt pool are upsampled to 40×40 and 80×80 resolutions respectively, and fused with the residual features with resolutions of 40×40×512 and 80×80×256 across levels, so that features of different scales are effectively integrated, and finally a multi-scale feature image of the melt pool containing rich multi-scale information is output. The image includes the geometric features of the melt pool edge, the dynamic morphology of the small holes and the thermal distribution information.

[0070] In some embodiments, the welding parameter feature embedding module includes a parameter embedding layer (Embedding) and a dimensionality expansion layer (unsqueeze & expand) connected in series, wherein the parameter embedding layer is composed of a fully connected layer, and the fully connected layer is used to map the welding parameter vector to a high-dimensional feature space with the same number of channels as the molten pool multi-scale feature image to obtain the mapped welding parameter feature; the dimensionality expansion layer is used to adjust the spatial size of the parameter feature by inserting an empty dimension (unsqueeze) and spatial replication expansion (expand), specifically to spatially expand the mapped welding parameter feature so that the spatial size of the mapped welding parameter feature is consistent with the molten pool multi-scale feature image, thereby obtaining the welding parameter feature and achieving dimensional alignment of the welding parameter and image features.

[0071] In some embodiments, the multimodal feature fusion module includes a channel splicing layer (torch.cat), a self-attention mechanism (Attention), a cross-stage local connection and splicing layer (Concat+CSPlayer), a downsampling layer, and a dynamic weight adjustment layer (Weight); the channel splicing layer is used to splice the image feature map and the welding parameter feature along the channel dimension, and fuse the two through the splicing operation to form a multimodal fusion feature map; the self-attention mechanism is composed of a convolution layer and a Sigmoid activation function, and the feature weight coefficient of each pixel position in the multimodal fusion feature map is calculated by the convolution layer plus the Sigmoid activation function, and the spliced ​​multimodal feature map is adjusted according to the feature weight coefficient. The fused features are weighted and the feature importance of each pixel position is dynamically adjusted; the cross-stage local connection and splicing layer is used to combine feature splicing and cross-stage local connection to optimize multi-scale feature fusion; the dynamic weight adjustment layer is used to dynamically allocate fusion weights according to the feature map level. In the specific implementation process, this embodiment dynamically learns the weight distribution of the molten pool area and the keyhole area through the self-attention mechanism based on the spliced ​​multimodal fusion feature map, suppresses the noise characteristics of the arc interference area, and optimizes the feature transfer path between feature maps of different scales (such as 80×80, 40×40 and 20×20) through cross-stage local connection and dynamic weight adjustment, and outputs an optimized feature map that integrates physical laws (welding parameters) and visual semantics (molten pool morphology).

[0072] In some embodiments, the melt penetration state classification module includes a YOLO detection head (YoloHead), which includes a convolutional layer and a classifier. The YOLO detection head is used to output the melt penetration state category and confidence.

[0073] During the specific implementation process, after the welding parameter features are adjusted to be consistent with the spatial size of the image feature map, this embodiment splices and fuses the adjusted welding parameter features with the image feature map in the channel dimension. Specifically, the image feature map comes from the second-level output of the image feature extraction module, and its dimension is [batch_size, base_channels, H, W]. It is aligned with the spatial dimension of the welding parameter features after spatial expansion (the shape is also [batch_size, base_channels, H, W]). This embodiment uses the torch.cat function of the channel-poor layer to splice the two feature maps along the channel dimension to generate a multimodal fusion feature map. The new feature map dimension is [batch_size, 2*base_channels, H, W]. At this point, the image features and the welding parameter features have been fused in the multimodal fusion feature map, which includes both the spatial information of the image and the welding parameter information.

[0074] Then, this embodiment uses a self-attention mechanism on the multimodal fusion feature map to achieve feature weighted fusion. The self-attention mechanism dynamically adjusts the feature weights of each position according to the context information of each pixel to enhance the influence of useful information. The self-attention mechanism designed in this embodiment is composed of a convolutional layer and a Sigmoid activation function to generate an attention map with the same shape as the multimodal fusion feature map. The attention map elements reflect the importance of the corresponding position of the input feature map, and the value range is between [0, 1]. Each pixel feature is weighted, that is, the pixel feature value is multiplied by the corresponding attention weight. This operation is intended to enable the network to automatically adjust the degree of influence of image features and welding parameters at each pixel position. Through the action of the attention mechanism, the network can learn the importance of welding parameters to the key areas of image understanding, and then dynamically adjust the influence of welding parameters on image features.

[0075] It should be noted that, in this embodiment, the fusion of image feature map and welding parameters is not limited to a single scale, but also involves multi-scale feature fusion. Specifically, on the feature maps of each level output by the image feature extraction module, the welding parameter feature embedding, spatial alignment and channel splicing operations are repeatedly performed to realize the fusion of image feature map and welding parameter feature map at different levels. In the fusion process of each scale, the effect of welding parameters is adjusted by dynamic weights to control the degree of influence at different scales. After the above steps, the image features and welding parameter features are successfully fused, and with the help of dynamic weighting and multi-scale fusion, the network can more efficiently utilize the features of different modalities. The welding parameters are no longer just static inputs, but are flexibly and dynamically fused with image features through weighting and attention mechanisms, thereby improving the model performance. The improved YOLO V5 network structure is shown as follows: Figure 3 As shown, it should be pointed out that Figure 3 The weight operation in

[15] is only performed on the feature branches of 256 channels and 512 channels to balance computational efficiency and feature expression capability.

[0076] S4. Inputting the standardized multimodal welding data into the penetration state recognition model to perform welding penetration state recognition, and obtaining a welding penetration state recognition result.

[0077] In this embodiment, the entire weld penetration state recognition model is built based on an improved YOLO X model. The image feature extraction module extracts high-dimensional features from the weld pool image. The welding parameter feature embedding module maps the welding parameters to a high-dimensional feature space and aligns them with the image features. The multimodal feature fusion module fuses and optimizes these two features. Finally, the weld penetration state classification module generates the weld penetration state recognition result. This structural design fully utilizes both image and welding parameter information to accurately identify the weld penetration state, providing an effective technical means for quality monitoring during K-TIG welding.

[0078] In summary, this embodiment uses a high-dynamic industrial camera to collect welding pool images in real time as the input data source of the improved YOLOv5 target detection network. At the same time, by introducing the two key process parameters of welding current and welding speed into the detection network, deep fusion of multimodal data is achieved. In addition, this embodiment improves the original jump connection structure, enhances the transmission efficiency of multi-scale features, and introduces a self-attention mechanism to achieve adaptive weight distribution of different modal features. Compared with traditional methods, this embodiment significantly improves the recognition accuracy of the penetration state during K-TIG welding, can effectively cope with the interference of complex welding environments such as strong arc light and spatter, and meets the needs of industrial sites for real-time monitoring of welding quality.

[0079] It should be noted that the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.

[0080] An embodiment of the present invention provides a method for identifying weld penetration status based on an improved YOLO X-model. The method collects a sequence of weld pool images and welding parameter data during deep penetration argon arc welding in real time to form an original multimodal welding dataset. The original multimodal welding dataset is preprocessed to obtain standardized multimodal welding data. A weld penetration status identification model based on the improved YOLO X-model is constructed. The weld penetration status identification model includes an image feature extraction module, a welding parameter feature embedding module, a multimodal feature fusion module, and a weld penetration status classification module, which are connected in sequence. The standardized multimodal welding data is input into the weld penetration status identification model to perform weld penetration status identification and obtain a weld penetration status identification result. Compared with the prior art, this method significantly improves the recognition accuracy of deep penetration argon arc welding weld penetration status by fusing multimodal data of welding current parameters, welding speed parameters, and weld pool images, combined with the multi-scale dynamic weighted fusion mechanism of the improved YOLO X-model. This method effectively monitors and controls the welding process, ensuring the stability and reliability of welding production.

[0081] The above-described embodiments merely represent several preferred implementations of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art could make several improvements and substitutions without departing from the technical principles of the present invention, and these improvements and substitutions should also be considered within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be based on the scope of protection of the claims.

Claims

1. A welding penetration state recognition method based on an improved YOLO X model, characterized in that: The method comprises the following steps: Real-time acquisition of a molten pool image sequence and welding parameter data during deep penetration argon arc welding to form an original multimodal welding data set; the welding parameter data includes at least welding current data and welding speed data; Preprocessing the original multimodal welding data set to obtain standardized multimodal welding data; Constructing a penetration state recognition model based on the improved YOLO X model; wherein the penetration state recognition model includes an image feature extraction module, a welding parameter feature embedding module, a multimodal feature fusion module, and a penetration state classification module connected in sequence; The standardized multimodal welding data is input into the penetration state recognition model to perform welding penetration state recognition to obtain a welding penetration state recognition result.

2. The welding penetration state recognition method based on the improved YOLO X model according to claim 1, characterized in that: The step of preprocessing the original multimodal welding data set to obtain standardized multimodal welding data includes: performing denoising processing on the melt pool image sequence to generate a standardized melt pool image; Normalizing the welding current data and the welding speed data to generate a welding parameter vector; The standardized molten pool image and the welding parameter vector are integrated with corresponding time stamps to obtain standardized multimodal welding data.

3. The welding penetration state recognition method based on the improved YOLO X model according to claim 2, characterized in that: The step of inputting the standardized multimodal welding data into the penetration state recognition model to perform welding penetration state recognition and obtain a welding penetration state recognition result comprises: The standardized molten pool image is input into the image feature extraction module for multi-level convolution processing to extract a multi-scale feature image of the molten pool containing geometric features of the molten pool edge, local morphological features of the small holes, and thermal distribution information; The welding parameter vector is mapped to a high-dimensional feature space by the welding parameter feature embedding module, and the welding parameter vector is spatially aligned with the molten pool multi-scale feature image to obtain a welding parameter feature; the number of output channels of the welding parameter feature is the same as the number of channels of the molten pool multi-scale feature image; The multi-modal feature fusion module splices the molten pool multi-scale feature image and the welding parameter feature in the channel dimension to obtain a multi-modal fusion feature map; Calculating a feature weight coefficient for each pixel position in the multimodal fusion feature map using a self-attention mechanism, and performing pixel-level weighting on the multimodal fusion feature using the feature weight coefficient to obtain an attention-optimized fusion feature; The attention optimization fusion feature is input into the penetration state classification module to perform penetration state classification and identification to obtain a welding penetration state identification result; the welding penetration state identification result includes an incomplete penetration state, a normal penetration state and an over-penetration state.

4. The method for identifying welding penetration status based on the improved YOLO X model according to claim 3, wherein: The image feature extraction module includes an input layer, a focusing layer, a convolution-batch normalization-activation function combination layer, a multi-level local residual block body and a global-local residual block body, which are connected in sequence.

5. The method for identifying welding penetration status based on the improved YOLO X model according to claim 4, characterized in that: The multi-level local residual block body adopts a multi-level serial architecture, and each level of the local residual block body includes a convolution-batch normalization-activation function combination layer and a cross-stage local connection layer connected in sequence, and the number of output channels between the convolution-batch normalization-activation function combination layer and the cross-stage local connection layer of the local residual block body at different levels increases with the increasing level; The global-local residual block includes a convolution-batch normalization-activation function combination layer, a spatial pyramid pooling layer, and a cross-stage local connection layer connected in series, wherein the spatial pyramid pooling layer is composed of a multi-scale maximum pooling layer and a convolution-batch normalization-activation function combination layer connected in series; the multi-scale maximum pooling layer includes multiple parallel pooling branches with different pooling sizes.

6. The method for identifying welding penetration status based on the improved YOLO X model according to claim 4, wherein: The step of inputting the standardized molten pool image into the image feature extraction module for multi-level convolution processing to extract a multi-scale feature image of the molten pool containing geometric features of the molten pool edge, local morphological features of the small holes, and thermal distribution information comprises: Using the focus layer to perform a slicing and reorganization operation on the standardized melt pool image to obtain spatial detail information of the melt pool area; The multi-level local residual block body gradually extracts the geometric features of the edge of the molten pool and the local morphological features of the small holes from the spatial detail information of the molten pool area, thereby obtaining the shallow residual features of the molten pool at different levels; Inputting the shallow residual features of the melt pool into the main body of the global-local residual block to perform pooling operations at different scales to obtain multi-scale pooling features; The multi-scale pooling feature is combined with the shallow residual feature of the melt pool to obtain the deep feature of the melt pool; The deep features of the melt pool are sequentially upsampled, and the deep features of the melt pool after each upsampling are fused with the shallow residual features of the melt pool at different levels through cross-level splicing operations to generate a multi-scale feature image of the melt pool.

7. The method for identifying welding penetration status based on the improved YOLO X model according to claim 3, wherein: The welding parameter feature embedding module includes a parameter embedding layer and a dimension expansion layer connected in series, wherein the parameter embedding layer is composed of a fully connected layer; The fully connected layer is used to map the welding parameter vector to a high-dimensional feature space with the same number of channels as the molten pool multi-scale feature image, to obtain the mapped welding parameter features; The dimension expansion layer is used to spatially expand the mapped welding parameter features so that the spatial size of the mapped welding parameter features is consistent with the multi-scale feature image of the molten pool, thereby obtaining the welding parameter features.

8. The method for identifying welding penetration status based on the improved YOLO X model according to claim 1, wherein: The multimodal feature fusion module includes a channel splicing layer, a self-attention mechanism, a cross-stage local connection and splicing layer, a downsampling layer and a dynamic weight adjustment layer.

9. The method for identifying welding penetration status based on the improved YOLO X model according to claim 8, characterized in that: The self-attention mechanism consists of a convolutional layer and a Sigmoid activation function.

10. The method for identifying welding penetration status based on the improved YOLO X model according to claim 1, characterized in that: The melt penetration state classification module includes a YOLO detection head, and the YOLO detection head includes a convolutional layer and a classifier.