A MIG welding seam condition monitoring system and method based on audio-visual dual mode
Through the audio-visual dual-modal MIG weld state monitoring system, combined with visual and sound data processing, and using the CNN network and attention mechanism, the problem that a single modal feature cannot meet high-precision recognition requirements is solved, and efficient and accurate recognition of the weld penetration state is achieved.
Patent Information
- Application Number
- CN202310575668.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-19
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-05-19
AI Technical Summary
In the existing technology, the single modal feature in weld penetration status monitoring cannot meet the high-precision recognition requirements.
A MIG weld state monitoring system based on audio-visual dual modality is adopted. Data is collected synchronously through the visual data acquisition unit and the sound data acquisition unit. The audio-visual data processing module is combined to perform noise reduction, modality unification and data standardization. The CNN network and attention mechanism are used to extract visual and sound features, and the weld penetration state is recognized through feature fusion.
It improves the accuracy and efficiency of weld penetration state identification, can achieve high-quality feature extraction and identification under complex working conditions, and adapts to welding quality monitoring under different working conditions.
Smart Images

Figure CN116604151B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent welding, and in particular to a system and method for monitoring the weld state of a MIG weld based on visual and audio dual modes. Background Art
[0002] In recent years, with the emergence of intelligent manufacturing and the rapid development of technologies such as artificial intelligence, the Industrial Internet, and sensors, more and more industrial sites have begun using intelligent systems to monitor production status during the production process, thereby preventing product quality problems from causing significant damage to the economy and society. Traditional welding quality inspection methods rely on post-weld testing, which involves performing density tests, magnetic particle inspections, ultrasonic inspections, and radiographic inspections on welds. However, these traditional post-weld inspection methods suffer from complex processes, high costs, and poor real-time performance, making them difficult to meet the current requirements of intelligent manufacturing for automated quality inspection.
[0003] Welding is a complex process involving electricity, light, sound, and heat. The enormous heat generated between the electrodes during welding can melt the metals, thereby connecting them. The connection between the metals is achieved through a solid-liquid-solid process. Melt inert gas arc welding (MIG) is a commonly used welding method for aluminum welding today. It uses a fusible welding wire as the electrode, an inert gas (Ar or He) as the shielding gas, and an arc between the welding wire and the workpiece as the heat source. During the welding process, the shielding gas is piped to the weld to isolate it from the outside air. The welding wire melts into molten droplets and enters the weld pool, where it cools to form a weld. Depending on the working conditions, the weld penetration state produced during the welding process is mainly divided into four categories: partial penetration, complete penetration, overpenetration, and weld-through. The weld penetration state has always been used as one of the most important indicators of welding process quality.
[0004] Over the years, many researchers have conducted in-depth research in the field of weld penetration monitoring. Traditional research often focuses on analyzing the characteristics of a single mode during the welding process. Wu Di et al., in their Chinese invention patent application "A method for determining weld penetration based on back-side keyhole features" (application number CN201610119705.4), effectively identified the weld penetration state by using keyhole images on the back side of the molten pool and an extreme learning machine model. Gao Yanfeng et al., in their Chinese invention patent application "A method for assisting welders in online judgment of weld penetration using arc sound" (CN202210291248.2), studied the frequency domain of welding arc sound through regression analysis and deep learning, identified the weld penetration state, and fed back the weld state to the welder through different vibration forms. The above studies mainly focused on single modalities such as electrical signals, acoustic signals, or visual signals during the welding process, and achieved certain results by applying machine learning and deep learning methods. However, in complex industrial production sites, single modalities are often interfered with by multiple factors, the signal-to-noise ratio is reduced, and the characteristic information related to the weld penetration state is also reduced. The final recognition effect is often not optimal.
[0005] Therefore, technicians in this field are committed to developing a new weld penetration state identification system and method to solve the problem in the existing technology that a single modal feature in weld penetration state monitoring cannot meet the high-precision identification requirements. Summary of the Invention
[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is how to overcome the problem that the weld penetration status monitoring cannot meet the high-precision recognition requirements due to the single modal characteristics.
[0007] To achieve the above objectives, the present invention provides a MIG welding seam condition monitoring system based on visual and audio dual modes, comprising:
[0008] An audio-visual data acquisition module, comprising a visual data acquisition unit, a sound data acquisition unit, and a synchronous acquisition and communication unit, for acquiring data generated during the welding process to obtain audio-visual data;
[0009] an audio-visual data processing module, the audio-visual data processing module being connected to the audio-visual data acquisition module and comprising a noise reduction unit, a modality unification unit, and a data standardization unit, for performing noise reduction, modality unification, and data standardization on the acquired audio-visual data, respectively;
[0010] A weld penetration state recognition module is connected to the audio-visual data processing module, and includes a visual feature extraction unit, a sound feature extraction unit, an audio-visual feature fusion unit, and a weld penetration state classification unit. The processed audio-visual data is subjected to feature extraction by the visual feature extraction unit and the sound feature extraction unit, and feature fusion is performed using the audio-visual feature fusion unit, and then enters the weld penetration state classification unit to realize the recognition of the weld penetration state.
[0011] Furthermore, in the audio-visual data acquisition module,
[0012] The visual data acquisition unit is a welding machine that communicates with an industrial CCD camera through the GigE Vision communication protocol. The industrial CCD camera is configured and data streams are transmitted through the GVCP protocol and the GVSP protocol at the communication protocol application layer, and a filter with a specific wavelength is added in front of the lens of the industrial CCD camera.
[0013] The sound data acquisition unit is a NI DAQExpress data acquisition module that is compatible with the sound pressure sensor and is controlled by LabVIEW image programming to perform real-time acquisition and storage of sound signals. The sound signals are collected by a microphone and then enter a signal conditioner to be converted into a CSV file of a sound pressure signal and stored in a specified file.
[0014] The synchronous acquisition communication unit uses the TCP transmission protocol to achieve synchronous communication between the welder and the data acquisition device, and controls the start and end of the operation of the industrial CCD camera and the sound pressure sensor by monitoring the presence or absence of the welding current of the welder.
[0015] Furthermore, in the audio-visual data processing module,
[0016] The noise reduction unit uses a median filter and an adaptive noise complete set empirical mode decomposition-wavelet packet threshold noise reduction method to remove noise from the sound signal;
[0017] The unified modality unit obtains a time-frequency diagram of a time series through short-time Fourier transform, and merges the sound pressure signal of the one-dimensional time series and the visual signal of the two-dimensional grayscale image into the same modality for subsequent processing.
[0018] Furthermore, in the weld penetration state identification module,
[0019] The visual feature extraction unit includes a CNN network and a spatial attention mechanism; the CNN network includes several input layers, convolution layers, activation functions, pooling layers and fully connected layers; the CNN network updates weights through backpropagation; the spatial attention mechanism performs channel attention and spatial attention processing on the feature layer of the CNN network;
[0020] The sound feature extraction unit includes the CNN network and a multi-scale convolution kernel combination convolution layer; the multi-scale convolution kernel combination convolution layer extracts more time domain and frequency domain feature information of the sound signal by adjusting the length and width of the convolution kernel and combining them; the multi-scale convolution kernel combination convolution layer is arranged after the input layer;
[0021] The audio-visual feature fusion unit flattens the deep visual features extracted by the visual feature extraction unit and the deep sound features extracted by the sound feature extraction unit into N*1 and M*1 feature vectors, respectively, and concatenates the feature vectors to form a (N+M)*1 fused feature vector;
[0022] The weld penetration state classification unit includes two fully connected layers, a softmax layer and an output layer; after passing through the two fully connected layers, the dimension of the fused feature vector is reduced to the number of categories to be classified, and then converted into the probability of the corresponding category by the softmax layer, and finally the weld penetration state is output by the output layer.
[0023] The present invention also provides a method for monitoring the weld state of a MIG weld based on audio-visual dual modes, the method comprising the following steps:
[0024] Step 1: Install a synchronous audio-visual data acquisition system to collect data and obtain audio-visual data;
[0025] Step 2: Data preprocessing: performing noise reduction, modality unification and data standardization on the collected audio-visual data;
[0026] Step 3: Build a weld penetration state recognition model based on audio-visual dual modalities;
[0027] Step 4: training the weld penetration state recognition model, adjusting and optimizing the network structure;
[0028] Step 5: Deploy the weld penetration state recognition model to the industrial site to realize real-time detection.
[0029] Furthermore, the step 1 includes the following sub-steps:
[0030] Step 1.1. Fix the CCD camera and sound pressure sensor to the welding machine arm, facing the contact point between the welding wire and the workpiece, ensuring that there is no obstruction in the middle and that a complete image of the molten pool can be captured;
[0031] Step 1.2, setting the sampling frequencies of the CCD camera and the sound pressure sensor, setting the sampling frequency of the CCD camera to 50 frames per second, setting the sampling frequency of the sound pressure sensor to 50 kHz, and starting the LabVIEW program for acquisition start control via TCP communication;
[0032] Step 1.3, setting the working parameters of the welding machine, including: arc starting current, welding current, and wire feeding speed; starting the welding machine to start working, the CCD camera and the sound pressure sensor to start working, collecting the molten pool image and sound pressure signal during the welding process, obtaining the audio-visual data and storing it.
[0033] Furthermore, step 2 includes the following sub-steps:
[0034] Step 2.1: Perform median filtering on the collected molten pool image to remove excess image noise and improve the signal-to-noise ratio:
[0035] g(x,y)=Med{f(a i ,b j ),(a i ,b j )∈A}
[0036] Among them, f, g are the values of the pixels before and after median filtering, and A is the filtering window;
[0037] Step 2.2, performing adaptive noise complete set empirical mode decomposition-wavelet packet threshold denoising on the collected sound pressure signal to remove noise from the sound signal;
[0038] Step 2.3: Determine the length L of each sound frame according to the sampling frequency settings of the CCD camera and the sound pressure sensor:
[0039]
[0040] Among them, f a , f i are the sampling frequencies of the sound pressure sensor and the CCD camera respectively;
[0041] Step 2.4: Convert the discrete digital sound pressure signal into a two-dimensional time-frequency spectrum through STFT:
[0042]
[0043] Wherein, a[n] is the sound pressure signal, and w[n] is the Hamming window function;
[0044] Step 2.5: resize the obtained molten pool image and the two-dimensional time-frequency spectrum image to an image of (100, 100, 3) to complete the preprocessing of the audio-visual data.
[0045] Furthermore, the step 2.2 includes the following sub-steps:
[0046] Step 2.2.1, decompose the sound pressure signal into multiple modal components by using adaptive noise complete set empirical mode decomposition, and record the sound pressure signal as y(t);
[0047] Step 2.2.2, Gaussian white noise v i (t) is added to the sound pressure signal y(t) to obtain a new signal No. 1:
[0048] y(t)+(-1) s εv i (t)
[0049] Where s = 1, 2; ε is the standard table of white noise, i = 1, 2, ... N, N is the number of components obtained by decomposition;
[0050] Perform empirical mode decomposition (EMD) on the new signal No. 1 to obtain the first-order intrinsic mode function (IMF) C1:
[0051]
[0052] Among them, r i ()for The residual signal outside
[0053] The obtained N IMFs are averaged to obtain the first IMF component for performing the adaptive noise complete set empirical mode decomposition.
[0054]
[0055] The residual signal after removing the first IMF component is r1(t):
[0056]
[0057] Add positive and negative paired Gaussian white noise to the residual signal r1(t) to obtain a new signal No. 2, perform EMD decomposition on the new signal No. 2 to obtain the first-order modal component D1, and then obtain the second IMF component obtained by performing the adaptive noise complete set empirical mode decomposition And the residual signal r2(t) after removing the second IMF component:
[0058]
[0059]
[0060] Repeat the above operation until the residual signal r is obtained K+1 () Monotonous, cannot be further decomposed;
[0061] At this time, K IMF components are obtained, and the sound pressure signal y(t) is decomposed into:
[0062]
[0063] At this point, the process of the adaptive noise complete set empirical mode decomposition is completed;
[0064] Step 2.2.3, comparing the correlation coefficient or variance contribution rate of each of the IMF components and the sound pressure signal y(t), and screening out the noisy components;
[0065] The correlation coefficient is:
[0066]
[0067] Among them, x i are the IMF components, and y is the sound pressure signal;
[0068] The variance contribution rate is:
[0069]
[0070] Where D(j) is the variance of the IMF component:
[0071]
[0072] Step 2.2.4: Perform wavelet packet threshold denoising; define subspace is the function U n The closure space of (), is the function U 2n () closure space, let U n ()satisfy:
[0073]
[0074]
[0075] Where h(k) and g(k) are low-pass and high-pass filters of length 2N, and g(k) = (-1) k h(k-1), the sequence constructed by the above formula is Deterministic orthogonal wavelet packets;
[0076] set up It can be expressed as:
[0077]
[0078] Get the wavelet packet decomposition result:
[0079]
[0080] The threshold method is set to:
[0081]
[0082] Reconstruct the signal after threshold denoising:
[0083]
[0084] At this point, the noise reduction of the sound pressure signal y(t) is completed.
[0085] Furthermore, step 3 includes the following sub-steps:
[0086] Step 3.1: The weld penetration state recognition model based on audio-visual dual modality includes a visual feature extraction convolutional neural network; the visual feature extraction convolutional neural network includes: four convolutional layers, four maximum pooling layers and two fully connected layers, and the activation function is the ReLU function:
[0087] f(x)=max(0,x)
[0088] Among them, the number of convolution kernels in the convolution layer is 8, 16, 32, and 64 respectively, the size of the convolution kernel is 5*5, and the step size is 1; the pooling kernel size of the maximum pooling layer is 3*3, and the step size is 2; the number of nodes in the fully connected layer is 3136 and 256, and the output feature size of the CNN network is 1*256;
[0089] Step 3.2, adding a visual attention mechanism to the visual feature extraction convolutional neural network, including a channel attention mechanism and a spatial attention mechanism;
[0090] The channel attention mechanism performs global average pooling and global maximum pooling on the input features, inputs the results of the global average pooling and the global maximum pooling into a shared network consisting of a multi-layer perceptron (MLP) and a hidden layer for processing, adds the two processed results and uses the sigmoid function to obtain the channel attention weight of the input feature, and multiplies the obtained channel attention weight by the input feature, that is:
[0091]
[0092] Wherein, Mc is the channel attention weight, F is the input feature, σ is the sigmoid function, W0 and W1 are the weights of the MLP, is the result of the global average pooling and the global maximum pooling;
[0093] The spatial attention mechanism calculates the maximum and average values on the channel of each input feature, stacks the convolutional layer with a channel number of 1, applies the sigmoid function to the result, and obtains the weight corresponding to each input feature. The corresponding weight obtained is multiplied by the input feature:
[0094]
[0095] Among them, Ms is the spatial attention weight, F is the input feature, σ is the sigmoid function, f 7×7 is a 7×7 convolution operation, is the result of spatial average pooling and maximum pooling;
[0096] Step 3.3: The audio-visual dual-modality weld penetration state recognition model also includes a sound feature extraction convolutional neural network. The structure of the sound feature extraction convolutional neural network is generally consistent with that of the visual feature extraction convolutional neural network. The sound feature extraction convolutional neural network uses a multi-scale convolution kernel combination convolution layer to replace the 5*5 convolution kernel in the network, and performs appropriate padding before convolution. By comparing the classification accuracy obtained from different combinations, the parallel channels corresponding to the 3*8, 4*6, 5*5, 6*4, and 8*3 convolution kernels are optimized.
[0097] Step 3.4: Perform feature fusion on the audiovisual features, concatenate the 256-dimensional feature vector extracted by the visual feature extraction convolutional neural network and the 256-dimensional feature vector extracted by the sound feature extraction convolutional neural network to obtain a 512-dimensional fused feature vector, and then input the obtained 512-dimensional fused feature vector into the classification unit, perform two full-connection operations to convert it into a 4-dimensional feature vector, and enter the softmax layer to complete the final weld penetration state classification task.
[0098] Furthermore, in step 4, the weld penetration state recognition model is trained, including setting the corresponding loss function, optimizer, and learning rate, batch size, epoch, dropout ratio, and num_classes parameters.
[0099] The present invention provides a MIG welding seam condition monitoring system and method based on audio-visual dual-mode, which has at least the following technical effects:
[0100] 1. In traditional machine learning, feature engineering is often complicated for large-scale data sets. Feature extraction and feature selection rely heavily on prior knowledge and need to be performed manually. In addition, the generalization performance of machine learning algorithms is poor and cannot adapt to the task of identifying the weld penetration state under different working conditions. The technical solution provided by the present invention can efficiently extract weld features by building a deep learning network, using CNN, attention mechanism, multi-scale convolution kernel combination convolution and other methods. At the same time, on this basis, the effect of the model under other working conditions can be improved through transfer learning, and it has strong generalization performance;
[0101] 2. In deep learning of a single modality, due to the harsh factory environment, the signal-to-noise ratio of the collected single modality data is often not very high, the data quality is low, and the recognition effect is often less than ideal. However, the technical solution provided by the present invention proposes deep learning of visual and auditory modal fusion. By synchronously collecting visual and auditory information and processing them accordingly, it can achieve the best advantages and minimize the shortcomings when extracting features, greatly improving the quality of feature quantities, and is more conducive to improving recognition effects. It can integrate multi-source information of the welding process and extract high-quality modal features, achieving better recognition of weld penetration status, greatly improving the efficiency and accuracy of welding quality monitoring.
[0102] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] Figure 1 This is a flow chart of identifying the weld penetration state according to a preferred embodiment of the present invention;
[0104] Figure 2 yes Figure 1 Flowchart of CEEMDAN-wavelet packet threshold joint denoising in the illustrated embodiment;
[0105] Figure 3 yes Figure 1 Schematic diagram of a weld penetration state identification model according to the embodiment shown;
[0106] Figure 4 yes Figure 1 Schematic diagram of the channel attention mechanism and spatial attention mechanism of CBAM in the illustrated embodiment;
[0107] Figure 5 1 is a schematic diagram of the multi-scale convolution kernel combination convolution of the embodiment shown in FIG. DETAILED DESCRIPTION
[0108] The following describes several preferred embodiments of the present invention with reference to the accompanying drawings to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0109] In response to the problem that the single modal feature in the existing weld penetration status monitoring cannot meet the needs of high-precision identification, the present invention provides a practical and effective identification method to monitor the quality of the welding process. By constructing an end-to-end weld penetration status identification model, it can effectively realize the accurate identification of the weld penetration status and achieve cost reduction and efficiency improvement of welding quality detection.
[0110] In view of the coexistence of multiple modal information, the concept of Multi Modal Deep Learning has been proposed in recent years and has achieved certain success in the fields of natural language processing, audio and video recognition, etc. It has now been applied to the fields of remote sensing images, medical imaging, human posture recognition, etc. Therefore, in an embodiment of the present invention, the concept of multimodal fusion is used to solve the above-mentioned problem of identifying the penetration state of the weld. The audio-visual data acquisition module collects synchronized visual and auditory data and fuses them at the feature level, which greatly enriches the feature quantity. In addition, the attention mechanism is introduced when extracting the visual information features of MIG welding, so as to realize the accurate extraction of the weld part features in the visual image. In the extraction of the auditory information features of MIG welding, the time-frequency spectrum obtained by STFT is subjected to multi-scale convolution kernel combination convolution to extract features, so as to realize richer feature extraction of the sound signal.
[0111] Example 1
[0112] An embodiment of the present invention provides a MIG welding weld state monitoring system based on audio-visual dual modality, including an audio-visual data acquisition module, an audio-visual data processing module and a weld penetration state recognition module. The audio-visual data acquisition module collects the data generated during the welding process in real time and synchronously, mainly collecting molten pool images and arc sound signals. The audio-visual data processing module performs corresponding processing on the collected raw data, including noise reduction, modality unification and data standardization. The weld penetration state recognition module uses convolutional neural networks (CNN), attention mechanisms and multi-scale convolution kernel combined convolution methods to extract features from the processed audio-visual data respectively, fuses the audio-visual features in the feature fusion layer, and finally enters the classification network to achieve effective recognition of the weld penetration state.
[0113] in,
[0114] The audio-visual data acquisition module includes a visual data acquisition unit, a sound data acquisition unit and a synchronous acquisition and communication unit, which collects the data generated during the welding process to obtain audio-visual data;
[0115] The audio-visual data processing module is connected to the audio-visual data acquisition module, and includes a noise reduction unit, a modality unification unit, and a data standardization unit, which respectively perform noise reduction, modality unification, and data standardization on the collected audio-visual data;
[0116] The weld penetration state recognition module is connected to the audio-visual data processing module, and includes a visual feature extraction unit, a sound feature extraction unit, an audio-visual feature fusion unit, and a weld penetration state classification unit. The processed audio-visual data is extracted by the visual feature extraction unit and the sound feature extraction unit, and the feature fusion is performed using the audio-visual feature fusion unit, and then enters the weld penetration state classification unit to realize the recognition of the weld penetration state.
[0117] Example 2
[0118] Based on Example 1, the visual data acquisition unit: During the visual data acquisition process, the welding machine host and industrial CCD camera communicate using the GigE Vision communication protocol. At the communication protocol application layer, the GVCP and GVSP protocols are used to configure the CCD camera and transmit data streams, respectively. Furthermore, a specific wavelength filter is placed in front of the lens to reduce the impact of the intense arc light during the welding process on the weld pool image.
[0119] Sound data acquisition unit: During the sound data acquisition process, LabVIEW graphical programming is used to control the NI DAQExpress data acquisition module that is paired with the sound pressure sensor to collect and store sound signals in real time. The sound signal is collected by the microphone and then enters the signal conditioner to be converted into a CSV file of the sound pressure signal and stored in a specified file.
[0120] Synchronous acquisition communication unit: TCP transmission protocol is used to realize synchronous communication between the welding machine and the data acquisition equipment, and the start and end of the operation of the industrial CCD camera and the sound pressure sensor are controlled by monitoring the presence or absence of the welding current of the welding machine.
[0121] The main function of audio-visual data processing is to process the collected raw data to obtain standard, high-quality data to be input into the weld penetration state recognition module for final recognition.
[0122] Noise reduction unit: Welding is a complex process involving electricity, light, sound, and heat, and noise is inevitable in this process. Visual information noise is mainly caused by arc light or metal spatter. The embodiment of the present invention uses a nonlinear filter called median filtering to remove noise points while protecting the edge information of the image. The noise of the sound signal is mainly caused by the operation of surrounding machines, molten pool oscillation, etc. The embodiment of the present invention uses a Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN)-wavelet packet threshold denoising method (such as Figure 2 The noise reduction module can effectively remove the noise from the sound signal. The original data can be greatly improved by the noise reduction module, thereby improving the final recognition accuracy.
[0123] Unified modal unit: Since the sound pressure signal is a one-dimensional time series, and the visual signal is a two-dimensional grayscale image, in order to eliminate cross-modal differences, the embodiment of the present invention obtains the time-frequency diagram of the time series through Short Time Fourier Transform (STFT), and then merges the two modes into the same mode for subsequent processing. At the same time, since the sound has short-term stability, the sound signal can be considered to be approximately unchanged within 10 to 30ms, so a short-time analysis method is adopted to analyze the sound signal. By formulating audio-visual data matching rules and determining the number of sound pressure samples corresponding to each picture according to the sampling frequency of the visual image, a one-to-one correspondence between the visual image and the sound pressure data segment can be achieved.
[0124] Data Normalization: Considering that images captured by the data acquisition device and time-frequency spectrograms obtained after STFT may be too large, slowing down training speed, or have inconsistent channel counts, the visual and time-frequency images are standardized. The images are resized and converted to standard RGB format before being input into the weld penetration recognition module. After data normalization, the visual and audio images can be used as input to the deep learning network for subsequent weld penetration recognition.
[0125] In the weld penetration state recognition module, the visual feature extraction unit: the visual feature extraction unit is composed of CNN and a spatial attention mechanism. CNN is composed of several input layers, convolutional layers, activation functions, pooling layers and fully connected layers. CNN performs weight updates through back propagation, and has been verified to be able to achieve good high-dimensional feature extraction effects on two-dimensional images. On this basis, since the image of the molten pool in the visual image input to the weld penetration state recognition module only occupies a part of the overall image, other information in the image, including the background, the unwelded part of the weld edge, etc., does not contribute much to the final penetration state recognition. However, during learning, they consume the same computing resources and have the same weights as the effective molten pool information, which often leads to reduced recognition efficiency and accuracy. In CNN, the introduction of the attention mechanism enables the network to pay attention to the molten pool area with richer information. In an embodiment of the present invention, a convolutional block attention module (CBAM) is added to the CNN network to perform channel attention and spatial attention processing on the input feature layer (such as Figure 4 shown).
[0126] Sound feature extraction unit: The sound feature extraction unit is composed of CNN and multi-scale convolution kernel combination convolution. Different from the actual molten pool image actually captured by the CCD camera, the time-frequency spectrum obtained by STFT represents the change trend in the time domain and frequency domain on the horizontal axis and vertical axis respectively. The use of conventional square convolution kernels often misses a lot of information in the time and frequency domains. The multi-scale convolution kernel combination convolution proposed in the embodiment of the present invention can extract richer time and frequency domain feature information of the sound signal by adjusting the length and width of the convolution kernel and combining them in parallel. The multi-scale convolution kernel combination convolution layer is arranged after the input layer.
[0127] Audio-visual feature fusion unit: The deep visual features extracted by the visual feature extraction unit and the deep sound features extracted by the sound feature extraction unit are flattened into N*1 and M*1 feature vectors, and the two feature vectors are connected to form a (N+M)*1 fusion feature vector, which is sent to the weld penetration state classification unit for classification.
[0128] Weld penetration classification unit: This unit consists of two fully connected layers, a softmax layer, and an output layer. The high-dimensional fusion features are reduced to the number of categories through the two fully connected layers. The softmax layer then converts them into probabilities for the corresponding categories, ultimately outputting the weld penetration status.
[0129] Example 3
[0130] This embodiment of the present invention provides a method for monitoring the weld state of a MIG weld based on audio-visual dual-modality. First, an audio-visual data acquisition device is constructed. A CCD camera and an acoustic pressure sensor are installed diagonally above the weld pool, moving with the welding arm, maintaining their relative positions. The collected raw data is then fed into an audio-visual data processing module using database technology for noise reduction, modality unification, and data standardization. The raw weld pool image and acoustic pressure data are processed into images of standard size and number of channels. The audio-visual data is labeled with four penetration states, and the training and test sets are divided into training and test sets in an 8:2 ratio. Finally, the data is fed into a weld penetration state recognition module for training. The network structure is adjusted and optimized. The final weld penetration state recognition result is obtained through feature extraction using a deep learning network and classification by a classification unit. Based on this, the system and method for MIG weld penetration state recognition based on audio-visual dual-modality deep learning is deployed in industrial sites to enable real-time monitoring of audio-visual data streams.
[0131] Specifically include the following steps, such as Figure 1 As shown:
[0132] Step 1: Install a synchronous audio-visual data acquisition system to collect data and obtain audio-visual data;
[0133] Step 2: Data preprocessing: noise reduction, modality unification and data standardization of the collected audio-visual data;
[0134] Step 3: Build a weld penetration state recognition model based on audio-visual dual modalities;
[0135] Step 4: Train the weld penetration state recognition model and adjust and optimize the network structure;
[0136] Step 5: Deploy the weld penetration state recognition model to the industrial site to realize real-time detection.
[0137] Example 4
[0138] Based on Example 3, step 1 includes the following sub-steps:
[0139] Step 1.1. Fix the CCD camera and sound pressure sensor to the welding machine arm, facing the contact point between the welding wire and the workpiece, ensuring that there is no obstruction in the middle and that a complete image of the molten pool can be captured;
[0140] Step 1.2: Set the sampling frequencies of the CCD camera and the sound pressure sensor to 50 frames per second and 50 kHz, respectively. Start the LabVIEW program that controls the acquisition start via TCP communication.
[0141] Step 1.3: Set the welding machine's operating parameters, including arc starting current, welding current, and wire feed speed; start the welding machine, and the CCD camera and sound pressure sensor start working to collect the molten pool image and sound pressure signal during the welding process, obtain audio-visual data, and store it.
[0142] Example 5
[0143] Based on Example 4, step 2 includes the following sub-steps:
[0144] Step 2.1: Perform median filtering on the collected melt pool image to remove excess image noise and improve the signal-to-noise ratio:
[0145] g(x,y)=Med{f(a i ,b j ),(a i ,b j )∈A}
[0146] Among them, f, g are the values of the pixels before and after median filtering, and A is the filtering window;
[0147] Step 2.2, performing CEEMDAN-wavelet packet threshold denoising on the collected sound pressure signal to remove noise from the sound signal;
[0148] Step 2.3. Determine the length L of each sound frame based on the sampling frequency settings of the CCD camera and sound pressure sensor:
[0149]
[0150] Among them, f a , f i are the sampling frequencies of the sound pressure sensor and CCD camera, respectively;
[0151] Step 2.4: Convert a discrete digital sound pressure signal into a two-dimensional time-frequency spectrum through STFT:
[0152]
[0153] Where a[n] is the sound pressure signal, w[n] is the Hamming window function;
[0154] Step 2.5: Resize the obtained melt pool image and two-dimensional time-frequency spectrum to an image of (100, 100, 3) to complete the preprocessing of the audio-visual data.
[0155] Wherein, step 2.2 includes the following sub-steps (such as Figure 2 shown):
[0156] Step 2.2.1. Decompose the sound pressure signal into multiple modal components using adaptive noise complete set empirical mode decomposition, and record the sound pressure signal as y(t);
[0157] Step 2.2.2, Gaussian white noise v i (t) is added to the sound pressure signal y(t) to obtain the new signal No. 1:
[0158] y(t)+(-1) s εv i (t)
[0159] Where s = 1, 2; ε is the standard table of white noise, i = 1, 2, ... N, N is the number of components obtained by decomposition;
[0160] Perform empirical mode decomposition (EMD) on the new signal No. 1 and obtain the first-order intrinsic mode function (IMF) C1:
[0161]
[0162] Among them, r i ()for The residual signal outside
[0163] Average the N IMFs obtained to obtain the first IMF component for adaptive noise complete set empirical mode decomposition
[0164]
[0165] The residual signal after removing the first IMF component is r1(t):
[0166]
[0167] Add positive and negative paired Gaussian white noise to the residual signal r1(t) to obtain the new signal No. 2, perform EMD decomposition on the new signal No. 2, and obtain the first-order modal component D1, and then obtain the second IMF component obtained by performing adaptive noise complete set empirical mode decomposition And the residual signal r2(t) after removing the second IMF component:
[0168]
[0169]
[0170] Repeat the above operation until the residual signal r is obtained K+1 () Monotonous, cannot be further decomposed;
[0171] At this time, K IMF components are obtained, and the sound pressure signal y(t) is decomposed into:
[0172]
[0173] At this point, the process of adaptive noise complete set empirical mode decomposition is completed;
[0174] Step 2.2.3, compare the correlation coefficient or variance contribution rate of each IMF component with the sound pressure signal y(t) to filter out the noisy components;
[0175] The correlation coefficient is:
[0176]
[0177] Among them, x i are the IMF components, y is the sound pressure signal;
[0178] The variance contribution rate is:
[0179]
[0180] Where D(j) is the variance of the IMF component:
[0181]
[0182] Step 2.2.4: Perform wavelet packet threshold denoising; define subspace is the function U n The closure space of (), is the function U 2n () closure space, let U n ()satisfy:
[0183]
[0184]
[0185] Where h(k) and g(k) are low-pass and high-pass filters of length 2N, and g(k) = (-1) k h(k-1), the sequence constructed by the above formula is Deterministic orthogonal wavelet packets;
[0186] set up It can be expressed as:
[0187]
[0188] Get the wavelet packet decomposition result:
[0189]
[0190]
[0191] The threshold method is set to:
[0192]
[0193] Reconstruct the signal after threshold denoising:
[0194]
[0195] At this point, the noise reduction of the sound pressure signal y(t) is completed.
[0196] Example 6
[0197] Based on Example 5, step 3 includes the following sub-steps: Figure 4 As shown:
[0198] Step 3.1: The weld penetration state recognition model based on audio-visual dual modality includes a visual feature extraction convolutional neural network. The visual feature extraction convolutional neural network includes: four convolutional layers, four maximum pooling layers, and two fully connected layers. The activation function is the ReLU function:
[0199] f(x)=max(0,x)
[0200] Among them, the number of convolution kernels in the convolution layer is 8, 16, 32, and 64, respectively. The size of the convolution kernel is 5*5 and the step size is 1. The pooling kernel size of the maximum pooling layer is 3*3 and the step size is 2. The number of nodes in the fully connected layer is 3136 and 256, and the output feature size of the CNN network is 1*256.
[0201] Step 3.2: Add the visual attention mechanism (CBAM) to the visual feature extraction convolutional neural network, including the channel attention mechanism and the spatial attention mechanism.
[0202] The channel attention mechanism performs global average pooling and global maximum pooling on the input features. The results of global average pooling and global maximum pooling are input into a shared network consisting of a multi-layer perceptron (MLP) and a hidden layer for processing. The two processed results are added together and the sigmoid function is used to obtain the channel attention weight of the input feature. The obtained channel attention weight is multiplied by the input feature, that is:
[0203]
[0204] Among them, Mc is the channel attention weight, F is the input feature, σ is the sigmoid function, W0 and W1 are the weights of MLP, is the result of global average pooling and global maximum pooling;
[0205] The spatial attention mechanism calculates the maximum and average values for each input feature channel, stacks the convolutional layer with a channel number of 1, and uses the sigmoid function to obtain the weight corresponding to each input feature. The corresponding weight obtained is multiplied by the input feature:
[0206]
[0207] Among them, Ms is the spatial attention weight, F is the input feature, σ is the sigmoid function, f 7×7 is a 7×7 convolution operation, is the result of spatial average pooling and maximum pooling;
[0208] Step 3.3, the weld penetration state recognition model based on audio-visual dual modality also includes a sound feature extraction convolutional neural network. The structure of the sound feature extraction convolutional neural network is roughly consistent with the visual feature extraction convolutional neural network. The sound feature extraction convolutional neural network uses a multi-scale convolution kernel combined convolution layer to replace the 5*5 convolution kernel in the network, and performs appropriate padding before convolution (such as Figure 5 By comparing the classification accuracy obtained by different combinations, the parallel channels corresponding to the convolution kernels of 3*8, 4*6, 5*5, 6*4, and 8*3 are optimized.
[0209] Step 3.4: Perform feature fusion on the audiovisual features, concatenate the 256-dimensional feature vector extracted by the visual feature extraction convolutional neural network and the 256-dimensional feature vector extracted by the sound feature extraction convolutional neural network to obtain a 512-dimensional fused feature vector, and then input the obtained 512-dimensional fused feature vector into the classification unit, perform two full-connection operations to convert it into a 4-dimensional feature vector, and enter the softmax layer to complete the final weld penetration state classification task.
[0210] In step 4, train the weld penetration state recognition model, including setting the corresponding loss function, optimizer, and learning rate, batch size, epoch, dropout ratio, and num_classes parameters.
[0211] The trained model is deployed to the host computer at the welding site, ultimately achieving real-time monitoring of the audio-visual data stream.
[0212] The preferred embodiments of the present invention have been described in detail above. It should be understood that numerous modifications and variations based on the concepts of the present invention are possible without inventive effort by those skilled in the art. Therefore, any technical solution that can be derived by one skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. A MIG welding seam condition monitoring system based on audio-visual dual mode, characterized in that: include: An audio-visual data acquisition module, comprising a visual data acquisition unit, an audio data acquisition unit, and a synchronous acquisition and communication unit, collects data generated during the welding process to obtain audio-visual data. The visual data acquisition unit enables the welding machine and the industrial CCD camera to communicate via the GigE Vision communication protocol, configures the industrial CCD camera and transmits data streams via the GVCP and GVSP protocols at the communication protocol application layer, and adds a filter of a specific wavelength in front of the lens of the industrial CCD camera. The sound data acquisition unit is a NI DAQExpress data acquisition module that is compatible with the sound pressure sensor and is controlled by LabVIEW image programming to perform real-time acquisition and storage of sound signals. The sound signals are collected by a microphone and then enter a signal conditioner to be converted into a CSV file of a sound pressure signal and stored in a specified file. The synchronous acquisition communication unit uses the TCP transmission protocol to achieve synchronous communication between the welding machine and the data acquisition device, and controls the start and end of the operation of the industrial CCD camera and the sound pressure sensor by monitoring the presence or absence of the welding current of the welding machine; an audio-visual data processing module, the audio-visual data processing module being connected to the audio-visual data acquisition module and comprising a noise reduction unit, a modality unification unit, and a data standardization unit, for performing noise reduction, modality unification, and data standardization on the acquired audio-visual data, respectively; A weld penetration state recognition module, the weld penetration state recognition module is connected to the audio-visual data processing module, and includes a visual feature extraction unit, a sound feature extraction unit, an audio-visual feature fusion unit, and a weld penetration state classification unit. The processed audio-visual data is subjected to feature extraction by the visual feature extraction unit and the sound feature extraction unit, and feature fusion is performed by the audio-visual feature fusion unit before entering the weld penetration state classification unit to realize the recognition of the weld penetration state. The visual feature extraction unit includes a CNN network and a spatial attention mechanism. The CNN network includes several input layers, convolutional layers, activation functions, pooling layers, and fully connected layers. The CNN network updates weights through back propagation. The spatial attention mechanism performs channel attention and spatial attention processing on the feature layer of the CNN network. The sound feature extraction unit includes the CNN network and a multi-scale convolution kernel combination convolution layer; the multi-scale convolution kernel combination convolution layer extracts more time domain and frequency domain feature information of the sound signal by adjusting the length and width of the convolution kernel and combining them; the multi-scale convolution kernel combination convolution layer is arranged after the input layer; The audio-visual feature fusion unit flattens the deep visual features extracted by the visual feature extraction unit and the deep sound features extracted by the sound feature extraction unit into and The feature vectors are connected to form The fusion feature vector of The weld penetration state classification unit includes two fully connected layers, a softmax layer and an output layer; after passing through the two fully connected layers, the dimension of the fused feature vector is reduced to the number of categories to be classified, and then converted into the probability of the corresponding category by the softmax layer, and finally the weld penetration state is output by the output layer.
2. The MIG welding seam condition monitoring system based on audio-visual dual mode according to claim 1 is characterized in that: In the audio-visual data processing module, The noise reduction unit uses a median filter and an adaptive noise complete set empirical mode decomposition-wavelet packet threshold noise reduction method to remove noise from the sound signal; The unified modality unit obtains a time-frequency diagram of a time series through short-time Fourier transform, and merges the sound pressure signal of the one-dimensional time series and the visual signal of the two-dimensional grayscale image into the same modality for subsequent processing.
3. A method for monitoring the weld condition of a MIG weld based on audio-visual dual-mode, using the MIG weld condition monitoring system based on audio-visual dual-mode according to claim 1 or 2, characterized in that: The method comprises the following steps: Step 1: Install a synchronous audio-visual data acquisition system to collect data and obtain audio-visual data; Step 2: Data preprocessing: performing noise reduction, modality unification and data standardization on the collected audio-visual data; Step 3: Build a weld penetration state recognition model based on audio-visual dual modalities; Step 4: training the weld penetration state recognition model, adjusting and optimizing the network structure; Step 5: Deploy the weld penetration state recognition model to the industrial site to realize real-time detection.
4. The MIG welding seam condition monitoring method based on audio-visual dual mode according to claim 3 is characterized in that: The step 1 includes the following sub-steps: Step 1.
1. Fix the CCD camera and sound pressure sensor to the welding machine arm, facing the contact point between the welding wire and the workpiece, ensuring that there is no obstruction in the middle and that a complete image of the molten pool can be captured; Step 1.2, setting the sampling frequencies of the CCD camera and the sound pressure sensor, setting the sampling frequency of the CCD camera to 50 frames per second, setting the sampling frequency of the sound pressure sensor to 50 kHz, and starting the LabVIEW program for acquisition start control via TCP communication; Step 1.3, setting the working parameters of the welding machine, including: arc starting current, welding current, and wire feeding speed; starting the welding machine to start working, the CCD camera and the sound pressure sensor to start working, collecting the molten pool image and sound pressure signal during the welding process, obtaining the audio-visual data and storing it.
5. The MIG welding seam condition monitoring method based on audio-visual dual modality according to claim 4 is characterized in that: Step 2 includes the following sub-steps: Step 2.1: Perform median filtering on the collected molten pool image to remove excess image noise and improve the signal-to-noise ratio: in, is the value of the pixel before and after median filtering, is the filter window; Step 2.2, performing adaptive noise complete set empirical mode decomposition-wavelet packet threshold denoising on the collected sound pressure signal to remove noise from the sound signal; Step 2.3: Determine the length of each sound frame based on the sampling frequency settings of the CCD camera and the sound pressure sensor. : in, are the sampling frequencies of the sound pressure sensor and the CCD camera respectively; Step 2.4: Convert the discrete digital sound pressure signal into a two-dimensional time-frequency spectrum through STFT: in, is the sound pressure signal, is the Hamming window function; Step 2.5: resize the obtained molten pool image and the two-dimensional time-frequency spectrum image to an image of (100, 100, 3) to complete the preprocessing of the audio-visual data.
6. The MIG welding seam condition monitoring method based on audio-visual dual mode according to claim 5 is characterized in that: The step 2.2 includes the following sub-steps: Step 2.2.1: Decompose the sound pressure signal into multiple modal components by using the adaptive noise complete set empirical mode decomposition, and record the sound pressure signal as ; Step 2.2.2: Gaussian white noise Add the sound pressure signal Get new signal No. 1: in, ; is the standard table of white noise, is the number of components finally obtained by decomposition; Perform empirical mode decomposition (EMD) on the new signal No. 1 to obtain the first-order intrinsic mode function (IMF) : in, for The residual signal outside The obtained N IMFs are averaged to obtain the first IMF component for performing the adaptive noise complete set empirical mode decomposition. : The residual signal after removing the first IMF component is : In the residual signal Add positive and negative paired Gaussian white noise to obtain new signal No. 2, perform EMD decomposition on the new signal No. 2, and obtain the first-order modal component , and then obtain the second IMF component obtained by performing the adaptive noise complete set empirical mode decomposition And the residual signal after removing the second IMF component : Repeat the above steps until the residual signal is obtained Monotonous, cannot be further decomposed; At this time, K IMF components are obtained, and the sound pressure signal is broken down into: At this point, the process of the adaptive noise complete set empirical mode decomposition is completed; Step 2.2.3: Compare each of the IMF components with the sound pressure signal The correlation coefficient or variance contribution rate of the noise component is screened out; The correlation coefficient is: in, are the IMF components, is the sound pressure signal; The variance contribution rate is: in, is the variance of the IMF component: Step 2.2.4: Perform wavelet packet threshold denoising; define subspace For function The closure space of For function The closure space of satisfy: in, are low-pass and high-pass filters of length 2N that satisfy , the sequence constructed by the above formula is Deterministic orthogonal wavelet packets; set up It can be expressed as: Get the wavelet packet decomposition result: The threshold method is set to: Reconstruct the signal after threshold denoising: At this point, the sound pressure signal is completed Noise reduction.
7. The MIG welding seam condition monitoring method based on audio-visual dual modality according to claim 6 is characterized in that: Step 3 includes the following sub-steps: Step 3.1: The weld penetration state recognition model based on audio-visual dual modality includes a visual feature extraction convolutional neural network; the visual feature extraction convolutional neural network includes: four convolutional layers, four maximum pooling layers and two fully connected layers, and the activation function is the ReLU function: Among them, the number of convolution kernels in the convolution layer is 8, 16, 32, and 64 respectively, the size of the convolution kernel is 5*5, and the step size is 1; the pooling kernel size of the maximum pooling layer is 3*3, and the step size is 2; the number of nodes in the fully connected layer is 3136 and 256, and the output feature size of the CNN network is 1*256; Step 3.2, adding a visual attention mechanism to the visual feature extraction convolutional neural network, including a channel attention mechanism and a spatial attention mechanism; The channel attention mechanism performs global average pooling and global maximum pooling on the input features, inputs the results of the global average pooling and the global maximum pooling into a shared network consisting of a multi-layer perceptron (MLP) and a hidden layer for processing, adds the two processed results and uses the sigmoid function to obtain the channel attention weight of the input feature, and multiplies the obtained channel attention weight by the input feature, that is: in, is the channel attention weight, is the input feature, is the sigmoid function, is the weight of the MLP, is the result of the global average pooling and the global maximum pooling; The spatial attention mechanism calculates the maximum and average values on the channel of each input feature, stacks the convolutional layer with a channel number of 1, applies the sigmoid function to the result, and obtains the weight corresponding to each input feature. The corresponding weight obtained is multiplied by the input feature: in, is the spatial attention weight, is the input feature, is the sigmoid function, for Convolution operation, is the result of spatial average pooling and maximum pooling; Step 3.3: The audio-visual dual-modality weld penetration state recognition model also includes a sound feature extraction convolutional neural network. The structure of the sound feature extraction convolutional neural network is generally consistent with that of the visual feature extraction convolutional neural network. The sound feature extraction convolutional neural network uses a multi-scale convolution kernel combination convolution layer to replace the 5*5 convolution kernel in the network, and performs appropriate padding before convolution. By comparing the classification accuracy obtained from different combinations, the parallel channels corresponding to the 3*8, 4*6, 5*5, 6*4, and 8*3 convolution kernels are optimized. Step 3.4: Perform feature fusion on the audiovisual features, concatenate the 256-dimensional feature vector extracted by the visual feature extraction convolutional neural network and the 256-dimensional feature vector extracted by the sound feature extraction convolutional neural network to obtain a 512-dimensional fused feature vector, and then input the obtained 512-dimensional fused feature vector into the classification unit, perform two full-connection operations to convert it into a 4-dimensional feature vector, and enter the softmax layer to complete the final weld penetration state classification task.
8. The method for monitoring the weld state of a MIG weld based on audio-visual dual modality according to claim 7, wherein: In step 4, the weld penetration state recognition model is trained, including setting the corresponding loss function, optimizer, and learning rate, batch size, epoch, dropout ratio, and num_classes parameters.
Citation Information
Patent Citations
Penetration state determination method based on small hole characteristic on back side
CN105741306A
Method for assisting welder to judge weld penetration state on line by utilizing electric arc sound
CN114633000A
Nondestructive detection method for weld penetration depths of laser non-penetration welding
CN112157368A
Welding quality real-time detection method and system based on high-frequency time sequence data
CN114722883A