Welding quality on-line monitoring method based on sound-light combination advanced prediction
By using a multimodal cross-attention fusion model combining acousto-optics, the problem of accurately predicting the dynamic evolution of the molten pool state in tungsten inert gas (TIG) welding was solved, enabling advanced prediction and stable control of welding quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF SCI & TECH
- Filing Date
- 2025-12-10
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies struggle to effectively address the complex dynamic evolution of the molten pool state in tungsten inert gas (TIG) welding. In particular, the combination of image prediction and regression methods suffers from error accumulation and feature degradation, making it difficult to achieve accurate welding quality prediction.
A multimodal cross-attention fusion model based on acoustic-optical combination is adopted. The molten pool image and sound signal are acquired through the welding acoustic-optical acquisition system. Multimodal feature fusion and temporal modeling are performed using CNN feature extraction module, attention module and Transformer temporal prediction module to achieve advanced prediction of welding quality.
It improves the accuracy and robustness of welding quality prediction, reduces error accumulation, enhances the ability to distinguish complex welding conditions, and achieves accurate prediction of future molten pool conditions.
Smart Images

Figure CN121289845B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multimodal imaging technology and relates to an online monitoring method for welding quality based on acoustic-optical combined advanced prediction. Background Technology
[0002] Tungsten inert gas (TIG) welding is widely used for high-precision welding of metals such as stainless steel, aluminum, and magnesium due to its high arc stability, excellent weld quality, aesthetically pleasing weld formation, and lack of spatter. It is particularly suitable for thin-plate welding, achieving extremely low heat input and avoiding material damage, thus playing a vital role in high-end manufacturing fields such as nuclear industry, aerospace, and food processing. However, its penetration depth is relatively shallow, efficiency is low, and it is susceptible to arc deflection and electromagnetic disturbances, making it difficult to guarantee weld stability. Especially in thick-plate welding, multiple operations are required, further increasing heat input and deformation risks. Therefore, achieving intelligent monitoring and control of the welding process has become a key direction for TIG process development.
[0003] To improve the real-time monitoring capability of the molten pool state, researchers have explored various sensing methods, including images, sound, and electricity. Among these, images, due to their ability to intuitively express changes in the molten pool morphology, are gradually becoming the most promising monitoring method. To overcome the imaging difficulties caused by strong arc light interference, various filter imaging systems have been proposed. One researcher developed a vision system that includes a CCD camera and a narrowband filter with a center wavelength of 655nm, a bandwidth of 40nm, and a transmittance of 85%, capturing images of the molten pool from below the workpiece. An improved narrowband filter with a wavelength of 660nm, a bandwidth of 35nm, and a transmittance of 90% has also been introduced, effectively suppressing arc light interference. Combined with new feature parameters such as the pinhole area and tilt angle for quantitatively describing the melt penetration state, this significantly improves the ability to acquire molten pool images under arc light interference conditions.
[0004] With the development of algorithms, traditional machine learning methods such as Random Forest (RF), Extreme Learning Machine (ELM), and Support Vector Machine (SVM) have been applied to the automatic identification of penetration status and weld defects in arc welding. For example, SVM and decision tree models have been used to identify and classify three types of weld defects—incomplete penetration, burn-through, and porosity—in gas shielded metal arc welding. Researchers have achieved effective differentiation between defective and good welds in surface defect images of friction stir welding based on kernel-function SVM. Other researchers have proposed a welder-intelligent enhanced deep random forest fusion method, WI-DRFF, which integrates deep features extracted by convolutional neural networks with manual keyhole features, using RF to predict penetration and significantly improve classification accuracy.
[0005] While traditional machine learning methods have achieved some success, their performance often relies on manual design and feature extraction, making them ill-suited to the complex backgrounds and unstructured interference encountered during welding. Furthermore, they suffer from information loss and insufficient generalization ability. Therefore, recent research has shifted towards end-to-end deep learning methods. Leveraging their powerful automatic feature learning capabilities, these methods extract high-level discriminative features from raw images or spectral signals. Combined with strategies such as image segmentation and attention mechanisms, they demonstrate higher accuracy and robustness in molten pool state recognition.
[0006] In existing technologies, researchers have proposed the Res-Seg multi-scale feature fusion semantic segmentation network to accurately extract the edge of the molten pool in TIG welding, significantly alleviating the difficulty of weak edge recognition caused by arc interference and reflection. By combining molten pool image preprocessing and contour analysis, the deviation information between the welding wire and the weld tip center corresponding to curvature extrema is mined, thereby achieving automated welding deviation detection. A multi-scale basis function modeling framework is constructed to sparsely represent the three-dimensional morphology of the molten pool, which can be used to predict size changes and identify spatter. These studies provide a deep characterization of the molten pool morphology at the image level, offering a data foundation for welding process state analysis.
[0007] However, the morphological changes of the molten pool during welding are not isolated and static, but rather a complex process of dynamic evolution with varying process parameters and physical behavior. Although CNN-based image feature extraction models have achieved good results in state recognition of single-frame images, they neglect the potential temporal dependencies in the image sequence, making it difficult to fully explore the evolutionary patterns of the molten pool state over time. Therefore, incorporating time-series information into the modeling process and jointly analyzing multiple frames not only helps improve the accuracy of current state recognition and more accurately reflects potential fluctuations and trends during welding, but also provides the possibility for predicting the molten pool state in future frames.
[0008] A prominent feature of molten pool images is the flow of liquid metal on its surface, which is closely related to the internal structure and heat input of the molten pool. Due to the effect of temperature accumulation, subsequent frames of molten pool images are often influenced by previous frames. Based on this, video prediction networks can be used to model its temporal evolution, providing support for early perception of welding quality. Common methods include ConvLSTM, PredNet, PredRNN, FutureGAN, MIMO-VP, iTransformer, etc. These methods, by predicting future image frames, can achieve regression prediction of quality indicators such as weld penetration or back weld reinforcement, providing a new path for proactive control of welding quality.
[0009] Researchers have used an improved PredNet prediction network to predict future molten pool images and combined it with a SEResnet regression network to predict the shape of the molten pool within 140ms, achieving an average accuracy of less than 0.3mm for weld reinforcement regression. However, while this method achieves early perception of the molten pool state, the two-stage approach combining image prediction and regression employs a frame-by-frame prediction strategy, which easily leads to strong error accumulation, affecting the accuracy of the final regression result. Furthermore, the serial "quadratic regression" is prone to feature degradation at the information level, further exacerbating the prediction error.
[0010] To enhance the modeling ability of temporal features in molten pool images, researchers have gradually introduced time-dimensional analysis strategies. RNNs and LSTMs possess powerful time modeling capabilities and have become effective methods to overcome the limitations of CNNs in extracting features from single-frame images. Some researchers have divided ultrasonic welding process signals into time segments and input them into an LSTM, achieving accurate classification and recognition of welding states. A top-view vision-based method for predicting weld width during laser welding has also been proposed. Using the arc and molten pool as model inputs, an ATT-LSTM (Attention-Long Short-Term Memory) post-weld width prediction model has been constructed, improving the accuracy and generalization ability of weld back width prediction. Furthermore, a two-stage segmentation-LSTM model for K-TIG welding penetration state has been proposed. This model uses a segmentation network to extract weld pool geometric features and then uses traditional algorithms to extract the weld gap. The features are then input into the model, achieving a prediction accuracy of 95.2% for the penetration state. In the scenario of laser powder bed melting (LPBF), researchers have constructed an RNN prediction framework that integrates images and process parameters, introduced data from the previous scan cycle, and used a GRU modeler to improve the ability to characterize the evolution trend of the molten pool, reducing the MAPE to 14.8%.
[0011] Besides visual information, sound signals, as an important representation channel of the molten pool state, exhibit unique advantages in judging weld depth and detecting anomalies. Different welding states correspond to different spectral structures and energy change patterns. For example, the arc sound in a normal welding state is relatively stable and the spectrum is concentrated, while when there is no fusion or spatter anomalies, the sound signal exhibits short-term high-frequency abrupt changes or abnormal energy fluctuations. Using acoustic feature extraction methods such as Short-Time Fourier Transform (STFT), the welding state can be effectively identified. In existing technologies, researchers have proposed an auditory recognition model combining attention mechanisms and Long Short-Term Memory (LSTM) networks, based on 15-dimensional acoustic features applied to weld penetration state identification, achieving an average accuracy of 95.32%. An evaluation model combining Support Vector Machine (SVM), grid search, and cross-validation has been constructed for online detection of weld depth defects. Furthermore, a convolutional neural network (TF-CNN) model based on time-frequency feature maps has been proposed, achieving the identification of four typical weld depth states with an accuracy of 98.2%. The formation mechanism of arc sound during pulsed GTAW welding was analyzed using a long short-term memory network (LSTM), and the depth of the weld was predicted, effectively revealing the intrinsic relationship between sound characteristics and weld pool evolution.
[0012] In summary, visual information excels at depicting the geometric evolution during the welding process, while auditory information is more sensitive to energy fluctuations and changes in internal structure. Integrating visual and auditory multimodal information is expected to significantly compensate for the shortcomings of single-modal representation and is an important development direction for improving the accuracy of molten pool state prediction.
[0013] To address the problem of multimodal feature fusion, existing technologies have proposed a cross-attention fusion network, CAFNet, which automatically integrates optical and acoustic information to classify welding quality without the need for time-frequency analysis or manual feature extraction, utilizing an interactive attention mechanism. A multispectral channel attention mechanism multimodal fusion network was designed, introducing parallel feature mapping to capture local and global dependencies, enhancing receptive field interaction and global modeling performance. It strengthens cross-channel information interaction, effectively fusing local high-frequency and global low-frequency features, achieving a classification accuracy of 98.8%. For the laser selective melting (LPBF) process, a dual-stream cross-modal fusion network was proposed, dividing high-quality images and acoustic features into "recoating stream" and "laser stream," respectively. Efficient feature integration was achieved through a residual bilinear cross-fusion strategy, demonstrating good robustness and generalization ability in complex scenarios.
[0014] Multimodal fusion has significant advantages in welding condition monitoring. Time series modeling methods can utilize information from multiple consecutive frames to form a contextual understanding of the current state and achieve forward-looking prediction of future trends.
[0015] Therefore, a new online monitoring method for welding quality based on combined acoustic and optical prediction is needed to solve the above problems. Summary of the Invention
[0016] The purpose of this invention is to provide an online monitoring method for welding quality based on acoustic-optical combined advanced prediction, so as to overcome the shortcomings of the prior art.
[0017] The technical solution of the present invention is as follows:
[0018] The online monitoring method for welding quality based on acoustic-optical combined advanced prediction adopts a welding acoustic-optical acquisition system, which includes an image acquisition device and a sound acquisition device.
[0019] The image acquisition device includes an infrared LED array, a camera, and a bandpass filter. Both the infrared LED array and the camera face the welding area, and the bandpass filter is positioned in front of the camera.
[0020] The sound acquisition device includes a microphone;
[0021] The online welding quality monitoring method includes the following steps:
[0022] 1) The welding acoustic-optical acquisition system is used to acquire the image sequence and sound sequence of the molten pool during the welding process. The sound sequence is converted into a spectrum diagram aligned with the time step of the image sequence by a short-time Fourier transform (STFT). According to the weld state, the weld is divided into incomplete penetration weld, normal penetration weld and burn-through weld. The corresponding image sequence, spectrum diagram and weld state are used to construct a welding process monitoring dataset.
[0023] 2) Input the welding process monitoring dataset into the multimodal cross-attention fusion model to complete the training of the multimodal cross-attention fusion model;
[0024] The multimodal cross-attention fusion model includes a CNN feature extraction module, an attention module, a gating fusion module, and a Transformer temporal prediction module.
[0025] The CNN feature extraction module uses the melt pool image sequence and spectrogram to extract melt pool image features and spectrogram features, respectively.
[0026] The attention module utilizes the features of the melt pool image to perform guided modeling of the spectrogram features, thereby obtaining cross-attention features;
[0027] The gated fusion module is used to perform gated fusion between the melt pool image features and the cross-attention features to obtain a fused feature sequence.
[0028] The fused feature sequence is input into the Transformer time-series prediction module for dynamic modeling, and the classification prediction result of the weld state is output.
[0029] 3) Input the image and sound sequences from the welding process into the multimodal cross-attention fusion model trained in step 2) to achieve classification and prediction of weld state.
[0030] Furthermore, the CNN feature extraction module in step 2) includes two feature extraction modules. Each feature extraction module comprises a CBR fusion layer, a 3*3 max pooling layer, a densely connected convolutional network, and a global average pooling layer, connected sequentially. The CBR fusion layer includes a Convolutional layer (Conv), a BN layer, and a first ReLU layer. The densely connected convolutional network includes multiple DenseNet-SE modules connected sequentially. Each DenseNet-SE module includes a DenseBlock module, a Transition module, and an SE channel attention module. The DenseBlock module includes multiple DenseLayer layers. The system consists of a first BatchNorm layer, a second ReLU layer, a first 1×1 convolutional layer, a second BatchNorm layer, a third ReLU layer, and a 3×3 convolutional layer connected in sequence. The Transition module includes a second 1×1 convolutional layer and a 2×2 average pooling layer. The SE channel attention module includes a global average pooling layer, a dimensionality-reduced 1×1 convolutional layer, a ReLU activation layer, an up-dimensional 1×1 convolutional layer, and a Sigmoid activation layer connected in sequence. The global average pooling layer is used to extract channel descriptions, and the dimensionality-reduced 1×1 convolutional layer, ReLU activation layer, up-dimensional 1×1 convolutional layer, and Sigmoid activation layer are used to model the non-linear relationship between channels and generate channel attention weights.
[0031] The attention module includes a self-attention module, a cross-attention module, and a feedforward network, which are connected in sequence.
[0032] The gated fusion module includes a channel splicing layer, two parallel gated networks, and a residual connection based on modal input, which are connected in sequence. The channel splicing layer is used to splice the features of the two modalities in the channel dimension. The two parallel gated networks perform nonlinear transformations on the spliced features to generate dynamic weights corresponding to the two modalities. The residual connection introduces the original modal features into the weighted fusion result.
[0033] The Transformer temporal prediction module includes an input projection layer, a position encoding embedding layer, a Transformer encoder stack layer, and a classification output layer connected in sequence.
[0034] Furthermore, the attention module described in step 2) is represented by the following formula:
[0035] ,
[0036] In the formula, The final output features represent the result of fusing self-attention, cross-attention, and feedforward networks. The input is the feature sequence of the main mode. The input auxiliary modality features are all of shape [B, T, D], where B is the batch size, T is the time series length, and D is the feature dimension. The representation of the master mode after it has completed information exchange within its own sequence. This represents the updated features of the main modality after cross-modal attention. , and These represent three LayerNormalization layers, with normalization performed before each sub-layer. This represents the self-attention mechanism within a modality, modeling the temporal relationships within the input sequence. This represents a cross-modal attention mechanism.
[0037] Furthermore, the gating fusion module described in step 2) is represented by the following formula:
[0038] ,
[0039] In the formula, This represents the output characteristics of two modes after gated weight fusion and residual connection. The feature vectors are from the image channels. Features resulting from cross-modal interaction This represents the concatenated features used to generate the gating weights. This represents the Sigmoid activation function. , These represent the gating weights of the two modalities, controlling the dynamic contribution of their respective features. and These are the linear transformation parameters, i.e., the weight matrix of the MLP.
[0040] Furthermore, the time series prediction module described in step 2) is represented by the following formula:
[0041] ,
[0042] Where, x∈R B×T×D Given the input sequence Z∈R B×3 This represents the class probability of the three welding states. This represents the position index number, ranging from [0, T]. The extra position is used to categorize the token. This indicates the position embedding module, which provides temporal information about the sequence. Indicates the input linear projection layer. Used to prevent overfitting It consists of several stacked TransformerEncoderLayer layers, performing encoding with multi-head self-attention and feedforward networks.
[0043] Furthermore, the angle between the microphone and the welding torch is 75°.
[0044] Furthermore, the bandpass filter is a 940nm bandpass filter.
[0045] Furthermore, both the image acquisition device and the sound acquisition device are mounted on the welding torch robotic arm.
[0046] Invention principle: The advantage of this invention comes from the ability of the image mode to perceive changes in spatial structure and the complementarity of the sound mode to the dynamic characteristics of energy disturbance. The two work together to improve the model's ability to distinguish complex welding states.
[0047] Beneficial effects: The online welding quality monitoring method based on acoustic-optical combination of the present invention uses CNN to extract molten pool image features and acoustic spectrum features, effectively integrates information in spatial and temporal dimensions, achieves deep interaction through cross-modal attention mechanism and gating fusion structure, and then uses Transformer to model temporal dependence to achieve advance prediction of future molten pool state. Attached Figure Description
[0048] Figure 1 This is a schematic diagram of the welding acoustic and optical acquisition system;
[0049] Figure 2 This is a structural schematic diagram of the base material for welding;
[0050] Figure 3 This is a schematic diagram of the DenseNet-SE module;
[0051] Figure 4 This is a schematic diagram of the Transformer time series prediction module;
[0052] Figure 5 This is a schematic diagram of the structure of a multimodal cross-attention fusion model;
[0053] Figure 6 It is the result of a combination of CNN and ViT;
[0054] Figure 7The classification accuracy is tested in a 0.5s future frame.
[0055] Figure 8 Predict confusion matrices for different OLs: (a) 0.05s, (b) 0.25s, (c) 0.5s, (d) 1s;
[0056] Figure 9 Visualization of t-SNE for different OL predictions: (a) 0.05s, (b) 0.25s, (c) 0.5s, (d) 1s. Detailed Implementation
[0057] 1. TIG welding acoustic and optical acquisition system
[0058] The platform and its acoustic and optical acquisition system, such as Figure 1 As shown. The platform includes a power cabinet, an argon arc welding machine (PI350 / 500AC / DC, a three-phase water-cooled welding machine suitable for MMA and TIG welding), a welding torch robotic arm, a protective gas cylinder (argon), a wire feeder, and a water cooling system.
[0059] During the welding process, the welding torch position is fine-tuned, with the distance between the tungsten electrode tip and the outer end face of the nozzle set to 4.0 mm, and the distance between the tungsten electrode tip and the workpiece surface maintained at 3.0 mm. To prevent oxidation and contamination during welding, the shielding gas flow rate is set to 15 L / min to provide sufficient protection. To maintain the continuity and uniformity of the welding process, the welding speed is set to 4 mm / s. A water-cooling system is also used to cool the welding torch, ensuring a stable operating temperature during long-term use and extending its service life.
[0060] During welding, traditional vision systems struggle to acquire clear images due to interference from factors such as strong arc light, high-temperature molten pool, and spatter. To improve the imaging quality of the molten pool area during welding, this system employs near-infrared active illumination. It illuminates the welding area using a 940nm infrared LED array and incorporates a 940nm bandpass filter in the camera's optical path to effectively shield against visible light interference from the arc, allowing only infrared reflections of the target wavelength to enter the imaging system. The sound acquisition system consists of an MP201 condenser microphone and a high-precision data acquisition card. The MP201 boasts excellent frequency response and low-noise performance, enabling high-sensitivity acquisition of broadband acoustic signals during welding. The acquired analog signals are converted to digital signals (ADC) and then input into a computer for subsequent acoustic feature extraction and multimodal analysis, providing reliable sound information support for molten pool condition identification and process monitoring.
[0061] The camera was controlled by the professional software MindVision, with parameters set to a resolution of 1280*1080, an exposure time of 4.3µs, and a frame rate of 200 frames per second, ensuring the continuity and accuracy of image capture. These parameters are widely accepted in actual manufacturing scenarios. Research found that a 75° angle between the microphone and the welding torch is the optimal angle for data acquisition. The distance between the GTAW welding torch and the microphone was 350mm, the microphone sampling rate was 51.2kHz, and the microphone and active illumination camera were used for synchronous data acquisition.
[0062] 2. Dataset
[0063] To simulate typical penetration depth variations during actual welding, this invention uses a dumbbell-shaped 304L stainless steel plate with a thickness of 3mm for TIG welding experiments. (See attached image.) Figure 2 The specimen structure is "wide at both ends and narrow in the middle," which can naturally form three typical penetration depth states in a single weld: "incomplete penetration," "complete penetration," and "burn-through." Specifically, in the initial stage, insufficient heat input easily leads to incomplete penetration at the ends; the middle section, due to its short heat diffusion path and poor heat dissipation, has moderate heat and often achieves complete penetration; while in the later stages of welding, excessive heat accumulation, coupled with the influence of grooves, easily leads to burn-through defects. This design provides an ideal experimental basis for modeling and judging multiple penetration depth states.
[0064] Image and sound data during the fusion welding process were collected using image acquisition device 1 and sound acquisition device 2 to construct a multimodal welding process monitoring dataset. Each data set has clear time sequence information and is divided into three categories according to the weld morphology: incomplete penetration, normal penetration, and burn-through.
[0065] Five sets of experimental data were collected, with three sets used for training, one for validation, and one for testing. The data contains approximately 44,250 images and their corresponding audio sequences: 26,571 images for model training, 9,177 images for validation, and 8,502 images for testing. To ensure the temporal continuity of the training, all data was retained without removing intermediate transition frames. Table 1 provides the parameter information for the dataset used in the model.
[0066] Table 1. Welding Parameters
[0067] Serial Number Current (A) Wire feeding speed (m / min) Welding speed (mm / s) 1 140 1.2 4 2 140 1.2 4 3 140 1.2 4 4 140 1.2 4 5 150 1.4 4
[0068] 3. Construction of CNN-Transformer Model
[0069] 3.1 CNN-Transformer Strategy
[0070] The model architecture proposed in this invention effectively integrates spatial and temporal information by introducing a modal feature extraction module and a temporal modeling module in the welding state prediction task.
[0071] In the feature extraction stage, the image modality uses a CNN structure to extract local spatial features from the weld pool image, focusing on capturing regional features closely related to welding stability, such as the weld pool edge and brightness distribution. Simultaneously, the acoustic modality inputs a spectrogram (e.g., STFT) into the CNN to extract the time-frequency features of the acoustic signal within the weld pool, which helps determine the weld pool state. Both modalities automatically learn key representations through deep convolutional operations, laying the foundation for subsequent multimodal fusion.
[0072] In the temporal modeling and prediction classification stage, a Transformer module is introduced to model the time series data, which can capture important dynamic change features across frames during the welding process. A cross-modal attention mechanism is used to interactively fuse image and audio modal features, and the fused features are fed into the Transformer for temporal modeling, finally outputting the welding state classification result for each frame.
[0073] This model fully combines the advantages of CNN in spatial modeling with the capabilities of Transformer in temporal modeling and feature fusion. It can not only accurately extract and utilize key feature information in the welding process, but also improve the accuracy and robustness of state prediction, providing a more reliable solution for intelligent identification of complex molten pool states.
[0074] 3.2. CNN Feature Extraction Module
[0075] The core idea of CNNs is to automatically extract spatial features from input data through local receptive fields, weight sharing, and multi-layer feature extraction mechanisms. They typically consist of modules such as convolutional layers, non-linear activation functions, pooling layers, and fully connected layers. In welding image analysis tasks, CNNs can automatically extract the geometric texture features of the molten pool region. By temporally modeling the features of multiple frames of images, CNNs can serve as spatial encoders for multimodal inputs, effectively assisting subsequent temporal prediction modules.
[0076] To fully extract spatial structural information from image modalities, this invention employs an improved DenseNet densely connected convolutional network structure, introducing a channel attention module to construct the DenseNet-SE model for spatial feature extraction of weld pool image sequences. The model structure is as follows: Figure 3 As shown.
[0077] DenseNet effectively mitigates gradient vanishing and promotes feature reuse by introducing a dense connection mechanism between convolutional layers, allowing features from previous layers to be directly passed to all subsequent layers. This structure solves the gradient vanishing and feature reuse problems in deep networks. Its basic components are the DenseBlock and Transition modules: the DenseBlock module consists of multiple DenseLayer modules, each of which sequentially performs BatchNorm, ReLU, 1×1 convolution, BatchNorm, ReLU, and 3×3 convolution, and achieves feature accumulation and enhancement through channel concatenation; the TransitionLayer performs feature dimensionality reduction and downsampling through 1×1 convolution dimensionality reduction and 2×2 average pooling.
[0078] To further improve the discriminative power of the features, an SE channel attention module is added after each Transition layer. This module first extracts channel descriptions using global average pooling (Squeeze), then models the non-linear relationship between channels using two 1×1 convolutions and generates attention weights (Excitation). Finally, these weights are applied to the original feature map for dynamic reweighting, thereby highlighting key channel information and enhancing the ability to identify welding states. The final extracted feature map is compressed into a global feature vector required for temporal modeling using global average pooling and serves as the input to the Transformer module.
[0079] 3.3. Transformer Time Series Prediction Module
[0080] To model the temporal evolution of multimodal fusion features and achieve advanced prediction of welding states in future frames, this invention designs a lightweight Transformer module based on the spatial features extracted by CNN. This module mainly includes an input projection layer, a positional encoding embedding layer, a stacked Transformer encoder layer, and a classification output layer. The model structure is as follows: Figure 4 As shown.
[0081] The input sequence is x∈R B×T×D Where B represents the batch size, T is the time series length, and D is the single-frame feature, derived from multimodal fusion extracted by the CNN module. A learnable class token [cls] is designed to capture the global representation of the entire sequence and serve as the basis for final classification. This token is concatenated at the beginning and end of the sequence, followed by a learnable positional encoding to enhance the model's ability to perceive temporal order.
[0082] (1)
[0083] Since the original input dimension D may differ from the hidden layer dimension of the Transformer module, this paper introduces a linear transformation layer for dimension mapping and uses Dropout to alleviate overfitting. The Transformer encoder adopts a standard structure, consisting of a multi-head self-attention mechanism and a feedforward network. Each layer contains residual connections and layer normalization. The projected sequence features are input into the stacked Transformer encoder:
[0084] (2)
[0085] Output tensor Z∈R T+1×B×d The first frame Z0∈R B×d The final representation of the learned [cls] token, representing the global features of the entire sequence, is projected onto a three-class classification space through Dropout and a fully connected layer, with the output being Z∈R. B×3 The probability of each of the three welding states corresponds to a category.
[0086] 3.4. MultiModel-CNN-Transformer Model
[0087] To achieve collaborative modeling of image and sound information during welding and temporal classification prediction of future states, this invention proposes a multimodal cross-attention fusion model, MultiModel-CNN-Transformer. This model takes images and STFTs (Spectroradiograms with Thin-Frame Theorem) as inputs, extracts their spatial features respectively, models the collaborative relationship between the two through a cross-modal attention mechanism, uses a gating mechanism to achieve intermodal information fusion, and finally achieves classification prediction of future welding states through a temporal modeling module. The overall model consists of the following key modules: a CNN feature extractor, a CrossAttention module, a GatedFusion module, and a Transformer temporal prediction module. The structure is as follows: Figure 5 As shown.
[0088] During the data acquisition phase, the system simultaneously acquires images and audio signals of the molten pool during the welding process. The audio signal is converted into a spectrogram aligned with the image time step through a Short Time Fourier Transform (STFT). The image sequence and the corresponding spectrogram are fed into two independent Convolutional Neural Networks (CNNs) to extract spatial features. After multiple convolution and pooling operations, a high-dimensional feature map is output, which is then compressed into a fixed-length vector using Global Average Pooling (GAP). Finally, a fully connected dimensionality reduction module unifies the input dimension for subsequent sequence modeling. The model employs a sliding window strategy to construct temporal samples; for example, frames 1 to 10 are extracted as the first sample, and then frames 2 to 11 are processed sequentially, and so on.
[0089] Traditional multimodal fusion methods often employ simple concatenation or weighted averaging, neglecting the information dependencies between different modalities. To enhance the image modality's ability to perceive key dynamic changes in the audio modality, this invention designs a cross-attention module for guided modeling of audio information from image features. This module integrates Self-Attention, Cross-Attention, and a feedforward MLP network.
[0090] First, the Self-Attention module normalizes the input image features and feeds them into the standard multi-head attention mechanism to capture the global dependencies between frames in the image sequence, enhancing the visual modality's ability to focus on internal temporally consistent regions. The formula is as follows:
[0091] (3)
[0092] Secondly, the Cross-Attention module uses image features as the query and sound modal features as the key and value. Through a cross-attention mechanism, it captures the dependency between the image and sound, guiding the image channel to focus on key acoustic regions that are temporally synchronized with it. The process is as follows:
[0093] (4)
[0094] Finally, to enhance nonlinear expressive power, an MLP module consisting of two perceptron layers is introduced, combined with residual connections to improve modeling capabilities:
[0095] (5)
[0096] Considering the potential differences in importance between different modalities in specific scenarios, this invention introduces the GatedFusion module to achieve gated fusion between image features and cross-attention features. Since the cross-attention mechanism guides the image modality to focus on information from key regions in the audio modality, it achieves deeper modality alignment and dependency modeling at the semantic level, using the cross-attention fusion features of image and audio as the primary focus and the original features of the image channels as auxiliary features. This module includes two parallel gating networks, gate1 and gate2, which respectively calculate the dynamic weights of the two modalities, thereby achieving fine-tuning of information contribution, and combining residual structures to enhance feature stability.
[0097] Given a feature vector x1∈R from an image channel B×T×D Features x2∈R after cross-modal interaction B×T×D First, the two are spliced together along the channel dimension:
[0098] (6)
[0099] Then, weighted attention is obtained through two gating branches:
[0100] (7)
[0101] (8)
[0102] in, This represents the Sigmoid activation function. The linear transformation parameters (i.e., the weight matrix of the MLP) are used to generate the gating weights for each mode, with an output dimension of D.
[0103] The specific calculation of the fusion features is as follows:
[0104] (9)
[0105] The residual connection mechanism enhances the stable transmission of information and alleviates the gradient vanishing problem during the training of deep models. The output features are further normalized by LayerNorm, which improves the convergence stability.
[0106] The fused feature sequences are input into the Transformer temporal modeling module for dynamic modeling. This module first introduces a learnable class token (CLStoken) before each sequence and adds positional encoding to preserve temporal order information. Subsequently, through stacked computation by the Transformer encoder, the model extracts local and global temporal dependencies layer by layer, finally outputting the global semantic information represented by the CLS token. The output features are mapped to the final classification result through Dropout layers and fully connected layers.
[0107] This framework effectively improves the accuracy of molten pool state recognition and the stability of temporal continuous prediction by fusing complementary information from image and sound modalities and leveraging the global modeling capabilities of the Transformer.
[0108] 3.5 Evaluation Indicators
[0109] To evaluate the temporal prediction and classification performance of the constructed MultiModel-CNN-Transformer model for the molten pool state (incomplete penetration, full penetration, burn-through) during TIG welding, this invention uses a three-class confusion matrix and its derived evaluation metrics to comprehensively analyze the model's prediction results.
[0110] In a three-class classification problem, the confusion matrix is a 3×3 lookup table that corresponds to the predicted and actual combinations of the three classes. For any class (e.g., "not fully fused"), it can be considered a positive class, and the other classes can be considered negative classes. This allows us to calculate its TruePositive (TP), FalseNegative (FN), FalsePositive (FP), and TrueNegative (TN). For example, TruePositive (TP) represents the number of samples correctly identified as "not fully fused" by the model; FalseNegative (FN) represents the number of samples that are actually "not fully fused" but were incorrectly predicted as other classes; FalsePositive (FP) represents the number of samples that are not actually "not fully fused" but were incorrectly predicted as "not fully fused"; and TrueNegative (TN) represents the number of samples that are neither actually "not fully fused" nor predicted as "not fully fused".
[0111] Based on the confusion matrix, the following performance metrics are further calculated to comprehensively evaluate the overall classification performance of the model: Accuracy, Precision, and F1 score. Precision and F1 score are calculated using a macroaverage method, that is, the arithmetic mean of the metrics for each of the three classes is calculated separately, without considering differences in the number of samples in each class, thus better reflecting the model's balanced performance across different classes. The relevant mathematical expressions are as follows:
[0112] (10)
[0113] (11)
[0114] (12)
[0115] (13)
[0116] Where N=3 represents the number of categories, and i represents the i-th category. Through comprehensive analysis of the above indicators, the model's generalization ability and practical application value in three-class classification tasks can be effectively evaluated.
[0117] The training and testing of this model were conducted on a computer equipped with an Intel Core i7 10700 CPU and an NVIDIA GeForce RTX 3090 GPU. The development environment was Python 3.10, implemented using the PyTorch framework. During training, to effectively train the constructed deep learning model, this paper employs a parameter optimization strategy based on the Adam algorithm, combined with a Warmup-Cosine learning rate scheduler to dynamically adjust the learning rate during training. This improves the model's convergence speed and generalization ability, effectively alleviating the problem of traditional SGD easily getting trapped in local optima or gradient vanishing when training deep neural networks. After multiple experimental verifications, the learning rate was set to 0.0005, and training was performed for a total of 80 iterations. In the testing phase, the model weights that performed best on the validation set were used for evaluation to ensure the reliability and accuracy of the results.
[0118] 4 Experiments
[0119] 4.1 Model Combination Selection
[0120] To evaluate the ability of the multimodal cross-attention fusion model to predict the melt pool state across different time ranges, this invention implemented seven training configurations. Specifically, 0.5 seconds was selected as the input IL (melt pool image and sound spectrum sequence samples), and the sample labels after 0.05s, 0.25s, 0.5s, 1s, 1.5s, 2s, and 2.5s were selected as the output OL (olive), which verifies the ability to predict the melt pool state at different prediction time points. This method helps to comprehensively evaluate the model's ability to capture temporal dependencies and its performance across different time ranges.
[0121] To ensure the overall reliability of the model, the effectiveness of Densenet_SE in spatial feature extraction was first verified, and comparative experiments were conducted with ResNet-34 and ViT models. The comparison results of feature extraction and classification accuracy of different CNN models are shown in Table 2. Experiments show that when Densenet_se201 is selected as the backbone network of the CNN, the model's test prediction accuracy can reach 93.38%. This advantage is attributed to the effective information reuse and feature transfer of the dense connection mechanism in the Densenet_se201 network structure, enabling the network to fully mine multi-level fine-grained features and enhance its ability to express complex textures and patterns in melt pool images and spectrograms. In contrast, Densenet_se121 has a slightly lower accuracy than se201, indicating that shallower dense networks can still capture spatiotemporal information well. ResNet34_se performs poorly, possibly due to its shallower residual structure and fewer feature paths limiting its ability to express complex multimodal signals and making it difficult to fully capture the detailed features of melt pool and spectrogram data.
[0122] In summary, this experiment verifies the superior performance of the temporal prediction architecture based on Densenet_se201 combined with Transformer in multimodal welding process state classification, demonstrating the synergistic effect of deep dense feature extraction and global temporal modeling, and providing strong support for subsequent multimodal industrial vision and sound fusion analysis.
[0123] Table 2. Test results of different CNNs under Transformer
[0124] Model Accuracy Precision f1-score DenseNet_se121 0.8989 0.8444 0.8548 DenseNet_se201 0.9338 0.9132 0.9196 ResNet34_se 0.8063 0.8215 0.7661
[0125] To further explore the synergistic relationship between visual and acoustic modalities in multimodal temporal classification tasks, we systematically compared the effects of four feature combination strategies (CNN+CNN, ViT+CNN, CNN+ViT, and ViT–ViT) and their combinations with Transformer. Specifically, we used CNN and ViT as feature extractors in the image and spectral channels, respectively, forming four combinations: CNN–CNN, ViT–CNN, CNN–ViT, and ViT–ViT. Experimental results are as follows: Figure 6 As shown.
[0126] Experimental results show that when both image and audio channels use Convolutional Neural Networks (CNN+CNN), most model combinations exhibit superior performance. For example, DenseNet-SE201 achieved the highest accuracy of 93.38%, highlighting the strong expressive power of the bimodal features extracted by CNNs in modeling local modes of images and spectrograms. This performance advantage is mainly attributed to the stable and dense local features provided by CNNs, which facilitate model training and fusion. Furthermore, the image and audio features extracted by CNNs have high consistency in spatial structure, which is beneficial for intermodal alignment. In addition, the Transformer architecture itself is sensitive to the consistency and distribution stability of input features, making it more suitable for the feature space constructed by dual CNNs. However, when ViT is introduced into the model structure (such as ViT+CNN or CNN+ViT), the overall performance decreases to varying degrees. The ViT-ViT combination performs very poorly and is not listed. In particular, the CNN+ViT combination achieves significantly lower accuracy than the dual-CNN architecture. This is because the STFT spectrogram, as a two-dimensional time-frequency representation of a sound signal, inherently retains local temporal continuity and frequency band distribution, exhibiting significant local structure. This makes it suitable for extracting short-term energy changes and frequency features through convolution operations. ViT, on the other hand, focuses more on global modeling and capturing long-range dependencies, lacking sensitivity to local textures, edges, and abrupt changes. This may lead to its difficulty in fully extracting discriminative features when processing images with strong local patterns, such as STFTs.
[0127] In summary, in the multimodal joint modeling task of melt pool images and spectrograms, the CNN–CNN architecture, with its good structural consistency and local modeling capabilities, combined with the dense connection mechanism of DenseNet-SE201, achieves optimal feature fusion performance. DenseNet's layer-by-layer connection method enhances feature propagation and reuse, facilitating information alignment and fusion between different modalities, making it the most suitable backbone structure for the current task. In contrast, ResNet and ViT still have certain shortcomings in modality fusion, inductive bias, and feature coordination. Further exploration of strategies such as structural symmetry optimization, modality alignment loss guidance, and adaptive fusion mechanisms can improve the model's representation and generalization capabilities under heterogeneous modality combinations.
[0128] 4.2 Model Time Series Prediction Results
[0129] The DenseNet_SE201+Transformer model was used to predict the future frame states of molten pool image sequences at different time points. The sample labels after 0.05s, 0.25s, 0.5s, 1s, 1.5s, 2s, and 2.5s were selected as the output logistic regression (OL). In the task of predicting the molten pool state at 0.05s, 0.5s, and 1s in the next few frames, the model achieved classification accuracies of 93.52%, 96.23%, and 94.73%, respectively, effectively distinguishing between incomplete melting, complete melting, and burn-through states in advance. By predicting new samples not used in training, accuracies of 89.64%, 93.38%, and 87.94% were obtained, demonstrating good generalization ability. The test set results are shown below. Figure 7 As shown.
[0130] Experimental results show that the model performs well in short-term predictions (0.05s~0.5s), with a maximum accuracy of 93.38%. However, as the prediction timeframe lengthens to over 1.0s, the accuracy gradually decreases, reaching a minimum of 82.52%. While the Transformer excels at modeling local and long-range dependencies, its prediction performance is still affected by the decay of forward information over long time intervals. Particularly when predicting future frames of 2.0s~2.5s, the correlation between the current frame and the target frame decreases significantly, making it difficult for the model to accurately capture effective discriminative features. Long-term prediction is essentially an extrapolation of the sequence trend; the model needs to extract higher-level structural information from the current input, rather than relying on instantaneous local features. Due to the nonlinear dynamics of the weld pool state, the longer the time, the greater the uncertainty. The Transformer's contextual modeling mechanism, which relies on existing inputs, amplifies the accumulated state prediction error over long time spans, resulting in a significant decrease in prediction accuracy after 2.0s. The results show that the model achieves its highest accuracy of 93.38% at the 0.5s prediction point, indicating that this time window may be within the optimal time range that the model can perceive. At this point, sufficient dynamic evolutionary clues are preserved, while the model has not yet entered an unpredictable stage with excessive feature decay and state changes.
[0131] To gain a deeper understanding of the model's feature discrimination capability during the testing phase, this paper introduces t-SNE (t-distributed Stochastic Neighbor Embedding) dimensionality reduction technology to visualize the multimodal fusion features extracted by the network in two-dimensional space, and represents the output results as a confusion matrix. Figure 8 and Figure 9 The confusion matrix and t-SNE visualization are shown for OL values of 0.05s, 0.25s, 0.5s, and 1s. The t-SNE visualization clearly shows the distribution characteristics of different welding states in the two-dimensional feature space, where most samples exhibit tight clustering within each class, while showing a clear separation between classes. Figure 9 (c) The phenomenon is most pronounced at 0.5s. This result fully demonstrates that the fusion feature extraction network designed in this paper has strong modality alignment and feature discrimination capabilities, providing a good feature foundation for downstream temporal classification tasks. Some overlapping areas exist in the visualization, possibly because the transition segment of the welding state reflects the critical penetration stage of the welding process. This is typically manifested at the image level as blurred boundaries and brightness and shape between complete and incomplete penetration; at the sound spectrum level, it also manifests as unstable or mixed distribution of feature amplitudes, making its representation in the feature space inherently "fuzzy" and "diverse," with some semantic overlap with the other two classes.
[0132] 4.3 Ablation Experiment
[0133] Table 3 Ablation results of unimodal and multimodal features
[0134] densenet_se201 Accuracy Precision F1 score ImageSpectrum 0.82430.9242 0.82820.8975 0.79280.9037 Multi-NoGFMulti-NoTransMulti-NoCA-GFMulti 0.87360.90010.85800.9338 0.86030.90100.83840.9132 0.87120.90330.83710.9196
[0135] To further verify the effectiveness of the multimodal fusion strategy in predicting the state of weld pool, a series of ablation experiments were designed to evaluate the performance of the model with separate inputs of the image modality (Image), the sound modality (Spectrum), the multimodal input (Multi), and the model after removing or replacing key modules, such as removing the gating module Multi-NoGF, replacing the Transformer with a simple average pooling + linear classifier Multi-NoTrans, and replacing the cross-attention and gating modules with a simple Concat (Multi-NoCA-GF). The experimental results are shown in Table 4.2.
[0136] When using only visual images as input, the model achieves a classification accuracy of 82.43%, indicating that the image modality effectively characterizes the geometry and surface evolution of the molten pool, serving as a crucial source of information for welding status identification. However, due to imaging instability factors such as blurred molten pool boundaries during Tig plate butt welding, visual modality recognition has certain blind spots, limiting its expressive power. In contrast, when using only acoustic spectrograms as input, the model achieves an accuracy of 92.42%, significantly higher than the image modality. This demonstrates that acoustic signals have a stronger representational ability in reflecting energy release and internal state disturbances during the fusion process, providing crucial supplementary information, especially when visual modality information is incomplete or interfered with.
[0137] In the multimodal infrastructure, the performance of Multi-NoGF model significantly decreased across all metrics after removing the GatedFusion gated fusion module, indicating that the gated fusion mechanism plays a crucial role in multimodal feature fusion. This module, by introducing a learnable gating mechanism, assigns dynamic weights to different features, which strengthens key information while suppressing redundancy and noise. Furthermore, combined with residual connections and normalization operations, it enhances the robustness and stability of the fused features. Compared to directly using CrossAttention output, the nonlinear structure of GatedFusion introduces a higher-order expressive power for modal interactions.
[0138] Furthermore, replacing the Transformer temporal prediction module with the simple average pooling linear classification module Multi-NoTrans also resulted in a decrease in model performance. This comparison shows that Transformer can more effectively model the temporal dependencies after modality fusion, especially in capturing the temporal dynamic changes between multi-frame fused pool images and spectrograms and the importance between keyframes, which is significantly better than static pooling schemes.
[0139] Furthermore, replacing the designed cross-attention + gating fusion mechanism with a simple Concat fusion (Multi-NoCA-GF) resulted in a further drop in model performance to 85.80%, making it the worst performer. This phenomenon indicates that simple concatenation fusion lacks the ability to model modal relationships, fails to uncover semantic dependencies or structural complementarity between images and sounds, and easily introduces information redundancy and modal misalignment issues. In contrast, CrossAttention can explicitly learn the response relationship between one modality and another, while GatedFusion further implements modal weighting and interference suppression, thereby achieving a more efficient fusion representation.
[0140] Based on the above experiments, the complete model employing a dual-modal input + cross-attention + gating fusion mechanism and using the transformer module for temporal prediction performed best among all structures, achieving an accuracy of 93.38%, thus validating the effectiveness of the multimodal fusion framework. Its advantage stems from the image modality's ability to perceive spatial structural changes and the audio modality's supplementary role in the dynamic features of energy perturbations. The synergistic expression of these two modalities enhances the model's ability to discriminate complex welding states, particularly demonstrating greater robustness in identifying boundary states such as "under-melting" and "burn-through."
[0141] Therefore, this ablation experiment fully demonstrates that the multimodal fusion strategy has significant advantages in improving model performance, enhancing working condition adaptability, and improving prediction robustness, and is one of the key technical paths to achieve high-precision prediction of the welding process.
[0142] 5. Conclusion
[0143] To address the challenge of predicting the molten pool penetration state in advance during TIG flat plate butt welding due to visual blurring and delayed response, this invention proposes a multimodal CNN-Transformer temporal classification model that integrates image and sound information for advanced prediction and discrimination of the molten pool state. This model effectively combines the local structure perception capabilities of CNNs with the advantages of Transformers in modeling long-term dependencies. By introducing welding sound as an auxiliary modality, it enhances the model's ability to perceive internal physical changes in the molten pool, thereby significantly improving prediction accuracy and foresight. Experimental results show that the model exhibits excellent performance across multiple time points, achieving an accuracy of 96.23% with a 0.5s lead time, providing a solid data foundation and methodological support for intelligent monitoring and process control of welding quality.
[0144] The main conclusions are as follows:
[0145] (1) In terms of model structure design, the multimodal model constructed in this invention uses DenseNet-SE201 as the backbone network to extract features of images and STFT sound respectively. By designing a cross-modal cross attention mechanism, the image branch can adaptively perceive and focus on key information regions in the sound modality. At the same time, dynamic feature weighted fusion is realized in the GatedFusion module, which enhances the multimodal collaborative modeling capability and enables the model to maintain stable prediction performance under complex welding conditions and multi-source interference.
[0146] (2) In the multi-time point prediction experiment, the model achieved high accuracy at different advance prediction time points from 0.05s to 2.5s, with a maximum of 96.23%, verifying that the model has good advance judgment ability and time series generalization. Among them, the short-term prediction accuracy is higher, while the long-term prediction accuracy is slightly lower, reflecting the increased uncertainty of the molten pool state over time.
[0147] (3) Through t-SNE embedding visualization analysis, it is shown that the three types of melt penetration states are clearly clustered in the high-dimensional feature space, which verifies that the model has good class discrimination ability. However, there is still some overlap in the state boundary region, which reveals that there is critical ambiguity and transition dynamics in the process of melt pool state evolution, suggesting that the model needs to further enhance its ability to capture fine-grained dynamic features.
[0148] In summary, the multimodal CNN-Transformer model proposed in this invention provides an effective and reliable solution for continuous monitoring and advanced prediction of the weld pool state, and is particularly suitable for controlling the ambiguous morphology of the weld pool in TIG welding.
Claims
1. A method for online monitoring of welding quality based on acoustic-optical combined advanced prediction, characterized in that, A welding acoustic and optical acquisition system is adopted, which includes an image acquisition device (1) and a sound acquisition device (2). The image acquisition device (1) includes an infrared LED array, a camera and a bandpass filter. The infrared LED array and the camera are both facing the welding area, and the bandpass filter is placed in front of the camera. The bandpass filter is a 940nm bandpass filter; The sound acquisition device (2) includes a microphone; Both the image acquisition device (1) and the sound acquisition device (2) are mounted on the welding torch robotic arm; The online welding quality monitoring method includes the following steps: 1) The welding acoustic-optical acquisition system is used to acquire the molten pool image sequence and sound sequence during the welding process. The sound sequence is converted into a spectrum diagram that is aligned with the time step of the molten pool image sequence through a short-time Fourier transform (STFT). According to the weld state, it is divided into incomplete penetration weld, normal penetration weld and burn-through weld. The corresponding molten pool image sequence, spectrum diagram and weld state are used to construct a welding process monitoring dataset. 2) Input the welding process monitoring dataset into the multimodal cross-attention fusion model to complete the training of the multimodal cross-attention fusion model; The multimodal cross-attention fusion model includes a CNN feature extraction module, an attention module, a gating fusion module, and a Transformer temporal prediction module. The CNN feature extraction module uses the melt pool image sequence and spectrogram to extract melt pool image features and spectrogram features, respectively. The attention module utilizes the features of the melt pool image to perform guided modeling of the spectrogram features, thereby obtaining cross-attention features; The gated fusion module is used to perform gated fusion between the melt pool image features and the cross-attention features to obtain a fused feature sequence. The fused feature sequence is input into the Transformer time-series prediction module for dynamic modeling, and the classification prediction result of the weld state is output. The CNN feature extraction module in step 2) includes two feature extraction modules. Each feature extraction module comprises a CBR fusion layer, a 3*3 max pooling layer, a densely connected convolutional network, and a global average pooling layer, connected sequentially. The CBR fusion layer includes a Convolutional layer (Conv), a BN layer, and a first ReLU layer. The densely connected convolutional network includes multiple DenseNet-SE modules connected sequentially. Each DenseNet-SE module includes a DenseBlock module, a Transition module, and an SE channel attention module. The DenseBlock module includes multiple DenseLayer layers, each consisting of a... The system consists of a first BatchNorm layer, a second ReLU layer, a first 1×1 convolutional layer, a second BatchNorm layer, a third ReLU layer, and a 3×3 convolutional layer. The Transition module includes a second 1×1 convolutional layer and a 2×2 average pooling layer. The SE channel attention module includes a global average pooling layer, a dimensionality-reduced 1×1 convolutional layer, a ReLU activation layer, an up-dimensional 1×1 convolutional layer, and a Sigmoid activation function connected in sequence. The global average pooling layer is used to extract channel descriptions, and the dimensionality-reduced 1×1 convolutional layer, ReLU activation layer, up-dimensional 1×1 convolutional layer, and Sigmoid activation layer are used to model the nonlinear relationship between channels and generate channel attention weights. The attention module includes a self-attention module, a cross-attention module, and a feedforward network, which are connected in sequence. The gated fusion module includes a channel splicing layer, two parallel gated networks, and a residual connection based on modal input, which are connected in sequence. The channel splicing layer is used to splice the features of the two modalities in the channel dimension. The two parallel gated networks perform nonlinear transformations on the spliced features to generate dynamic weights corresponding to the two modalities. The residual connection introduces the original modal features into the weighted fusion result. The Transformer temporal prediction module includes an input projection layer, a position encoding embedding layer, a Transformer encoder stack layer, and a classification output layer connected in sequence. The attention module mentioned in step 2) is represented by the following formula: , In the formula, The final output features represent the result of fusing self-attention, cross-attention, and feedforward networks. The input is the feature sequence of the main mode. The input auxiliary modality features are all of shape [B, T, D], where B is the batch size, T is the time series length, and D is the feature dimension. The representation of the master mode after it has completed information exchange within its own sequence. This represents the updated features of the main modality after cross-modal attention. , and These represent three LayerNormalization layers, with normalization performed before each sub-layer. This represents the self-attention mechanism within a modality, modeling the temporal relationships within the input sequence. This represents a cross-modal attention mechanism; The gating fusion module mentioned in step 2) is represented by the following formula: , In the formula, This represents the output characteristics of two modes after gated weight fusion and residual connection. The feature vectors are from the image channels. Features resulting from cross-modal interaction This represents the concatenated features used to generate the gating weights. This represents the Sigmoid activation function. , These represent the gating weights of the two modalities, controlling the dynamic contribution of their respective features. and These are the linear transformation parameters, i.e., the weight matrix of the MLP; The time series prediction module described in step 2) is represented by the following formula: , Where, x∈R B×T×D Given the input sequence Z∈R B×3 This represents the class probability of the three welding states. This represents the position index number, ranging from [0, T]. The extra position is used to categorize the token. This indicates the position embedding module, which provides temporal information about the sequence. Indicates the input linear projection layer. Used to prevent overfitting It consists of several stacked TransformerEncoderLayer layers, performing multi-head self-attention and feedforward network encoding; 3) Input the image sequence and sound sequence of the molten pool during the welding process into the multimodal cross-attention fusion model trained in step 2) to achieve classification and prediction of the weld state.
2. The online welding quality monitoring method based on combined acoustic and optical prediction as described in claim 1, characterized in that, The angle between the microphone and the welding torch is 75°.