Real-Time Spatial Attention Module and Its Application in Weld Nugget Quality Detection
By introducing a real space-time attention module to the quality detection of welding joint melting cores, the real space-time correlation characteristics of multi-channel signals are deeply explored, and the problems of low efficiency and low automation of existing detection methods are solved, and high-precision and intelligent solder melting core quality detection is achieved.
Patent Information
- Application Number
- CN202310137106.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-02-20
AI Technical Summary
The existing welding joint melting core quality detection methods have problems such as high time and labor consumption, waste of resources, low detection efficiency and low degree of automation, and it is difficult to meet the development needs of high-speed automation in the automobile manufacturing industry.
A real space-time attention module is proposed, through the space-time and time-processing module, the real space-time correlation characteristics of multi-channel signals are deeply explored, and the detection accuracy of the intelligent model is improved. The module includes a real spatial attention unit and a global temporal context attention unit, which can effectively process the spatial and temporal information of multi-channel signals.
Through the application of the real space-time attention module, the accuracy and efficiency of welding joint melting core quality detection are significantly improved, intelligent, efficient and high-precision detection is achieved, and high-speed automation needs in the automobile manufacturing industry are met.
Smart Images

Figure CN116050477B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent manufacturing, and specifically relates to a real-time and space attention module and its application in the detection of the nugget quality of solder joints. Background Art
[0002] Resistance spot welding is a resistance welding process applied to welding various metal sheet parts. Due to its advantages of small workpiece deformation, high welding efficiency, low process cost, simple operation during the welding process, and recyclability, it is widely used in the welding manufacturing of white body in high-speed automation. According to statistics, more than 90% of the assembly volume of a typical white body is completed by resistance spot welding. The quality of solder joints greatly affects the quality of the white body, especially the nugget quality of solder joints directly determines the mechanical strength of solder joints, and further determines the safety performance and service life of the white body. Therefore, the detection of the nugget quality of solder joints is crucial for controlling the processing quality of the white body.
[0003] In current actual engineering, the detection of the nugget quality of solder joints mostly adopts the strategy of manual sampling inspection, and uses destructive chiseling inspection and ultrasonic technology to sample and detect solder joints. Destructive chiseling inspection directly observes the nugget structure by destroying the solder joint connection as the basis for quality discrimination; ultrasonic detection uses the pulse waveform formed by the reflection of sound waves at the boundary surfaces of different structures as the basis for quality discrimination. Although these two detection methods have been widely used in the field of industrial manufacturing, they also have obvious limitations: destructive chiseling inspection requires workers to manually chisel the solder joints, which consumes a large amount of time and labor during the chiseling process, and the destruction of samples will cause waste of resources; ultrasonic detection requires professional technicians to detect each solder joint one by one, the process is cumbersome and it is easy to have repetitions or omissions during the detection process. These limitations result in low sampling rates, low detection efficiency, and low automation levels of these two methods, and it is difficult to meet the development needs of high-speed automation in the current automotive manufacturing industry.
[0004] In recent years, with the development of industrial big data and computer technology, the data accumulation and calculation speed have been continuously improved. Some intelligent models based on industrial signals have been studied and applied in the field of solder joint quality detection. For example, the non-linear implicit approximate functional solution between welding current and nugget diameter is mined through the universal approximation characteristic of the BP neural network; the potential correlation between dynamic resistance signals and the mechanical strength of solder joints is explored using a recurrent neural network; the transient dynamic behavior of the physical field during the welding force-thermal coupling process is simulated through a residual or dense network structure; the appearance image features of solder joints are deeply mined through a convolutional neural network and accurately corresponding to the appearance quality state of solder joints, etc. These studies have all achieved preliminary success, bringing new opportunities and inspirations to the theory and method of solder joint quality detection based on industrial big data, and also bringing the possibility of the implementation and application of a new generation of convenient and efficient solder joint quality detection technology.
[0005] The core algorithm of the intelligent detection model can be designed specifically according to the characteristics of the measurement signal. Through neural network layers with different functions in multiple layers, data operations such as segmentation, aggregation, dimension elevation and reduction of the original signal are performed, and high-dimensional semantic features of the signal are extracted layer by layer as the basis for subsequent linear regression or nonlinear classification. In the model training process, the gradient of the output result is backpropagated, and the optimizer of machine learning is used to iteratively optimize the model parameters, so that the output result of the model continuously approaches the real situation, and finally the implicit nonlinear relationship between the vibration response signal and the health state of the mechanical system is learned. The signal analysis and processing process of the model does not require artificial feature engineering preprocessing, and can improve the calculation speed through GPU parallel computing, meet the development needs of high-speed automation in the automotive manufacturing industry, and realize intelligent, efficient and high-precision detection of the nugget quality of solder joints. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a real spatio-temporal attention module and its application in the detection of nugget quality of solder joints. The real spatio-temporal attention module deeply mines the real spatio-temporal correlation features of multi-channel signals, can effectively improve the detection accuracy of the intelligent model, has strong versatility, can be applied to various industrial fields to process different signals, and can be flexibly embedded into various models during application to improve performance.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] The present invention first proposes a real spatio-temporal attention module, which includes a spatial processing module and a temporal processing module;
[0009] The spatial processing module includes a real spatial attention unit for mining the potential real spatial correlation information of multi-channel data. The real spatial attention unit includes a global average pooling layer, a global standard deviation pooling layer, a multi-layer perceptron and a residual connection block; the global average pooling layer and the global standard deviation pooling layer are respectively used to encode the input multi-channel signal and obtain two sets of aggregated features. The two sets of aggregated features respectively use the multi-layer perceptron to mine the hidden spatial relationship of each channel of data and generate adaptive spatial weight vectors respectively. After the two spatial weight vectors are added, the Softmax function is used for activation to obtain the spatial attention weight vector. The spatial attention weight vector corrects the original feature map through the residual connection block and obtains the real spatial feature map;
[0010] The temporal processing module includes a parallel temporal short sequence attention unit and a global temporal context attention unit;
[0011] The temporal short - sequence attention unit includes two branches for performing operations at different spatial scales respectively; the first branch sequentially performs average pooling and downsampling operations on the input real - space feature map. After average pooling, irregular padding is used to keep the size of the pooled feature map unchanged. After downsampling, linear interpolation is performed on the obtained hidden - space feature map to expand its size to the original spatial size, obtaining an interpolated feature map; the interpolated feature map and the real - space feature map are element - wise summed using residual connection and an activation function is used to generate a temporal short - sequence attention weight matrix; the second branch performs a convolution operation on the real - space feature map at the original spatial scale to extract the feature information of the real - space feature map and generate an original - scale feature map; the original - scale feature map is corrected using the temporal short - sequence attention weight matrix to obtain the output of the temporal short - sequence attention unit.
[0012] The global - time context attention unit includes an information - collection branch and an information - distribution branch arranged in parallel; the global - time context attention unit respectively compresses the number of channels of the real - space feature map to half of the original using two convolution operations, and respectively obtains an information - collection feature map and an information - distribution feature map; the information - collection branch sequentially performs a convolution operation and a reshape operation on the information - collection feature map; after the convolution operation, the number of channels of the information - collection feature map expands to be equal to the product of its height and width, obtaining a first global - time attention matrix; the reshape operation is performed to adjust the size of the first global - time attention matrix, and the Hadamard product is used to distribute the weights of the mutual influence between signals to the information - collection feature map to correct the information - collection feature map, obtaining the output result of the information - collection branch; the information - distribution branch sequentially performs a convolution operation and a reshape operation on the information - distribution feature map; after the convolution operation, the number of channels of the information - distribution feature map expands to be equal to the product of its height and width, obtaining a second global - time attention matrix; the reshape operation is performed to adjust the size of the second global - time attention matrix, and the Hadamard product is used to distribute the weights of the mutual influence between signals to the information - distribution feature map to correct the information - distribution feature map, obtaining the output result of the information - distribution branch;
[0013] The output of the global - time context attention unit is obtained by concatenating the output results of the information - collection branch and the information - distribution branch along the channel dimension; the output of the time - processing module is obtained by concatenating the output of the temporal short - sequence attention unit and the output of the global - time context attention unit along the channel dimension.
[0014] Furthermore, the principle of the real - space attention unit is:
[0015] The global average pooling layer and the global standard deviation pooling layer respectively encode the input multi-channel signals and obtain two sets of aggregated features:
[0016]
[0017]
[0018] Among them, x m represents the feature map of the m-th channel, represents a three-dimensional tensor, H represents the height, W represents the width, and C represents the channel; respectively represent the aggregated features obtained by global average pooling and global standard deviation pooling; represents the i-th data of the c-th channel feature map; represents the average value of all data in the c-th channel; GAP(·) and GSP(·) respectively represent global average pooling and global standard deviation pooling;
[0019] The two sets of aggregated features respectively mine the spatial relationships hidden in the data of each channel through the multi-layer perceptron and respectively generate adaptive spatial weight vectors. After adding the two spatial weight vectors, the Softmax function is used for activation to obtain the spatial attention weight vector:
[0020]
[0021] Among them, w c represents c the weight of the channel; and respectively represent and the spatial weight vectors generated by the multi-layer perceptron; σ(·) represents the Softmax function; MLP(·) represents the multi-layer perceptron;
[0022] Introduce a residual connection, perform the Hadamard product on the weight of the corresponding channel and the feature map of the corresponding channel to correct the signals of different channels and obtain the real spatial feature map:
[0023]
[0024] Among them, represents c the real spatial feature map of the channel.
[0025] Furthermore, the principle of the first branch of the temporal short sequence attention unit is:
[0026] For the input real spatial feature map First, average pooling is performed on the original feature map using a 1×k kernel to aggregate short-range features, and irregular padding is used to ensure that the size of the feature map after pooling remains unchanged. Then, a 1×k convolutional kernel is used to further extract high-level features, resulting in a hidden space feature map:
[0027]
[0028] W′ = (W - k) + 1
[0029] where h represents the hidden space feature map; h m represents the feature map of the m-th channel in the hidden space feature map;
[0030] Interpolate linearly along the W′ dimension to the original spatial size uniformly to obtain an interpolated feature map:
[0031]
[0032] where Y represents the interpolated feature map; y m represents the feature map of the m-th channel in the interpolated feature map;
[0033] Use residual connection to perform element-wise summation of the interpolated feature map and the real space feature map, embed the correction value extracted from the hidden space into the real space feature map, and use an activation function to generate a short temporal sequence attention weight matrix:
[0034]
[0035] where w represents the short temporal sequence attention weight matrix; σ(·) represents the Sigmoid activation function.
[0036] Furthermore, the principle of the second branch of the short temporal sequence attention unit is as follows:
[0037] The second branch performs a 1×1 convolution operation on the real space feature map at the original spatial scale to extract the feature information of the original input and generate an original scale feature map
[0038]
[0039] where k represents the size of the convolutional kernel; ω represents the convolutional kernel weight vector, p represents the p-th element in the real space feature map and s represents the -th element in the original scale feature map obtained by convolution s
[0040] Furthermore, the output of the short temporal sequence attention unit is:
[0041]
[0042] Among them, Z 1 represents the output of the short-term sequential attention unit; represents the Hadamard product.
[0043] Furthermore, the global temporal context attention unit respectively uses two 1×1 convolutional operations to compress the channels of the real space feature map to half of the original, and respectively obtains the information acquisition feature map and the information distribution feature map
[0044] Furthermore, the information acquisition branch performs 1×1 convolution on the information acquisition feature map to expand the number of channels to H×W, obtaining the first global temporal attention matrix w 1 , then for signal i, the value at position i on the j-th channel feature map represents the influence degree of signal j on signal i
[0045]
[0046] Among them, Conv 1×1 represents the 1×1 convolution operation; represents the information acquisition feature map;
[0047] Performs a reshape operation to adjust the size of the first global temporal attention matrix, and uses the Hadamard product to distribute the weights of the mutual influence between signals to the information acquisition feature map to correct the information acquisition feature map, obtaining the output result of the information acquisition branch and;
[0048]
[0049] Among them, represents the matrix composed of all information acquisition feature maps.
[0050] Furthermore, the information distribution branch performs 1×1 convolution on the information distribution feature map to expand the number of channels to H×W, obtaining the second global temporal attention matrix w 2 , then for signal i, the value at position i on the j-th channel feature map represents the influence degree of signal j on signal i
[0051]
[0052] Among them, Conv 1×1 represents the 1×1 convolution operation; represents the information distribution feature map;
[0053] Perform a reshape operation to adjust the size of the second global temporal attention matrix, and use the Hadamard product to assign the weights of the mutual influence between signals to the information allocation feature map to correct the information allocation feature map, obtaining the output result of the information allocation branch And;
[0054]
[0055] Wherein, represents the matrix composed of all information allocation feature maps.
[0056] Furthermore, the output of the global temporal context attention unit is:
[0057]
[0058] Wherein, Z 2 represents the output of the global temporal context attention unit; F 1 represents the output of the information acquisition branch; F 2 represents the output of the information allocation branch; represents the concatenation operation along the channel dimension;
[0059] The output of the temporal processing module is:
[0060]
[0061] Wherein, Z represents the output of the temporal processing module; Z 1 represents the output of the temporal short sequence attention unit.
[0062] The present invention also proposes an application of the real spatio-temporal attention module in the detection of the quality of solder joint fusion cores. The real spatio-temporal attention module as described above is embedded in a convolutional neural network to construct an intelligent detection model, and the intelligent detection model is used to detect the quality of solder joint fusion cores.
[0063] The beneficial effects of the present invention are as follows:
[0064] The real spatiotemporal attention module of the present invention utilizes the real space attention unit included in the space processing module to mine the potential real space correlation information of multi-channel data, encodes the input multi-channel signal through global average pooling GAP and global standard deviation pooling GSP, generates two sets of aggregated features through shared MLP processing, and finally corrects the original feature map through residual connection; the real space feature map output by the space processing module is used as the input of the time processing module, and the time series short sequence attention unit is used to perform convolution operations in two spaces of different scales, and the convolution operation of the original scale space keeps the original features of the signal as much as possible; the downsampled small scale space deeply mines the correlation between short time series data through convolution kernels of special sizes, and embeds it back into the real space feature map through interpolation operations. row correction; the global time context attention unit integrates all global information, realizes the interaction of all information by collecting and distributing multiple convolution and reshape operations of the two branches, and obtains the information interaction matrix, and finally uses the information interaction matrix to re-correct the real space feature map to obtain the output feature map; then the outputs of the temporal short sequence attention unit and the global time context attention unit are spliced to obtain the output of the time processing module, which is the output of the real space-time attention module of the present invention; in summary, the real space-time attention module of the present invention deeply mines the real space-time correlation characteristics of multi-channel signals, which can effectively improve the detection accuracy of the intelligent model, and has strong versatility, can be applied to a variety of industrial fields to process different signals, and can also flexibly embed a variety of models to improve performance when applied. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to make the purpose, technical solution and beneficial effects of the present invention clearer, the present invention provides the following drawings for illustration:
[0066] Figure 1 A flow chart of the application of the real spatiotemporal attention module of the present invention in the weld nugget quality detection;
[0067] Figure 2 It is a structural principle diagram of an embodiment of the real spatiotemporal attention module of the present invention;
[0068] Figure 3 This is the structural schematic diagram of the real spatial attention unit;
[0069] Figure 4 This is the structural principle diagram of the time-series short sequence attention unit;
[0070] Figure 5 The structural schematic diagram of the global temporal context attention unit;
[0071] Figure 6 This is the structural schematic diagram of the TSA module embedded in ResNet for solder nugget quality detection;
[0072] Figure 7 It is the confusion matrix obtained by validating the ResNet+TSA intelligent detection model on the solder joint nugget quality data set. Specific implementation manners
[0073] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited are not intended to limit the present invention.
[0074] As Figure 1 shown, the application of the real spatio-temporal attention module in the solder joint nugget quality detection in this embodiment takes the front hood of a white vehicle body as an example. A plurality of sensors are arranged around the solder joint to collect vibration response signals, and at the same time, the true quality results of the solder joint are obtained. A data set is constructed with the vibration response signals and quality true results of the solder joint. At the same time, the real spatio-temporal attention module (TSA) is embedded into the convolutional neural network to construct an intelligent detection model. The intelligent detection model is trained with the data set, and the trained intelligent detection model is used to measure the quality of the solder joint nugget to realize the intelligent detection of the quality of the solder joint nugget. Specifically, the real spatio-temporal attention module is the core of the intelligent detection model, and its principle is as follows.
[0075] I. Real spatio-temporal attention module
[0076] The real spatio-temporal attention module (TSA) deeply mines the potential correlation features of multi-channel input signals from both spatial and temporal perspectives. As Figure 2 shown, the real spatio-temporal attention module in this embodiment includes a spatial processing module and a temporal processing module. The spatial processing module in this embodiment includes a real spatial attention unit for mining the potential real spatial correlation information of multi-channel data. The temporal processing module in this embodiment includes a parallel short-time sequence attention unit and a global temporal context attention unit.
[0077] 1.1. Real spatial attention unit
[0078] Process parameters are diverse. Different process parameters reflect the working conditions of the system from different angles, and the same process parameters reflect the working conditions of the system to different degrees. Taking the multi-channel solder joint vibration response signals in this embodiment as an example, the multi-channel data is collected by sensors at different positions. Although the signal properties are the same, all are vibration response signals, but due to the different positions of the sensors, the response degrees of each channel signal to the quality of the solder joint nugget are different. The sensors closer to the solder joint have stronger response capabilities to the nugget quality, and the signals collected by the sensors far from the solder joint will reduce the response capabilities to the nugget quality due to factors such as energy loss during transmission. Therefore, it is necessary to extract the potential real spatial relationships of each channel data and assign reasonable weights to represent the expression capabilities of different channels for the nugget quality.
[0079] As shown Figure 3 in the figure, the real - space attention unit of this embodiment includes a global average pooling layer, a global standard - deviation pooling layer, a multi - layer perceptron, and a residual connection block; the global average pooling layer and the global standard - deviation pooling layer are respectively used to encode the input multi - channel signal and obtain two sets of aggregated features. The two sets of aggregated features respectively use the multi - layer perceptron to mine the spatial relationships hidden in each channel's data and generate adaptive spatial weight vectors. After adding the two spatial weight vectors, the Softmax function is used for activation to obtain the spatial attention weight vector, and the spatial attention weight vector corrects the original feature map through the residual connection block to obtain the real - space feature map.
[0080] Specifically, in this embodiment, for the feature map x m representing the feature map of the m - th channel, where represents a three - dimensional tensor, H represents height, W represents width, and C represents channels. Considering that the original data is vibration signals with both positive and negative values, using traditional global average pooling and global max pooling will lose all negative - value data. In view of this, this embodiment uses the data distribution and dispersion degree of each channel as the parameters for feature aggregation, and designs a hybrid pooling method of global average pooling GAP plus global standard - deviation pooling GSP to obtain two sets of aggregated features of the global information within each channel. That is, the global average pooling layer and the global standard - deviation pooling layer respectively encode the input multi - channel signal and obtain two sets of aggregated features:
[0081]
[0082]
[0083] Among them, x m represents the feature map of the m - th channel, represents a three - dimensional tensor, H represents height, W represents width, and C represents channels; respectively represent the aggregated features obtained by global average pooling and global standard - deviation pooling; represents the i - th data of the c - th channel's feature map; represents the average value of all data in the c - th channel; GAP(·) and GSP(·) respectively represent global average pooling and global standard - deviation pooling.
[0084] The two sets of aggregated features respectively use the multi - layer perceptron to mine the spatial relationships hidden in each channel's data and generate adaptive spatial weight vectors. After adding the two spatial weight vectors, the Softmax function is used for activation to obtain the spatial attention weight vector:
[0085]
[0086] Among them, w c represents the weight of the c-channel; and respectively represent and the spatial weight vectors generated by the multi-layer perceptron; σ(·) represents the Softmax function; MLP(·) represents the multi-layer perceptron.
[0087] To avoid overcorrection, a residual connection is introduced to perform the Hadamard product on the weight w of the corresponding channel c and the feature map x of the corresponding channel c to correct the signals of different channels and obtain the real spatial feature map:
[0088]
[0089] Among them, represents the real spatial feature map of the c-channel.
[0090] 1.2. Temporal Short Sequence Attention Unit
[0091] The vibration response signal used in this example is a typical temporal periodic signal. The vibration change trend within a short temporal range contains the information of the solder joint nugget quality. Extracting the potential features of these change trends helps the model learn more feature representations to improve the accuracy of the detection results. However, excessive focus on short-range information will lose the overall grasp of the entire temporal data. In view of this, a data embedding method is used to perform a relatively gentle correction on the real spatial feature map.
[0092] As Figure 4 shown, the temporal short sequence attention unit of this embodiment includes two branches for performing operations at different spatial scales respectively; the first branch sequentially performs average pooling and downsampling operations on the input real spatial feature map. After average pooling, irregular padding is used to keep the size of the pooled feature map unchanged. After downsampling, linear interpolation is performed on the obtained hidden spatial feature map to expand the size to the original spatial size to obtain the interpolated feature map; the residual connection is used to perform element-wise summation of the interpolated feature map and the real spatial feature map and an activation function is used to generate the temporal short sequence attention weight matrix. The second branch performs a convolution operation on the real spatial feature map at the original spatial scale to extract the feature information of the real spatial feature map and generate the original scale feature map; the original scale feature map is corrected using the temporal short sequence attention weight matrix to obtain the output of the temporal short sequence attention unit;
[0093] Specifically, the principle of the first branch of the temporal short sequence attention unit is as follows:
[0094] For the input real-space feature map First, use a 1×k kernel to perform average pooling on the original feature map to aggregate short-range features, and adopt irregular padding to ensure that the size of the feature map after pooling remains unchanged. Then, use a 1×k convolutional kernel to further extract high-level features without using padding, obtaining the hidden-space feature map:
[0095]
[0096] W′ = (W - k) + 1
[0097] where h represents the hidden-space feature map; h m represents the feature map of the m-th channel in the hidden-space feature map.
[0098] At this time, the height H is the same as the size of the original feature map, while the width W is compressed to W′ due to the uneven kernel convolution operation. Since the size of the hidden-space feature map has changed and cannot be directly embedded, it is uniformly linearly interpolated along the W′ dimension to the original space size to obtain the interpolated feature map:
[0099]
[0100] where Y represents the interpolated feature map; y m represents the feature map of the m-th channel in the interpolated feature map;
[0101] Use residual connection to perform element-wise summation of the interpolated feature map and the real-space feature map, embed the correction value extracted from the hidden space into the real-space feature map, and use an activation function to generate the temporal short-sequence attention weight matrix:
[0102]
[0103] where w represents the temporal short-sequence attention weight matrix; σ(·) represents the Sigmoid activation function.
[0104] The principle of the second branch of the temporal short-sequence attention unit is as follows:
[0105] The second branch performs a 1×1 convolution operation on the real-space feature map at the original space scale to extract the feature information of the original input and generate the original-scale feature map
[0106]
[0107] where k represents the size of the convolutional kernel; ω represents the convolutional kernel weight vector, p represents the p-th element in the real-space feature map and s represents the s-th element in the original-scale feature map obtained by convolution in the original-scale feature map.
[0108] Use the short - term time - series attention weight matrix \(w\) to correct the original scale feature map Obtain the final output of the short - term time - series attention unit:
[0109]
[0110] Among them, \(Z\) 1 represents the output of the short - term time - series attention unit; represents the Hadamard product.
[0111] 1.3. Global time - context attention unit
[0112] The structure of the global time - context attention unit is as Figure 5 shown. The global time - context attention unit adopts a two - branch design. The two branches have the same structure and perform acquisition and distribution tasks respectively, that is, they are used to extract the influence of other signals on signal \(i\) and the influence of signal \(i\) on other signals respectively.
[0113] To ensure the fairness of information acquisition and distribution and reduce the computational overhead as much as possible, the feature map input to the global time - context attention unit is compressed by two 1×1 convolutions to half of the original number of channels, obtaining the information acquisition feature map and the information distribution feature map and are fed into the information acquisition branch and the information distribution branch respectively.
[0114] Specifically, the information acquisition branch performs a 1×1 convolution on the information acquisition feature map to expand the number of channels to \(H\times W\), obtaining the first global time attention matrix \(w\) 1 , and at this time, the channels represent the degree of influence between each point. Specifically, for signal \(i\), the value at position \(i\) on the \(j\) - th channel feature map represents the degree of influence of signal \(j\) on signal \(i\)
[0115]
[0116] Among them, \(Conv\) 1×1 represents the 1×1 convolution operation; represents the information acquisition feature map;
[0117] Then perform a reshape operation to adjust the size of the first global time attention matrix, and use the Hadamard product to distribute the weights of the mutual influence between signals to the information acquisition feature map to correct the information acquisition feature map, obtaining the output result of the information acquisition branch and;
[0118]
[0119] Among them, represents the matrix composed of all information acquisition feature maps.
[0120] The structure and process of the information distribution branch are exactly the same as those of the information acquisition branch. Specifically, the information distribution branch performs 1×1 convolution on the information distribution feature map to expand the number of channels to H×W, obtaining the second global temporal attention matrix w 2 , then for signal i, the value at position i on the j-th channel feature map represents the influence degree of signal j on signal i
[0121]
[0122] Among them, Conv 1×1 represents the 1×1 convolution operation; represents the information distribution feature map;
[0123] Perform a reshape operation to adjust the size of the second global temporal attention matrix, and use the Hadamard product to distribute the weights of the mutual influence between signals to the information distribution feature map to correct the information distribution feature map, obtaining the output result of the information distribution branch and;
[0124]
[0125] Among them, represents the matrix composed of all information distribution feature maps.
[0126] Finally, the corrected feature maps of the information acquisition branch and the information distribution branch are concatenated along the channel dimension to obtain the output of the global temporal context attention unit as:
[0127]
[0128] Among them, Z 2 represents the output of the global temporal context attention unit; F 1 represents the output of the information acquisition branch; F 2 represents the output of the information distribution branch; represents the concatenation operation along the channel dimension.
[0129] This two-way mapping operation of acquisition and distribution can complementarily encode the correlation information at different positions, fully mine the hidden features between the global signal remote contexts, and achieve global temporal context attention.
[0130] After the parallel temporal short sequence attention unit and the global temporal context attention unit are processed, the attention maps Z 1 and Z2 , the two are concatenated along the channel dimension to obtain the final output result of the time processing module, which is also the final output result of the real spatio-temporal attention module:
[0131]
[0132] Among them, Z represents the output of the time processing module; Z 1 represents the output of the short-time sequence attention unit for time series.
[0133] 1.4. Application
[0134] The application of the real spatio-temporal attention module in this embodiment in the detection of the quality of the welding nugget is to embed the real spatio-temporal attention module into the convolutional neural network to construct an intelligent detection model, so as to realize the intelligent detection of the quality of the welding nugget. As Figure 6 shown, in this example, the Residual Network ResNet is selected as the basic model, and the real spatio-temporal attention module is embedded before the residual block of ResNet for pre-processing feature extraction.
[0135] II. Experimental Verification
[0136] To fully test the performance of the real spatio-temporal attention module (TSA), based on the welding nugget quality dataset (RSW), multiple groups of experiments were designed to test the overall performance of the module, and finally the generality of TSA was tested based on the publicly available dataset of Northeastern University SEU.
[0137] 2.1. Dataset and Experimental Settings
[0138] The welding nugget quality dataset (RSW) used in the experiment was made based on the multi-channel vibration response signals of the welding nugget quality, including 2800 groups of samples for each of the five categories of weld bead, burn-through, lack of fusion, crack and normal. The length of each sample is 1024, corresponding to a real time of 1 second, and the training set, test set and validation set are divided according to the ratio of 5:1:1.
[0139] The publicly available dataset of Northeastern University SEU consists of two sub-datasets, namely the gearbox fault dataset SEU-G and the bearing fault dataset SEU-B. The gearbox fault dataset SEU-G includes five categories: normal, missing teeth, tooth surface wear, tooth root crack and tooth profile damage; the bearing fault dataset SEU-B includes five categories: normal, roller wear, outer ring wear, inner ring wear and compound wear. The data is segmented using a sliding window of size 1024 to obtain 14000 groups of samples, and the training set, test set and validation set are divided according to the ratio of 5:1:1.
[0140] During the training process of the intelligent detection model, the Batch size was set to 128, the Learning rate was set to 0.01, the model weights were initialized using Kaiming and the weight decay was set to 0.0005 to avoid model overfitting. The CrossEntropy loss function and the Adam optimization algorithm were used to iteratively optimize the model, and the number of training iterations epochs was set to 100 times.
[0141] The intelligent detection neural network used in the experiment was written based on the PyTorch deep learning framework and ran on a server configured with an Nvidia RTX 3080 GPU.
[0142] 2.2. Experimental Results
[0143] Traditional ResNet, DenseNet, VGG, and AlexNet were used as the base models, and TSA was embedded for comparative experiments on the RSW dataset. The experimental results are shown in Table 1.
[0144] Table 1 TSA Overall Performance
[0145]
[0146] In Table 1, "-" in the TSA row indicates that the TSA module is not embedded in the model, and the "+" sign indicates that the TSA module is embedded in the model. It can be seen that embedding the TSA module can improve the accuracy of the original model to varying degrees. Among them, the ResNet model with the embedded TSA module has the highest accuracy, reaching 96.35%. The confusion matrix obtained by validating the ResNet+TSA intelligent detection model on the solder joint nugget quality dataset is shown in Figure 7 . And higher accuracies can also be obtained after embedding the TSA module in other classic models, fully demonstrating the performance of the TSA module.
[0147] To verify the generality of TSA, the above four traditional models were still used as the base models. Each model was divided into two types: embedded with TSA and not embedded with TSA, and then comparative experiments were conducted on the SEU-G and SEU-B datasets. The experimental results are shown in Table 2.
[0148] Table 2 TSA General Performance
[0149]
[0150] It can be clearly seen from Table 2 that all base models achieved higher accuracies after embedding TSA on the SEU-G and SEU-B datasets, fully demonstrating the good generality of the TSA module.
[0151] III. Conclusion
[0152] The TSA module designed for the intelligent detection task of the fusion core quality of solder joints can effectively improve the performance of the basic model. The prediction accuracy of the ResNet model embedded with the TSA module reaches 96.35%, and it has a small number of parameters. It can be deployed on the enterprise production line to achieve accurate and rapid detection of the fusion core quality of solder joints, improve the degree of automation, and reduce labor costs. In addition, the TSA module has good versatility and can be flexibly embedded into different basic models and used in other fields.
[0153] The above-described embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.
Claims
1. A method for detecting the nugget quality of solder joints based on a real spatio-temporal attention module, characterized in that: A plurality of sensors are arranged around the solder joint to collect vibration response signals, and at the same time, the true quality result of the solder joint is obtained. A data set is constructed with the vibration response signal and the true quality result of the solder joint. At the same time, the real spatio-temporal attention module is embedded into the convolutional neural network to construct an intelligent detection model. The intelligent detection model is trained with the data set, and the nugget quality of the solder joint is measured by using the trained intelligent detection model to realize the intelligent detection of the nugget quality of the solder joint; The real spatio-temporal attention module includes a spatial processing module and a temporal processing module; The spatial processing module includes a real spatial attention unit for mining the potential real spatial correlation information of multi-channel data. The real spatial attention unit includes a global average pooling layer, a global standard deviation pooling layer, a multi-layer perceptron, and a residual connection block. The global average pooling layer and the global standard deviation pooling layer are respectively used to encode the input multi-channel signals and obtain two sets of aggregated features. The two sets of aggregated features respectively use the multi-layer perceptron to mine the spatial relationships hidden in each channel of data and generate adaptive spatial weight vectors. After the two spatial weight vectors are added, the Softmax function is used for activation to obtain a spatial attention weight vector. The spatial attention weight vector corrects the original feature map through the residual connection block and obtains a real spatial feature map; The temporal processing module includes a parallel temporal short sequence attention unit and a global temporal context attention unit; The temporal short sequence attention unit includes two branches for performing operations at different spatial scales. The first branch sequentially performs average pooling and downsampling operations on the input real spatial feature map. After average pooling, irregular padding is used to keep the size of the pooled feature map unchanged. After downsampling, linear interpolation is performed on the obtained hidden spatial feature map to expand the size to the original spatial size to obtain an interpolated feature map. The interpolated feature map and the real spatial feature map are element-wise summed using a residual connection and an activation function is used to generate a temporal short sequence attention weight matrix; The second branch performs a convolution operation on the real spatial feature map at the original spatial scale to extract the feature information of the real spatial feature map and generate an original scale feature map; The original scale feature map is corrected using the temporal short sequence attention weight matrix to obtain the output of the temporal short sequence attention unit; The global temporal context attention unit includes an information collection branch and an information allocation branch arranged in parallel; the global temporal context attention unit uses two convolution operations to compress the number of channels of the real space feature map to half of the original number, and obtains an information collection feature map and an information allocation feature map respectively; the information collection branch performs convolution operations and reshape operations on the information collection feature map in turn; after performing the convolution operation, the number of channels of the information collection feature map is expanded to be equal to the product of its height and width, and a first global temporal attention matrix is obtained; a reshape operation is performed to adjust the size of the first global temporal attention matrix, and Hadamard The information allocation branch performs convolution and reshape operations on the information allocation feature map in turn; after the convolution operation, the number of channels of the information allocation feature map is expanded to be equal to the product of its height and width to obtain a second global time attention matrix; the reshape operation is performed to adjust the size of the second global time attention matrix, and the Hadamard product is used to allocate the weights of the mutual influence between the signals to the information allocation feature map to modify the information allocation feature map, and the output result of the information allocation branch is obtained; The output of the information acquisition branch and the output of the information allocation branch are spliced along the channel dimension to obtain the output of the global time context attention unit; the output of the temporal short sequence attention unit and the output of the global time context attention unit are spliced along the channel dimension to obtain the output of the time processing module.
2. According to the method for detecting weld nugget quality based on real spatiotemporal attention module of claim 1, Features: The principle of the real spatial attention unit is: The global average pooling layer and the global standard deviation pooling layer encode the input multi-channel signal and obtain two sets of aggregated features: Among them, x m represents the feature map of the m-th channel, represents a three-dimensional tensor, H represents height, W represents width, and C represents channels; respectively represent the aggregated features obtained by global average pooling and global standard deviation pooling; represents the i-th data of the feature map of the c-th channel; represents the average value of all data in the c-th channel; GAP(·) and GSP(·) respectively represent global average pooling and global standard deviation pooling; The two sets of aggregated features are respectively used to mine the hidden spatial relationship of each channel data through the multi-layer perceptron and generate adaptive spatial weight vectors respectively. After adding the two spatial weight vectors, the Softmax function is used to activate and obtain the spatial attention weight vector: Among them, w c represents c the weight of the channel; and respectively represent and the spatial weight vectors generated by the multi-layer perceptron; σ(·) represents the Softmax function; MLP(·) represents the multi-layer perceptron; Residual connection is introduced to perform Hadamard product on the weight of the corresponding channel and the feature map of the corresponding channel to correct the signals of different channels and obtain the real space feature map: Among them, represents c the real-space feature map of the channel.
3. According to the solder nugget quality detection method based on real spatiotemporal attention module of claim 1, Features: The principle of the first branch of the temporal short sequence attention unit is: For the input real-space feature map First, use a 1×k kernel to perform average pooling on the original feature map to aggregate short-range features, and use irregular padding to ensure that the size of the feature map after pooling remains unchanged; then use a 1×k convolutional kernel to further extract high-level features to obtain the hidden-space feature map: W′=(Wk)+1 Among them, h represents the hidden space feature map; h m represents the feature map of the m-th channel in the hidden space feature map; Uniformly linearly interpolate along the W′ dimension to the original spatial size to obtain the interpolated feature map: Among them, Y represents the interpolated feature map; y m represents the feature map of the m-th channel in the interpolated feature map; The interpolated feature map is element-wise summed with the real space feature map using the residual connection, the correction value extracted from the hidden space is embedded into the real space feature map, and the activation function is used to generate the time series short sequence attention weight matrix: Among them, w represents the attention weight matrix of the short time series; σ(·) represents the Sigmoid activation function.
4. The method for detecting the quality of the solder joint fusion core based on the real spatio-temporal attention module according to claim 3, characterized in that: The principle of the second branch of the short time series attention unit is: The second branch performs a 1×1 convolution operation on the real space feature map at the original spatial scale to extract the feature information of the original input and generate the original scale feature map Among them, k represents the size of the convolutional kernel; ω represents the convolutional kernel weight vector, and p represents the p-th element in the real-space feature map ; s represents the original-scale feature map obtained by convolution and the s -th element in it.
5. The method for detecting the quality of the solder joint fusion core based on the real spatio-temporal attention module according to claim 3, characterized in that: The output of the short time series attention unit is: Among them, Z 1 represents the output of the temporal short sequence attention unit; represents the Hadamard product.
6. The method for detecting the quality of the solder joint fusion core based on the real spatio-temporal attention module according to claim 5, characterized in that: The global time context attention unit respectively uses two 1×1 convolution operations to compress the channels of the real space feature map to half of the original, and respectively obtains an information acquisition feature map and an information distribution feature map 7. The method for detecting the quality of the solder joint fusion core based on the real spatio-temporal attention module according to claim 6, characterized in that: The information acquisition branch performs 1×1 convolution on the information acquisition feature map to expand the number of channels to H×W, obtaining the first global temporal attention matrix w 1 , then for signal i, the value at position i on the j-th channel feature map represents the influence degree of signal j on signal i Among them, Conv 1×1 represents a 1×1 convolution operation; represents an information acquisition feature map; Perform a reshape operation to adjust the dimensions of the first global temporal attention matrix, and use the Hadamard product to assign the weights of the interaction between signals to the information acquisition feature map to correct the information acquisition feature map, obtaining the output result of the information acquisition branch and; Among them, represents the matrix composed of all information acquisition feature maps.
8. The method for detecting the quality of the solder joint fusion core based on the real spatio-temporal attention module according to claim 6, characterized in that: The information distribution branch performs 1×1 convolution on the information distribution feature map to expand the number of channels to H×W, obtaining the second global temporal attention matrix w 2 , then for signal i, the value at position i on the j-th channel feature map represents the influence degree of signal j on signal i Among them, Conv 1×1 represents a 1×1 convolution operation; represents an information distribution feature map; Perform a reshape operation to adjust the size of the second global temporal attention matrix, and use the Hadamard product to assign the weights of the mutual influence between signals to the information allocation feature map to correct the information allocation feature map, obtaining the output result of the information allocation branch and; Among them, represents the matrix composed of all information distribution feature maps.
9. The method for detecting the quality of the solder joint fusion core based on the real spatio-temporal attention module according to claim 6, characterized in that: The output of the global time context attention unit is as follows: Among them, Z 2 represents the output of the global time context attention unit; F 1 represents the output of the information collection branch; F 2 represents the output of the information distribution branch; represents a concatenation operation along the channel dimension; The output of the time processing module is: Among them, Z represents the output of the time processing module; Z 1 represents the output of the short time series attention unit.