Welding defect detection method and system based on visual mixed attention mechanism

The context perception and feature utilization of welding defect detection model are enhanced through visual hybrid attention mechanism, solving the problems of feature loss and location information loss in traditional solutions, improving detection accuracy and generalization capabilities, and reducing labor costs.

CN120259280AActive Publication Date: 2025-07-04NINGDE SKEQI INTELLIGENT EQUIP CO LTD

Patent Information

Application Number
CN202510716942.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-04
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Existing computer vision and deep learning solutions have problems such as poor generalization, loss of feature, loss of position information transmission and low detection accuracy in welding defect detection, especially when detecting small defects and low resolution images.

Method used

The visual hybrid attention mechanism is adopted, combined with the multi-head self-attention and coordinate attention mechanism, and through feature extraction, optimization and fusion, the context perception ability and feature utilization range of the model are enhanced, and a hybrid attention feature pyramid network architecture (HAFPN) is constructed to solve the problems of feature loss and location information loss.

Benefits of technology

It improves the identification accuracy and generalization ability of welding defect detection, reduces the inconsistency of manual standards and labor costs, and achieves efficient defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259280A_ABST
    Figure CN120259280A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and discloses a welding defect detection method and system based on a visual mixed attention mechanism, and the method comprises the steps: firstly, capturing a welding spot image of an object to be detected through an industrial camera, and inputting the image at a fixed resolution; an input image is transmitted to a feature extraction layer for feature extraction, the input image is set as x, and feature optimization is performed on a feature map output by the feature extraction layer; and enabling the optimized feature map to enter an output layer, and carrying out target detection. According to the method, the context sensing capability of the model is enhanced, the feature utilization range is widened, the model has higher robustness, and different types of welding defects can be effectively overcome. With the help of an attention mechanism, the problems of feature loss and position information transmission loss of a traditional DNN scheme are solved, and the problem that the detection rate of unobvious defects is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and specifically to a welding defect detection method and system based on a visual hybrid attention mechanism. Background Art

[0002] Traditional manual welding defect detection is no longer suitable for modern high-efficiency intelligent factories due to low efficiency, inconsistent evaluation, high cost, lack of real-time data, etc., and is gradually being replaced by computer vision-based automated detection.

[0003] Detection solutions based on computer vision have become mainstream in industrial scenarios such as solder joint defect detection due to advantages such as real-time performance, continuity, and non-contact. However, existing computer vision and deep learning defect detection solutions have the following problems: First, some solutions rely on manually extracted rules and do not have good generalization for different industrial parts, requiring expert prior knowledge; Second, some solutions based on Deep Neural Networks (DNN) have good real-time performance and fast detection speed, but have poor detection effects for parts with small defect areas and low-resolution images; Third, some solutions based on models of YOLO (You Only Look Once) and multi-layer DNNs have improved accuracy, but have poor real-time performance and lose some features, which makes it difficult to cover high-precision scenarios.

[0004] In order to meet higher-efficiency production and manufacturing, a more intelligent new solution is needed to solve the problems of low precision, high false detection rate, and computational cost in solder joint defect detection in surface mount technology in industrial scenarios; Therefore, for the above problems, a welding defect detection method and system based on a visual hybrid attention mechanism are needed. Summary of the Invention

[0005] The purpose of the present invention is to provide a welding defect detection method and system based on a visual hybrid attention mechanism. The present invention enhances the model's ability to perceive context, improves the utilization range of features, makes the model more robust, and can effectively target different types of welding defects. With the help of the attention mechanism, the problems of feature loss and loss of position information transmission in traditional DNN solutions are solved, and the problem of low detection rate for unobvious defects is solved.

[0006] The present invention is implemented as follows: The present invention provides a welding defect detection method based on a visual hybrid attention mechanism, which is specifically executed according to the following steps: S1: For the object to be detected, the industrial camera captures the solder joint image and inputs it at a fixed resolution; the fixed resolution is 1280x1280 pixels or 1920x1280 pixels; S2: For the input image, it is passed to the feature extraction layer for feature extraction. Let the input image be , as shown in Equation (1); Equation (1); Among them, the size is H×W×C , where H is the height, W is the width, C is the number of channels; Specifically, it is executed according to the following steps: S2.1: First, the input image is subjected to feature extraction through a series of convolutions and normalizations to extract the primary features of the image, including edges and textures, and then the features are integrated through pyramid pooling; S2.2: The input image passes through the convolutional layer. The convolutional layer consists of multiple learnable filter convolution kernels. The filter slides on the image, calculates the dot product of the filter and the local area of the image, and outputs the feature map, as shown in Equation (2); Equation (2); Among them represents the input image, the original image or the output of the previous layer, with a size of H×W×C, k represents the convolution kernel matrix, obtained by model training, the model is the HAFPN feature pyramid extraction module, with a size of k×k×C, b represents the bias term, which is a constant value added to the result of the convolution operation, with a size of 1×1×m, and * is the convolution calculation; The model training is specifically executed according to the following steps: S2.2.1: First, perform positioning: align the upper left corner of the convolution kernel with the upper left corner of the input image; S2.2.2: Then calculate the dot product of the convolution kernel and the local area of the input image; S2.2.3: Then slide the convolution kernel one pixel to the right and repeat the dot product calculation until the end of the current row is reached; S2.2.4: Move the convolution kernel to the start position of the next row and repeat steps S2.2.2 and S2.2.3 until the entire input image is covered; S2.2.5: Obtain the result of the last dot product and finally output a new feature map through the convolution operation.

[0007] S2.3: Perform normalization processing. A batch normalization layer follows the convolutional layer to stabilize the training and accelerate convergence, as shown in Equation (3); Equation (3); where x’ represents the output of the current layer, represents the mean of the output of the current layer, σ represents the variance of the output of the current layer, , preventing division by zero; S2.4: The normalized feature map passes through the SiLu activation function to introduce non-linearity, as shown in Equation (4); Equation (4); where, represents the input image; S2.5: Use the C3 convolutional layer with the CSP structure to further extract features, as shown in Equation (5); Equation (5); where, represents the input image, and CSP(x) is the feature map after being processed by the Cross Stage Partial structure; S2.6: Perform pyramid pooling to integrate feature information at different scales, reduce the size of the feature map, and reduce the subsequent computational amount, as shown in Equation (6); Equation (6); where, represents the input image, and pool(x, k) represents using pooling kernels of different sizes k to perform a pooling operation. Continue to perform the pooling operation by sliding a fixed-size window over the input image and selecting the maximum value within each window as the output.

[0008] S3: Optimize the feature map output by the feature extraction layer; specifically, execute according to the following steps: S3.1: Process the feature map through depthwise separable convolution DWConv and layer normalization LN, and then further optimize the feature map through the enhanced multi-head self-attention EMSA and coordinate attention CA mechanisms; Through multi-head self-attention, through the multi-scale and multi-head self-attention mechanisms, capture global context information and enhance the expressive ability of features, as shown in Equation (7); Equation (7); where, X input represents the input feature, X output represents the output feature, X m and X nRepresents the intermediate feature, Q, K, and V represent the query matrix, key matrix, and value matrix respectively. Linear is a linear transformation operation, SiLU is the activation function of Linear, FC represents the fully connected layer, and d is a scalar factor.

[0009] S3.2: Perform feature fusion. The feature map after being processed by the hybrid attention mechanism is fused through a multi-layer perceptron layer MLP to obtain an optimized feature map.

[0010] The feature map after being processed by the hybrid attention mechanism is fused through a multi-layer perceptron layer MLP to obtain an optimized feature map; First, the input features achieve parameter sharing through a depthwise separable convolutional residual block and enhance the learning ability for local features, then are processed by a normalization layer. Its output is processed by a multi-head attention mechanism and a coordinate attention mechanism respectively, normalized by LN, and finally obtained through a multi-layer perceptron layer.

[0011] S4: Import the optimized feature map into the output layer to output the result, as shown in Equation (8); Equation (8); Among them, X input represents the input feature, X output represents the output feature, X m , X n and X q are intermediate features, DWconv represents depthwise separable convolution, LN represents normalization, CA represents the coordinate attention mechanism, EMSA represents the enhanced multi-head self-attention mechanism, and MLP represents the multi-layer perceptron.

[0012] For object detection, first, after feature optimization, a feature map F : H x W x D is obtained, where H and W are the height and width of the feature map, and C is the number of channels. Based on this, the defective bounding box can be obtained; Among them: Center point x: ; Center point y: ; Width w: ; Height h: ; Among them, σ is the sigmoid function, W and H are the width and height of the feature map respectively, e x and e yis the coordinate offset, which is the offset of the predicted feature map coordinates compared to the original input size, P w and P h are the predefined width and height respectively, which can be defined as the mean value of the defect size in the scenario described in the present invention. and are the offsets used to convert the predicted value to the actual size; Then calculate the defect probability; for the output feature map of the fully connected layer defined as f, the defect probability can be obtained through the sofxtmax layer, as shown in Equation (9); Equation (9); Among them, K is the total number of defect categories, which is 2 when only detecting the existence, P is the probability distribution of each category, and the category with the largest corresponding value is taken as the result.

[0013] Furthermore, the present invention provides a welding defect detection system based on a visual hybrid attention mechanism, including a data input module for receiving the solder joint image captured by an industrial camera; a feature extraction module for extracting features from the input image through convolution and normalization; a feature optimization module for importing the feature map output by the feature extraction module into the feature optimization module: an output module for importing the optimized feature map into the output module for target detection and outputting the defect probability result.

[0014] Furthermore, the present invention provides a computer-storable medium, which includes an embedded processing system and a stored program. When the program runs under the control of the embedded system, it controls the traffic perturbation generation method based on packet feature inversion to execute a welding defect detection method according to any one of the above.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. By designing a multi-head self-attention module and a position attention module, the scope of feature utilization of the network is enhanced, and the calculation speed of the model is accelerated. Based on the multi-head self-attention module and the position attention module, a hybrid attention mechanism (HAM) module is designed to improve the local feature learning ability. On this basis, a hybrid attention feature pyramid network architecture (HAFPN) module is constructed to enhance the ability of FPN to perceive context information and solve the problem of accuracy decline caused by the loss of position information. Finally, based on the HAFPN module, a feature detection system based on the visual hybrid attention mechanism is implemented, which improves the recognition accuracy of welding defect detection in the industrial field while ensuring the detection speed. In addition, the solution based on the visual hybrid attention mechanism of this system can also reduce the problems of inconsistent manual standards and high labor costs.

[0016] 2. The proposed solution of the present invention is based on the hybrid attention mechanism, which enhances the model's ability to perceive context, improves the scope of feature utilization, makes the model more robust, and can effectively target different types of welding defects. With the help of the attention mechanism, the problems of feature loss and loss of position information transmission in the traditional DNN solution are solved, and the problem of low detection rate of unobvious defects is solved.

[0017] 3. By combining the self-attention mechanism with the coordinate attention mechanism, a hybrid attention network is designed to solve the problem of feature loss encountered by traditional deep neural networks in the field of industrial defect detection. Applying the hybrid attention mechanism to the YOLO detection model solves the problem of low detection accuracy of subtle defects and improves the generalization ability of the model.

[0018] 4. The hybrid attention mechanism designed in the present invention enhances the network's ability to perceive long-distance position information and learn local features, and improves the recognition generalization and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0020] Figure 1 is the method flow chart of the present invention; Figure 2 is the system structure diagram of the present invention; Figure 3 is the system structure flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but is merely for the selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] Please refer to Figures 1 - 3 , the present invention provides a welding defect detection method based on a visual hybrid attention mechanism, which is specifically executed according to the following steps: S1: For the object to be detected, a solder joint image is captured by an industrial camera and input at a fixed resolution; the fixed resolution is 1280x1280 pixels or 1920x1280 pixels; S2: For the input image, it is passed to the feature extraction layer for feature extraction. Let the input image be , as shown in Equation (1); Equation (1); where the size is H×W×C , where H is the height, W is the width, C is the number of channels; Specifically, it is executed according to the following steps: S2.1: The input image is first subjected to feature extraction through a series of convolutions and normalizations to extract the primary features of the image, including edges and textures, and then the features are integrated through pyramid pooling; S2.2: The input image passes through a convolutional layer. The convolutional layer consists of multiple learnable filter convolution kernels. The filter slides on the image, calculates the dot product of the filter and the local area of the image, and outputs a feature map, as shown in Equation (2); Equation (2); where represents the input image matrix, the original image or the output of the previous layer, with a size of H×W×C, k represents the convolution kernel matrix, obtained by model training, the model is the HAFPN feature pyramid extraction module, with a size of k×k×C, b represents the bias term, a constant value added to the result of the convolution operation, with a size of 1×1×m, and * represents the convolution calculation; The model training is specifically executed according to the following steps: S2.2.1: First, perform positioning: align the upper left corner of the convolutional kernel with the upper left corner of the input image; S2.2.2: Then, calculate the dot product of the convolutional kernel and the local region of the input image; S2.2.3: Then, slide the convolutional kernel one pixel to the right and repeat the dot product calculation until reaching the end of the current row; S2.2.4: Move the convolutional kernel to the start position of the next row and repeat steps S2.2.2 and S2.2.3 until covering the entire input image; S2.2.5: Obtain the result of the last dot product and finally output a new feature map through the convolution operation.

[0023] S2.3: Perform normalization processing. A batch normalization layer follows the convolutional layer, which is used to stabilize training and accelerate convergence, as shown in Equation (3); Equation (3); Where x’ represents the output of the current layer, represents the mean of the output of the current layer, σ represents the variance of the output of the current layer, , to prevent division by zero; S2.4: The normalized feature map passes through the SiLu activation function to introduce non-linearity, as shown in Equation (4); Equation (4); Where, represents the input image; S2.5: Use the C3 convolutional layer with the CSP structure to further extract features, as shown in Equation (5); Equation (5); Where, represents the input image, and CSP(x) is the feature map processed by the Cross Stage Partial structure; S2.6: Perform pyramid pooling to integrate feature information of different scales, reduce the size of the feature map, and reduce the subsequent calculation amount, as shown in Equation (6); Equation (6); Where, represents the input image, and pool(x, k) represents using pooling kernels of different sizes k to perform a pooling operation. Continue to perform the pooling operation on the input image Slide a window of a fixed size upwards and select the maximum value within each window as the output.

[0024] S3: Optimize the feature map output by the feature extraction layer; specifically, execute according to the following steps: S3.1: Process the feature map through depthwise separable convolution DWConv and layer normalization LN, and then further optimize the feature map through enhanced multi-head self-attention EMSA and coordinate attention CA mechanisms; Through multi-head self-attention, through multi-scale and multi-head self-attention mechanisms, capture global context information and enhance the expressive ability of features, as shown in Equation (7); Equation (7); Among them, X input represents the input feature, X output represents the output feature, X m and X n represent intermediate features. Q, K, and V represent the query matrix, key matrix, and value matrix respectively. Linear is a linear transformation operation, SiLU is the activation function of Linaer, FC represents a fully connected layer, and d is a scalar factor.

[0025] First, linearly transform the original features through the Q, K, and V components of the fully connected layer behavior, multiply the Q and K matrices and then perform a series of non-linear transformations, input into the Silu activation function after passing through the fully connected layer, and then use Tanh for processing after passing through another fully connected layer. The output result is the matrix after multiplying with the V component matrix of the linear transformation. Finally, use the fully connected layer to fuse with the original input features to obtain the final output result; Compared with the original MSA, EMSA has more non-linear transformations, which can make the context perception ability stronger, expand the utilization range of features by the model network, and make the network expression ability stronger; And through coordinate attention, enhance the model's perception ability of position information; first, perform pooling on the height and width of the image to obtain feature maps with dimensions of and 1×W×C. Subsequently, connect the feature maps and reduce the dimension through a shared convolution to obtain a feature map with dimensions of . Then, perform non-linear transformation to enhance its expressive ability. Subsequently, use convolution to restore the original dimension. Finally, use HardSigmoid. Compared with SigMoid, HardSigmoid does not require exponentiation, so its calculation speed is faster.

[0026] S3.2: Perform feature fusion. The feature map processed by the hybrid attention mechanism is fused through a multi-layer perceptron layer MLP to obtain an optimized feature map.

[0027] The feature map processed by the hybrid attention mechanism is fused through a multi-layer perceptron layer MLP to obtain an optimized feature map. First, the input features achieve parameter sharing through a depthwise separable convolutional residual block and enhance the learning ability for local features. Then, it is processed by a normalization layer. After the outputs are processed by a multi-head attention mechanism and a coordinate attention mechanism respectively, they are normalized by LN. Finally, the output result is obtained through a multi-layer perceptron layer, as shown in Equation (8); Equation (8); Among them, X input represents the input features, X output represents the output features, X m , X n and X q are intermediate features. DWconv represents depthwise separable convolution, LN represents normalization, CA represents coordinate attention mechanism, EMSA represents enhanced multi-head self-attention mechanism, and MLP represents multi-layer perceptron.

[0028] S4: Import the optimized feature map into the output layer for object detection. First, after feature optimization, a feature map F : H x W x D is obtained, where H and W are the height and width of the feature map, and C is the number of channels. Based on this, the defective bounding box can be obtained. Among them: Center point x: ; Center point y: ; Width w: ; Height h: ; Among them, σ is the sigmoid function, W and H are the width and height of the feature map respectively, e x and e y are the coordinate offsets, predicting the offset of the feature map coordinates compared to the original input size, p w and p h are the predefined width and height respectively, which can be defined as the mean of the defect size in the scenario described in the present invention, and is the offset used to convert the predicted value into the actual size; Then calculate the defect probability; for the image output by the fully connected layer Defined as f, the defect probability can be obtained through the sofxtmax layer, as shown in Equation (9); Equation (9); Where K is the total number of defect categories, which is 2 when only detecting the existence, P is the probability distribution of each category, and the category with the largest corresponding value is taken as the result.

[0029] In this embodiment, the present invention provides a welding defect detection system based on a visual hybrid attention mechanism, including a data input module for receiving the solder joint image captured by an industrial camera; A feature extraction module for extracting features from the input image through convolution and normalization; A feature optimization module for importing the feature map output by the feature extraction module into the feature optimization module: An output module for importing the optimized feature map into the output module for target detection and outputting the defect probability result.

[0030] In the second embodiment provided in this embodiment, the present invention provides a computer-storable medium, and the computer-readable storage medium includes an embedded processing system and a stored program, which controls the execution of the traffic disturbance generation method based on packet feature inversion to perform any one of the above-mentioned welding defect detection methods based on a visual hybrid attention mechanism when the program runs under the control of the embedded system.

[0031] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, there are various changes and modifications to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A welding defect detection method based on a visual hybrid attention mechanism, characterized in that: The specific implementation steps are as follows: S1: For the object to be detected, a soldering point image is captured by an industrial camera and input at a fixed resolution; S2: The input image is passed to the feature extraction layer for feature extraction. Let the input image be , as shown in Equation (1); Formula (1); Among them, the size is H×W×C , where H is the height, W is the width, C is the number of channels; S3: Feature optimization is performed on the feature map output by the feature extraction layer; S4: The optimized feature map is imported into the output layer for object detection.

2. The welding defect detection method based on the visual hybrid attention mechanism according to claim 1, wherein: In step S2, the specific implementation steps are as follows: S2.1: The input image is first subjected to feature extraction through a series of convolutions and normalizations to extract primary features of the image, including edges and textures, and then the features are integrated through pyramid pooling; S2.2: The input image passes through a convolutional layer, which consists of multiple learnable filter convolution kernels. The filter slides on the image, calculates the dot product between the filter and the local area of the image, and outputs a feature map, as shown in Equation (2); Formula (2); Among them represents the input image, the original image or the output of the previous layer, with a size of H×W×C. k represents the convolutional kernel matrix, which is obtained through model training. The model is the HAFPN feature pyramid extraction module, with a size of k×k×C. b represents the bias term, which is a constant value added to the result of the convolutional operation, with a size of 1×1×m, and * represents the convolutional calculation; S2.3: Normalization processing is performed. A batch normalization layer follows the convolutional layer to stabilize training and accelerate convergence, as shown in Equation (3); Formula (3); Among them x’ represents the output of the current layer, represents the mean of the output of the current layer, σ represents the variance of the output of the current layer, , to prevent division by zero; S2.4: The normalized feature map passes through the SiLu activation function to introduce non-linearity, as shown in Equation (4); Formula (4); Among them, represents the input image; S2.5: The C3 convolutional layer with a CSP structure is used to extract features, as shown in Equation (5); Formula (5); Among them, represents the input image, and CSP(x) is the feature map after being processed by the Cross Stage Partial structure; S2.6: Pyramid pooling is performed to integrate feature information at different scales, reduce the size of the feature map, and reduce the subsequent computational amount, as shown in Equation (6); Formula (6); Among them, represents the input image, and pool(x, k) represents using pooling kernels of different sizes k to perform a pooling operation.

3. A welding defect detection method based on a visual hybrid attention mechanism according to claim 1, characterized in that: In step S2.2, the model training is specifically implemented as follows: S2.2.1: First, perform positioning: align the upper left corner of the convolution kernel with the upper left corner of the input image; S2.2.2: Then calculate the dot product between the convolution kernel and the local area of the input image; S2.2.3: Then slide the convolution kernel one pixel to the right and repeat the dot product calculation until the end of the current row is reached; S2.2.4: Move the convolution kernel to the start position of the next row and repeat steps S2.2.2 and S2.2.3 until the entire input image is covered; S2.2.5: Obtain the result of the last dot product and finally output a new feature map through convolution operation.

4. The welding defect detection method based on the visual hybrid attention mechanism according to claim 2, characterized in that: In step S2.6, continue with the pooling operation, sliding a window of a fixed size over the input image and selecting the maximum value within each window as the output.

5. A welding defect detection method based on a visual hybrid attention mechanism according to claim 1, characterized in that: In step S3, the specific implementation steps are as follows: S3.1: Process the feature map through depthwise separable convolution DWConv and layer normalization LN, and then optimize the feature map through the enhanced multi-head self-attention EMSA and coordinate attention CA mechanisms; S3.2: Perform feature fusion. The feature map processed by the hybrid attention mechanism passes through a multi-layer perceptron layer MLP for fusion to obtain the optimized feature map.

6. A welding defect detection method based on a visual hybrid attention mechanism according to claim 5, characterized in that: In step S3.1, through multi-head self-attention, through multi-scale and multi-head self-attention mechanisms, global context information is captured to enhance the expression ability of features, as shown in Equation (7); Formula (7); Among them, X input represents the input feature, X output represents the output feature, X m and X n represent intermediate features, Q, K, and V represent the query matrix, key matrix, and value matrix respectively, Linear is a linear transformation operation, SiLU is the activation function of Linaer, FC represents a fully connected layer, and d is a scalar factor.

7. A welding defect detection method based on a visual hybrid attention mechanism according to claim 5, characterized in that: In step S3.2, the feature map processed by the hybrid attention mechanism is fused through a multi-layer perceptron layer MLP to obtain an optimized feature map; first, the input features achieve parameter sharing through a depthwise separable convolutional residual block and enhance the learning ability for local features, then are processed using a normalization layer, and their outputs are respectively processed through a multi-head attention mechanism and a coordinate attention mechanism, normalized using LN, and finally the output result is obtained through a multi-layer perceptron layer, as shown in Equation (8); Formula (8); Among them, X input represents the input feature, X output represents the output feature, X m , X n and X q are intermediate features, DWconv represents depthwise separable convolution, LN represents normalization, CA represents coordinate attention mechanism, EMSA represents enhanced multi-head self-attention mechanism, and MLP represents multi-layer perceptron.

8. The welding defect detection method based on the visual hybrid attention mechanism according to claim 1, characterized in that: In step S4, after feature optimization, a feature map is obtained first F : H x W x D, where H and W are the height and width of the feature map, and C is the number of channels. Based on this, the defective bounding box can be obtained; where: Center point x: ; Center point y: ; Width w: ; Height h: ; Among them, σ is the sigmoid function, W and H are the width and height of the feature map, respectively, e x and e y are the coordinate offsets, representing the offsets of the predicted feature map coordinates compared to the original input size, P w and P h are the predefined width and height, respectively, and are the offsets used to convert the predicted values to the actual size; Then calculate the defect probability; for the image output by the fully connected layer is defined as , and the defect probability can be obtained through the softmax layer, as shown in Equation (9); Formula (9); Among them, is the total number of defect categories, which is 2 when only detecting the existence, is the probability distribution of each category, and the category with the largest corresponding value is taken as the result.

9. A welding defect detection system based on a visual hybrid attention mechanism, characterized in that: It includes a data input module for receiving the solder joint image captured by an industrial camera; A feature extraction module for extracting features from the input image through convolution and normalization; A feature optimization module for importing the feature map output by the feature extraction module into the feature optimization module: An output module for importing the optimized feature map into the output module for target detection and outputting the defect probability result.

Citation Information

Patent Citations

  • Semantic segmentation method based on feature pyramid attention and mixed attention cascading

    CN112651973A

  • Wafer surface defect mode detection method based on deep attention network

    CN113362320A

  • Welding defect detection method and device based on attention fusion

    CN116843657A

  • Intelligent machine vision detection method and system based on image processing and storage medium

    CN119205719A

  • Chip defect visual inspection method

    CN119515875A

Cited By

  • Welding quality detection method and device based on machine vision

    CN120908185A