A method for identifying defects of automotive gears based on machine vision
By integrating self-attention, spatial attention and channel attention mechanisms in vehicle gear defect recognition, the problem of low recognition accuracy in the prior art is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202410689881.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-05-30
AI Technical Summary
The prior art is difficult to effectively extract subtle features of automotive gear defects, resulting in low recognition accuracy and susceptible to environmental factors.
The multi-attention mechanism fusion method is adopted, combining the self-attention mechanism, spatial attention mechanism and channel attention mechanism, and through convolution and element-by-element addition, the fusion effect of feature map weights is improved, thereby improving the recognition accuracy of automobile gear defects.
It significantly improves the accuracy and robustness of vehicle gear defect identification, and enhances quality control and safe production levels.
Smart Images

Figure CN118587487B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a method for identifying defects of automotive gears based on machine vision. Background Art
[0002] With the rapid development of industrial automation, machine vision technology has been increasingly widely applied in the field of quality inspection. Especially in the automotive manufacturing industry, gears, as key transmission components, their quality is directly related to the performance and safety of automobiles. Therefore, accurate identification of automotive gear defects is particularly important. Traditional methods for gear defect identification mostly rely on manual visual inspection, which is not only inefficient but also easily affected by subjective factors, resulting in missed or misjudged inspections. Therefore, using machine vision technology to achieve automated and high-precision identification of automotive gear defects has become a research hotspot.
[0003] In the field of machine vision, deep learning models such as convolutional neural networks (CNNs) have been proven to have excellent performance in tasks such as image classification and object detection. However, for complex gear defect identification tasks, a simple CNN model often has difficulty fully extracting the subtle features of defects. This is mainly because there are a wide variety of gear defect types, and the defect sizes and shapes are different. At the same time, it may also be affected by environmental factors such as lighting and noise. Therefore, how to improve the model's ability to extract defect features has become the key to improving the accuracy of gear defect identification.
[0004] In recent years, the attention mechanism has received extensive attention in the field of deep learning. By simulating the selective attention mechanism of the human visual system, it enables the model to automatically focus on the key regions in the image, thereby improving the pertinence and efficiency of feature extraction. Multi-attention fusion is to combine different types of attention mechanisms to make full use of their respective advantages and further improve the performance of the model. For example, the spatial attention mechanism can focus on information at different positions in the image, while the channel attention mechanism can emphasize the dependence between different channels. By fusing these two attention mechanisms, it is possible to better capture the local and global features of gear defects, thereby improving the accuracy of identification.
[0005] In summary, applying the combination of machine vision technology and multi-attention mechanism to automotive gear defect identification can not only improve the degree of automation and efficiency of identification, but also effectively improve the accuracy and robustness of identification, which is of great significance for improving the quality control and safety production level of the automotive manufacturing industry. Summary of the Invention
[0006] The present invention provides a method for identifying defects in automotive gears based on machine vision, aiming to integrate multiple attention mechanisms to improve the recognition effect of machine vision in the scenario of automotive gear defect recognition.
[0007] Based on the machine vision recognition, the present invention aims to integrate multiple attention mechanisms and provides a method for identifying defects in automotive gears based on machine vision, including the following steps:
[0008] S1. Construct a machine vision detection module. Input the feature map into Branch 1, Branch 2, Branch 3, and Branch 4. In Branch 1, perform self-attention mechanism operation on the input feature map to obtain a self-attention feature map. In Branch 2, perform convolution and spatial attention mechanism operations on the input feature map in sequence to obtain a spatial attention feature map. In Branch 3, perform convolution and channel attention mechanism operations on the input feature map in sequence to obtain a channel attention feature map. Add the self-attention feature map, spatial attention feature map, and channel attention feature map element by element to obtain an intermediate feature map. Perform convolution operation on the intermediate feature map to obtain a convolutional intermediate feature map. Add the convolutional intermediate feature map and the input feature map of Branch 4 element by element to obtain an output feature map. Input the output feature map into the detection head to obtain the detection result.
[0009] S2. Prepare the automotive gear defect dataset. The recognized categories are divided into two types: defective and non-defective. For the defective category, collect defective automotive gears. Place the defective side of the gear horizontally and rotate it by a certain angle in sequence. Take pictures of the defective automotive gears from directly above multiple times, and label the entire automotive gear in these pictures as defective. Collect normal automotive gears. Place one side of the gear horizontally and rotate it by a certain angle in sequence. Take pictures of the defective automotive gears from directly above multiple times, and label the entire automotive gear in these pictures as non-defective. After obtaining the automotive gear defect dataset, perform Mixup data augmentation on the pictures in the dataset.
[0010] S3. Use the machine vision algorithm to train the automotive gear defect dataset. The machine vision algorithm includes an input end, a backbone network, a machine vision detection module, and an output end. The input end inputs pictures and annotation files. Pass the pictures through the backbone network, which includes convolutional layers, residual blocks, batch normalization layers, activation function layers, pooling layers, and linear layers. After the pictures pass through the backbone network, output the corresponding feature maps of the pictures. Input the feature maps into the machine vision detection module to obtain the detection result. After training is completed, obtain the detection model.
[0011] S4. Use the detection model to detect the defects in automotive gears. The input of the detection model is a single picture or a frame extracted from a camera video. After the image is input into the detection model, it passes through the backbone network and the machine vision detection module in sequence, and then is input into the detection head to obtain the detection target position and classification. The classification includes defective or non-defective. If the classification is defective, it is determined that the gears of the current vehicle have defects.
[0012] Preferably, in step S1, given an input feature map , where H and W represent the width and height respectively, the input image enters branches 1, 2, 3, and 4 of the module. Branch 1 performs operations through the self-attention mechanism to obtain a self-attention feature map , branch 2 performs operations through convolution and the spatial attention mechanism to obtain a spatial attention feature map , branch 3 performs operations through convolution and the channel attention mechanism to obtain a channel attention feature map , the self-attention feature map, the spatial attention feature map, and the channel attention feature map are added element-wise to obtain an intermediate feature map , the intermediate feature map is subjected to convolution operations to obtain a convolutional intermediate feature map , the convolutional intermediate feature map and the input feature map of branch 4 are added element-wise to obtain an output feature map .
[0013] Preferably, in step S2, the side with gear defects is placed horizontally and rotated by a certain angle in sequence. The angle is set to 30 degrees. In this way, multiple pictures of gear defects in different positions can be obtained. For Mixup data augmentation, new sample data is obtained through linear interpolation, including flipping, rotation, scaling, and shifting. While generating new pictures, new annotation data is generated, and the new samples and new annotation data are integrated into the automotive gear defect dataset.
[0014] Preferably, in step S3, the convolutional layer is used to extract image features, the residual block is used to construct a deep network, the batch normalization layer is used to accelerate network training and improve performance, the activation function layer uses functions such as ReLU to increase the non-linear expression ability of the network, the pooling layer is used to reduce the dimension of the feature map and reduce the amount of calculation, and the linear layer is used to map the features extracted previously to the label space of the samples. At the same time, in the backbone network, the convolutional layer is part of the residual block, and the output of the convolutional layer will be added to the input to form a residual connection. The calculation method of the residual block is , where y is the output of the residual block and x is the input of the residual block is part of the residual block, including multiple convolutional layers, batch normalization layers, and activation function layers. The residual connection is formed by directly adding the input x to , enabling the algorithm to learn the residual mapping between the input and the output.
[0015] Compared with the prior art, the present invention has the following technical effects:
[0016] The technical solution provided by the present invention proposes a machine vision detection module that integrates self-attention mechanism, spatial attention mechanism, and channel attention mechanism. By performing operations such as convolution and element-wise addition on these attention mechanisms, the weights of the feature maps are appropriately fused, thereby improving the recognition effect of machine vision in the scenario of automotive gear defect recognition. Description of the Drawings
[0017] Figure 1 It is a flowchart of automotive gear defect recognition provided by the present invention.
[0018] Figure 2 It is a structural diagram of the machine vision detection module provided by the present invention. Detailed Embodiments
[0019] The present invention aims to propose a method for automotive gear defect recognition based on machine vision, and proposes a machine vision detection module that integrates self-attention mechanism, spatial attention mechanism, and channel attention mechanism. By performing operations such as convolution and element-wise addition on these attention mechanisms, the weights of the feature maps are appropriately fused, thereby improving the recognition effect of machine vision in the scenario of automotive gear defect recognition.
[0020] Please refer to Figure 1 As shown, a method for automotive gear defect recognition based on machine vision in an embodiment of the present application:
[0021] S1. Construct a machine vision detection module. As Figure 2 shown, input the feature map into Branch 1, Branch 2, Branch 3, and Branch 4. In Branch 1, perform self-attention mechanism operation on the input feature map to obtain a self-attention feature map. In Branch 2, perform convolution and spatial attention mechanism operations on the input feature map in sequence to obtain a spatial attention feature map. In Branch 3, perform convolution and channel attention mechanism operations on the input feature map in sequence to obtain a channel attention feature map. Add the self-attention feature map, spatial attention feature map, and channel attention feature map element-wise to obtain an intermediate feature map. Perform convolution operation on the intermediate feature map to obtain a convolution intermediate feature map. Add the convolution intermediate feature map and the input feature map of Branch 4 element-wise to obtain an output feature map. Input the output feature map into the detection head to obtain the detection result;
[0022] S2. Prepare an automotive gear defect dataset, and classify the recognized types into two categories: defective and non-defective. For the defective category, collect defective automotive gears, place the defective side of the gear horizontally and rotate it by a certain angle in sequence, take pictures of the defective automotive gears from directly above multiple times, and label the entire automotive gears in these pictures as defective. Collect normal automotive gears, place one side of the gear horizontally and rotate it by a certain angle in sequence, take pictures of the defective automotive gears from directly above multiple times, and label the entire automotive gears in these pictures as non-defective. After obtaining the automotive gear defect dataset, perform Mixup data augmentation on the pictures in the dataset;
[0023] S3. Use a machine vision algorithm to train the automotive gear defect dataset. The machine vision algorithm includes an input end, a backbone network, a machine vision detection module, and an output end. The input end inputs pictures and annotation files, and passes the pictures through the backbone network. The backbone network includes a convolutional layer, a residual block, a batch normalization layer, an activation function layer, a pooling layer, and a linear layer. After the pictures pass through the backbone network, the corresponding feature maps of the pictures are output. The feature maps are input into the machine vision detection module to obtain detection results. After training is completed, a detection model is obtained;
[0024] S4. Use the detection model to detect automotive gear defects. The input of the detection model is a single picture or a frame extracted from a camera video. After the image is input into the detection model, it passes through the backbone network and the machine vision detection module in sequence, and is input into the detection head to obtain the detection target position and classification. The classification includes defective or non-defective. If the classification is defective, it is determined that the gears of the current vehicle are defective.
[0025] Further, in step S1, given an input feature map , where H and W represent width and height respectively, input the input image into branches 1, 2, 3, and 4 of this module. Branch 1 performs operations through the self-attention mechanism to obtain a self-attention feature map , branch 2 performs operations through convolution and the spatial attention mechanism to obtain a spatial attention feature map , branch 3 performs operations through convolution and the channel attention mechanism to obtain a channel attention feature map , add the self-attention feature map, the spatial attention feature map, and the channel attention feature map element-wise to obtain an intermediate feature map , perform convolution operations on the intermediate feature map to obtain a convolutional intermediate feature map , add the convolutional intermediate feature map and the input feature map of branch 4 element-wise to obtain an output feature map .
[0026] Further, in S2, place the side with gear defects horizontally and rotate it by a certain angle in sequence. The angle is set to 30 degrees. In this way, multiple pictures of gear defects in different positions can be obtained. For Mixup data augmentation, new sample data is obtained through linear interpolation, including flipping, rotation, scaling, and shifting. While generating new pictures, new annotation data is generated, and the new samples and new annotation data are integrated into the automotive gear defect dataset.
[0027] Further, in S3, the convolutional layer is used to extract image features, the residual block is used to construct a deep network, the batch normalization layer is used to accelerate network training and improve performance, the activation function layer uses functions such as ReLU to increase the non-linear expression ability of the network, the pooling layer is used to reduce the dimension of the feature map and reduce the amount of calculation, and the linear layer is used to map the features extracted previously to the label space of the samples. At the same time, in the backbone network, the convolutional layer is part of the residual block, and the output of the convolutional layer will be added to the input to form a residual connection. The calculation method of the residual block is , where y is the output of the residual block and x is the input of the residual block. is part of the residual block, including multiple convolutional layers, batch normalization layers, and activation function layers. The residual connection is achieved by directly adding the input x to , enabling the algorithm to learn the residual mapping between the input and the output.
[0028] Further, in S3, the Retinanet network is used as the backbone network, and the Retinanet detection head is used as the detection head. For the detection result, if the detection head does not output a defect recognition box, it is regarded as having no target. If a defect recognition box is output, it is judged whether there is a defect or no defect according to the classification of the recognition box.
[0029] The above is only the preferred implementation mode of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the creative concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A method for identifying automobile gear defects based on machine vision, characterized in that: The following steps are involved: S1. Build a machine vision detection module, input feature maps to branches 1, 2, 3, and 4, perform self-attention operation on the input feature map in branch 1 to obtain a self-attention feature map, perform convolution and spatial attention operation on the input feature map in branch 2 to obtain a spatial attention feature map, perform convolution and channel attention operation on the input feature map in branch 3 to obtain a channel attention feature map, add the self-attention feature map, the spatial attention feature map, and the channel attention feature map element by element to obtain an intermediate feature map, perform convolution operation on the intermediate feature map to obtain a convolution intermediate feature map, add the convolution intermediate feature map and the input feature map of branch 4 element by element to obtain an output feature map, input the output feature map to the detection head, and obtain the detection result; S2. Prepare a dataset of automobile gear defects, and classify the identification types into defective and non-defective types. For defective types, collect defective automobile gears, place the defective side of the gear horizontally and rotate it by a certain angle in sequence, take pictures of automobile gear defects from above multiple times, and mark the automobile gears in these pictures as defective as a whole. Collect normal automobile gears, place one side of the gear horizontally and rotate it by a certain angle in sequence, take pictures of automobile gear defects from above multiple times, and mark the automobile gears in these pictures as non-defective as a whole. After obtaining the automobile gear defect dataset, perform Mixup data enhancement on the pictures in the dataset. S3. Use machine vision algorithm to train the automotive gear defect data set. The machine vision algorithm includes input end, backbone network, machine vision detection module and output end. The input end inputs the picture and annotation file, and passes the picture through the backbone network. The backbone network includes convolution layer, residual block, batch normalization layer, activation function layer, pooling layer and linear layer. After the picture passes through the backbone network, the feature map corresponding to the picture is output. The feature map is input into the machine vision detection module to obtain the detection result. After the training is completed, the detection model is obtained; S4. Use the detection model to detect automobile gear defects. The input of the detection model is a single picture or a frame of camera video. After the image is input into the detection model, it passes through the backbone network and the machine vision detection module in sequence and is input into the detection head to obtain the detection target position and classification. The classification includes defective or non-defective. If it is classified as defective, it is determined that the current vehicle gear has defects.
2. The method for identifying automobile gear defects based on machine vision according to claim 1, characterized in that: In step S1, given the input feature map , H and W represent height and width respectively, input feature map to branches 1, 2, 3 and 4 in this module, branch 1 is operated by self-attention mechanism to obtain self-attention feature map , branch 2 obtains the spatial attention feature map through convolution and spatial attention mechanism operation , branch 3 obtains the channel attention feature map through convolution and channel attention mechanism operation , add the self-attention feature map, spatial attention feature map and channel attention feature map element by element to get the intermediate feature map , perform convolution operation on the intermediate feature map to obtain the convolution intermediate feature map , add the convolution intermediate feature map and the branch 4 input feature map element by element to get the output feature map .
3. The method for identifying automobile gear defects based on machine vision according to claim 1, characterized in that: In step S2, the defective side of the gear is placed horizontally and rotated at a certain angle in turn. The angle is set to 30 degrees. In this way, multiple pictures of gear defects in different positions can be obtained. For Mixup data enhancement, new sample data is obtained through linear interpolation, including flipping, rotation, scaling and shifting. New annotation data is generated while generating new pictures, and the new samples and new annotation data are integrated into the automotive gear defect dataset.
4. The method for identifying automobile gear defects based on machine vision according to claim 1, characterized in that: In step S3, the convolution layer is used to extract image features, the residual block is used to build a deep network, the batch normalization layer is used to accelerate network training and improve performance, the activation function layer uses the ReLU function to increase the nonlinear expression ability of the network, the pooling layer is used to reduce the dimension of the feature map and reduce the amount of calculation, and the linear layer is used to map the previously extracted features to the sample label space. At the same time, the convolution layer in the backbone network is part of the residual block. The output of the convolution layer will be added to the input to form a residual connection. The calculation method of the residual block is: , where y is the output of the residual block and x is the input of the residual block. It is part of the residual block, which contains multiple convolutional layers, batch normalization layers, and activation function layers. The residual connection is connected by adding the input x directly to , allowing the algorithm to learn the residual mapping between input and output.
Citation Information
Patent Citations
Workpiece defect detection method and device fusing multi-attention mechanism
CN113822885A
Light-weight gear surface defect detection method based on MSTA-YOLOv5
CN115953386A