Wheel part detection method based on multi-scale dynamic fusion

By introducing the MSA-T module, DWFA module and DAD-Head module in wheel component detection, the problem of difficulty in capturing changes in multi-dimensional components and long-distance dependencies in the context of complexity in the prior art is solved, and higher detection robustness and accuracy are achieved.

CN120071000AActive Publication Date: 2025-05-30ZHILIAN INFORMATION TECH CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510180882.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-30
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the size changes of multi-size parts and the long-distance dependence in complex backgrounds in wheel parts detection, resulting in mis-checking and missed inspection problems, and it is difficult to balance small-target detection and large-target semantic understanding.

Method used

The MSA-T module, DWFA module and DAD-Head module are proposed. The MSA-T module extracts local features through multi-scale convolution kernels and captures long-distance dependence and global context information. The DWFA module adjusts the feature fusion ratio through a dynamic weight generation mechanism, and the DAD-Head module optimizes the feature requirements of classification and regression tasks through a dynamic weight allocation mechanism.

Benefits of technology

It significantly improves the robustness and accuracy of wheel component detection, reduces false detection and missed detection problems, can accurately detect multi-size components in complex contexts, and maintain high-level semantic information in small object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071000A_ABST
    Figure CN120071000A_ABST
Patent Text Reader

Abstract

The invention provides a wheel part detection method based on multi-scale dynamic fusion, relates to the technical field of image recognition, provides a wheel part detection process, and comprises an image preprocessing module, a backbone network module, a feature fusion module, a target detection module, an MSA-T module, a DWFA module and a DAD-Head module. The MSA-T module extracts local features of different receptive fields through a multi-scale convolution kernel, the self-attention mechanism can capture long-distance dependency and global context information, and the DWFA module adaptively adjusts a fusion proportion according to contribution degrees of different input features through a dynamic weight generation mechanism; according to the DAD-Head module, alignment and fusion of multi-scale features are achieved, high-level semantic information is not lost while the fine granularity of a small target is captured, and the DAD-Head module is specially optimized for feature requirements of a classification task and a regression task through a dynamic weight distribution mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a method for detecting wheel components based on multi-scale dynamic fusion. Background Art

[0002] In vehicle detection, the detection of wheel components is a key link for vehicle qualification. The MSA-T module extracts local features with different receptive fields through multi-scale convolutional kernels, enabling the model to effectively capture the size changes of wheel components at multiple scales. At the same time, combined with the self-attention mechanism of Transformer, MSA-T can capture the long-range dependencies and global context information between wheel components. This mechanism is particularly suitable for the situation where multiple components are closely arranged or partially occluded in complex scenes, and can effectively reduce false detection and missed detection problems, improving the robustness of object detection.

[0003] The DWFA module realizes the balance between the fine-grained information of small targets and the high-level semantic information of large targets through a dynamic weight generation mechanism, which adaptively adjusts the feature fusion ratio according to the contribution degree of different input features. The multi-scale feature alignment and fusion enable the DWFA module to perform excellently in detecting small-sized components, while ensuring that the semantic understanding of large-sized components is not lost. Its dynamic characteristics enable the model to flexibly adjust the weight of feature usage for different scenarios, significantly enhancing the detection ability of wheel components in complex backgrounds.

[0004] The DAD-Head module optimizes the classification and regression tasks separately through a dynamic weight allocation mechanism, enabling the classification branch to focus on improving the category discrimination ability, while the regression branch more precisely fits the target bounding box. For the wheel component detection scenario, this design can more accurately identify the categories of different components and improve the accuracy of target localization. Especially in small target detection, the regression branch dynamically optimizes the spatial detail features, greatly reducing the missed detection rate; in complex backgrounds, the task-specific optimization of the classification branch can effectively reduce false detection.

[0005] The combination of the MSA-T module, DWFA module and DAD-Head module not only improves the adaptability to multi-size and multi-category targets, but also significantly enhances the detection accuracy and robustness of the model in complex scenes, thus meeting the requirements of industrial detection for real-time performance, high precision and high reliability. Summary of the Invention

[0006] The present invention proposes a wheel component detection method based on multi-scale dynamic fusion, aiming to propose an MSA-T module, a DWFA module, and a DAD-Head module. The MSA-T module extracts local features with different receptive fields through multi-scale convolutional kernels, and the self-attention mechanism can capture long-range dependencies and global context information. The DWFA module adaptively adjusts the fusion ratio according to the contribution degrees of different input features through a dynamic weight generation mechanism; the alignment and fusion of multi-scale features capture fine-grained small targets without losing high-level semantic information. The DAD-Head module is optimized for the feature requirements of classification tasks and regression tasks through a dynamic weight allocation mechanism.

[0007] The present invention aims to improve on the basis of image detection and provides a wheel component detection method based on multi-scale dynamic fusion, including the following steps: S1. Production of a wheel component image data set: Take pictures of wheel components using an industrial camera to obtain wheel component images, mark the positions to be detected, where the positions to be detected are the spokes and pores, and mark other objects as irrelevant targets. After marking, a wheel component image data set is obtained; S2. Construct an image preprocessing module, including image scaling and normalization processing; S3. Construct a backbone network module, including a RepVGG module and an MSA-T module. The MSA-T module includes multi-scale convolutional feature extraction MSC, Flatten operation, Transformer self-attention calculation, and multi-scale feature reduction; S4. Construct a feature fusion module, including a DWFA module. The DWFA module includes feature alignment, dynamic weight generation, and dynamic weighted feature fusion to enhance the output features; S5. Construct an object detection module, including a DAD-Head module and a post-processing module. The DAD-Head module includes dynamic feature task allocation, a classification branch, and a regression branch.

[0008] Preferably, in step S2, for the image preprocessing module, input an image , perform normalization and size adjustment on the image to obtain the preprocessed input features, , is the preprocessing function.

[0009] Preferably, in step S3, for the backbone network module, input the image feature , where , are the height and width of the feature respectively, is the number of channels, and extract multi-level local features through RepVGG, , where Contains multi-scale features and feeds the image features into the MSA-T module. The MSA-T module includes multi-scale convolutional feature extraction MSC, Flatten operation, Transformer self-attention calculation, and multi-scale feature restoration. The image features are fed into the MSC module. Three convolutional kernels of different scales are used to extract local features respectively, generating three branch features, , , , where, , , , representing the sizes of different convolutional kernels. The features extracted by different convolutional kernels are fused into multi-scale features through channel-wise concatenation, , where ; is adjusted to have d channels using convolution and flattened into a sequence form, , where ; is projected into query Q, key K, and value V, , , , where, , , is a learnable projection matrix. The dot product of query and key is calculated and normalized by to obtain the attention weights, , where is the attention matrix. The enhanced features are obtained by weighting value V with the attention weights A, , where the weights are multiplied by the matrix , where ; The enhanced features are restored from the sequence form to a two-dimensional feature map, , where , and the input features and the enhanced features are integrated using residual connection, , where is the output of the backbone network module, containing multi-scale features and global self-attention enhanced information. The input features and the enhanced features are integrated using residual connection, is the output of the backbone network module, containing multi-scale features and global self-attention enhanced information.

[0010] Preferably, in step S4, for the feature fusion module, the output features of the backbone network module are used as input features, and the MST-A module outputs multi-scale features, , where is the low-level fine-grained feature, is the middle-level semantic feature, is the high-level semantic feature. Then, , , are input into the DWFA module to adjust the features of different scales to the same resolution and number of channels. , , , , where is the downsampling operation using a convolutional kernel k = 3, is the upsampling operation using bilinear interpolation; the aligned features are concatenated and dynamic weights are generated. , , and global pooling and a fully connected layer are used to generate fusion weights. , where represents the dynamic weights of each scale, satisfying . The features are weighted and summed using the dynamic weights. , where is the output feature after fusion. The fused features are adjusted in terms of the number of channels through convolution and a residual connection is added. . The final output feature of the feature fusion module is .

[0011] Preferably, in step S5, for the object detection module, the output features of the feature fusion module are input into the DAD-Head module. The features contain both spatial information and the contextual semantic information required for classification and regression. The DAD-Head module distributes to the classification and regression task branches through a dynamic weight mechanism, and uses global pooling and convolution to generate the feature weights for the classification and regression tasks. , is the per-channel weight, is the Sigmoid activation function used to normalize to [0, 1]. According to the generated weights, are respectively weighted to the classification branch and the regression branch. , , is the per-channel dot product, is the feature for classification, is the feature for regression. The classification branch is responsible for predicting the class probabilities of each pixel point. , where , where K is the number of target categories, and for each pixel the probability for the k-th category is , where is the Logit value of the k-th category. For each pixel, the index of the category prediction is , and the confidence is ; The regression branch predicts the target bounding box parameters corresponding to each pixel, , where , and for each pixel the output is , where , are the coordinates of the center point of the bounding box, and w and h are the width and height of the bounding box. Map the bounding box parameters on the feature map to the original image, , , , , where , are the height and width of the original image. Enter the post-processing stage, and filter out the prediction boxes with confidence lower than the threshold , , is . For the boxes with high confidence, sort them from high to low according to the confidence, and calculate the intersection over union (IoU) between the candidate boxes , . For the boxes with an IoU , only keep the box with a higher confidence , is . The final output is , where N is the number of targets after post-processing. Each target object contains the category index , the confidence , and the bounding box parameters .

[0012] Compared with the prior art, the present invention has the following technical effects: The technical solution provided by the present invention proposes an MSA-T module, a DWFA module, and a DAD-Head module. The MSA-T module extracts local features with different receptive fields through multi-scale convolutional kernels, and the self-attention mechanism can capture long-range dependencies and global context information. The DWFA module, through the dynamic weight generation mechanism, adaptively adjusts the fusion ratio according to the contribution degree of different input features; the alignment and fusion of multi-scale features capture the fine-grained details of small targets while not losing high-level semantic information. The DAD-Head module, through the dynamic weight allocation mechanism, is optimized specifically for the feature requirements of the classification task and the regression task. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is the flowchart for detecting wheel components provided by the present invention.

[0014] Figure 2 It is the backbone network structure diagram provided by the present invention.

[0015] Figure 3 It is the feature fusion structure diagram provided by the present invention.

[0016] Figure 4 It is the object detection structure diagram provided by the present invention. Specific implementation manners

[0017] The present invention proposes a method for detecting wheel components based on multi-scale dynamic fusion, aiming to propose an MSA-T module, a DWFA module, and a DAD-Head module. The MSA-T module extracts local features with different receptive fields through multi-scale convolutional kernels, and the self-attention mechanism can capture long-range dependencies and global context information. The DWFA module adaptively adjusts the fusion ratio according to the contribution degrees of different input features through a dynamic weight generation mechanism; the alignment and fusion of multi-scale features capture the fine-grained details of small targets while not losing high-level semantic information. The DAD-Head module is optimized specifically for the feature requirements of classification tasks and regression tasks through a dynamic weight allocation mechanism.

[0018] Please refer to Figure 1 as shown, the method for detecting wheel components based on multi-scale dynamic fusion in the embodiments of the present application: S1. Production of the wheel component image dataset: Take pictures of the wheel components using an industrial camera to obtain 3000 wheel component images. Locate the air holes in the wheel components, and the position to be detected is the spoke air holes. Label other objects as irrelevant targets. After labeling, obtain a wheel component image dataset with a total of 5000 images, and divide the dataset and the validation set according to a ratio of 7 to 3; S2. Construct an image preprocessing module, including image scaling and normalization processing; S3. Construct a backbone network module, including a RepVGG module and an MSA-T module. The MSA-T module includes multi-scale convolutional feature extraction MSC, Flatten operation, Transformer self-attention calculation, and multi-scale feature restoration; S4. Construct a feature fusion module, including a DWFA module. The DWFA module includes feature alignment, dynamic weight generation, and dynamic weighted feature fusion to enhance the output features; S5. Construct an object detection module, including a DAD-Head module and a post-processing module. The DAD-Head module includes dynamic feature task assignment, a classification branch, and a regression branch; Input the image into the image preprocessing module, perform image scaling and normalization on the image, input the image features into the backbone network module, extract local features with different receptive fields through the MSA-T module, the self-attention mechanism can capture long-range dependencies and global context information, and then through the DWFA module's dynamic weight generation mechanism, adaptively adjust the fusion ratio according to the contribution degree of different input features; then perform alignment and fusion of multi-scale features, capture the fine-grained details of small targets while not losing high-level semantic information, and then through the DAD-Head module, optimize according to the feature requirements of classification tasks and regression tasks through the dynamic weight allocation mechanism, and finally, output the image.

[0019] Further, in step S2, for the image preprocessing module, input an image , perform normalization and size adjustment on the image to obtain the preprocessed input features: is the preprocessing function, H, W, and C represent the height, width, and number of channels of the input features, the height and width are both 640, and the number of channels is 224.

[0020] Further, in step S3, for the backbone network module, its structure is as Figure 2 shown, input the image features , where , are the height and width of the features respectively, is the number of channels, extract multi-level local features through RepVGG, , where contains multi-scale features, input the image features into the MSA-T module, the MSA-T module includes multi-scale convolutional feature extraction MSC, Flatten operation, Transformer self-attention calculation, multi-scale feature reduction, the image features are input into the MSC module, use three different scales of convolutional kernels to extract local features respectively, and generate three branch features, , , , where , , , represent the sizes of different convolutional kernels, fuse the features extracted by different convolutional kernels through channel-wise concatenation into multi-scale features, , where ; for use convolution to adjust the number of channels to d and flatten it into a sequence form, , where ; input The projection is the query Q, key K, and value V, , , , where, , , is a learnable projection matrix that calculates the dot product of the query and the key and normalizes it through to obtain the attention weights, , where is the attention matrix, and the weighted value V using the attention weights A gives the enhanced features, , the weights and the matrix are multiplied, where ; the enhanced features are restored from the sequence form to a two-dimensional feature map, , where , and the input features and the enhanced features are integrated using a residual connection, , where is the output of the backbone network module, containing multi-scale features and global self-attention enhanced information.

[0021] Furthermore, in step S4, for the feature fusion module, its structure is as Figure 3 shown, taking the output features of the backbone network module as the input features, the MST-A module outputs multi-scale features, , where is the low-level fine-grained features, is the middle-level semantic features, is the high-level semantic features, and , , are input into the DWFA module to adjust the features of different scales to the same resolution and number of channels , , , , where, is the downsampling operation using a convolutional kernel k = 3, is the upsampling operation using bilinear interpolation; the aligned features are concatenated and a dynamic weight is generated, , and the fusion weight is generated using global pooling and a fully connected layer, where, represents the dynamic weight of each scale, satisfying , and the features are weighted and summed using the dynamic weight, , where, is the output feature after fusion. The fused feature is adjusted in terms of the number of channels through convolution and a residual connection is added. , and the final output feature of the feature fusion module is .

[0022] Furthermore, in step S5, for the object detection module, its structure is as Figure 4 shown. The output feature of the feature fusion module is input into the DAD-Head module. The feature contains both spatial information and the context semantic information required for classification and regression. The DAD-Head module distributes to the two task branches of classification and regression through a dynamic weight mechanism, and uses global pooling and convolution to generate the feature weights for the classification and regression tasks. , is the per-channel weight, is the Sigmoid activation function, which is used to normalize to [0, 1]. According to the generated weights, is weighted to the classification branch and the regression branch respectively. , , is the per-channel dot product, is the feature for classification, is the feature for regression. The classification branch is responsible for predicting the class probability of each pixel point. , where , K is the number of object classes, and for each pixel point the probability for the k-th class is , where is the Logit value for the k-th class. For each pixel point, the index of the class prediction is , and the confidence is ; the regression branch predicts the target bounding box parameters corresponding to each pixel point. , where , and for each pixel point the output is , where , are the coordinates of the center point of the bounding box, and w, h are the width and height of the bounding box. The bounding box parameters on the feature map are mapped to the original image. , , , , where , are the height and width of the original image. Entering the post-processing stage, the prediction boxes with a confidence lower than the threshold are filtered out. , is For the bounding boxes with high confidence, sort them in descending order of confidence and calculate the intersection over union (IoU) between the candidate bounding boxes. , For the bounding boxes with an IoU , only retain the bounding boxes with higher confidence. , is The final output is , where N is the number of targets after post-processing, and each target object contains a class index , confidence , and bounding box parameters .

[0023] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention.

Claims

1. A wheel parts detection method based on multi-scale dynamic fusion, characterized in that: The following steps are involved: S1. Preparation of wheel parts image dataset: Use an industrial camera to take pictures of wheel parts to obtain wheel parts images, and mark the positions that need to be detected. The positions that need to be detected are the spoke holes, and other objects are marked as irrelevant targets. After the marking is completed, a wheel parts image dataset is obtained; S2, constructing an image preprocessing module, including image scaling and normalization processing; S3, build the backbone network module, including the RepVGG module and the MSA-T module. The MSA-T module includes multi-scale convolution feature extraction MSC, Flatten operation, Transformer self-attention calculation, and multi-scale feature restoration; S4. Construct a feature fusion module, including a DWFA module. The DWFA module includes feature alignment, dynamic weight generation, dynamic weighted feature fusion, and enhanced output features; S5. Build a target detection module, including a DAD-Head module and a post-processing module. The DAD-Head module includes dynamic feature task allocation, classification branch, and regression branch. In step S2, for the image preprocessing module, input an image , normalize and resize the image to obtain the preprocessed input features, , It is a preprocessing function.

2. The wheel component detection method based on multi-scale dynamic fusion according to claim 1 is characterized in that: In step S3, for the backbone network module, the image features are input. ,in, , are the height and width of the feature, is the number of channels, and multi-level local features are extracted through RepVGG. ,in Contains multi-scale features, image features Input into the MSA-T module, which includes multi-scale convolution feature extraction MSC, Flatten operation, Transformer self-attention calculation, multi-scale feature restoration, image features The input is sent to the MSC module, and three convolution kernels of different scales are used to extract local features and generate three branch features. , , ,in, , , , represents the size of different convolution kernels, and the features extracted by different convolution kernels are fused into multi-scale features by channel-by-channel splicing. ,in ;right use The convolution adjusts the number of channels to d and flattens it into a sequence form. ,in ; Input The projection is query Q, key K, value V, , , ,in, , , is a learnable projection matrix, computing the query and key The dot product of Normalized to get the attention weight, ,in is the attention matrix, using the attention weight A weighted value V to obtain the enhanced features, , weight and matrix Multiply ; Restore the enhanced features from sequence form to a two-dimensional feature map, ,in , using residual connections to integrate input features and enhanced features, ,in The output of the backbone network module contains multi-scale features and global self-attention enhancement information.

3. The wheel component detection method based on multi-scale dynamic fusion according to claim 1, characterized in that: In step S4, for the feature fusion module, the output features of the backbone network module are used as input features, and the MST-A module outputs multi-scale features. ,in is a low-level fine-grained feature, is a mid-level semantic feature, is a high-level semantic feature. , , Input into the DWFA module to adjust the features of different scales to the same resolution and number of channels , , , ,in, It is a downsampling operation using convolution kernel k=3. It uses bilinear interpolation for upsampling; concatenates the aligned features and generates dynamic weights , , using global pooling and fully connected layers to generate fusion weights ,in, Represents the dynamic weight of each scale, satisfying ; Use dynamic weights to perform weighted summation of features, ,in, is the fused output feature; the fused feature is convolved to adjust the number of channels and add a residual connection. , the final output feature of the feature fusion module is .

4. The wheel component detection method based on multi-scale dynamic fusion according to claim 1, characterized in that: In step S5, for the target detection module, the feature fusion module outputs the feature Input into the DAD-Head module, the features contain both spatial information and contextual semantic information required for classification and regression; the DAD-Head module uses a dynamic weight mechanism to Assigned to two task branches, classification and regression, using global pooling and Convolution generates feature weights for classification and regression tasks, , is the channel-by-channel weight, is the Sigmoid activation function, which is used to normalize to [0,1]. According to the generated weights, Weighted to the classification branch and regression branch respectively, , , is the channel-wise dot product, is the feature used for classification, is the feature used for regression; The classification branch is responsible for predicting the category probability of each pixel. ,in , K is the number of target categories, each pixel The probability for the kth class is ,in is the Logit value of the kth class. For each pixel, the index of the class prediction is , the confidence level is ; The regression branch predicts the target bounding box parameters corresponding to each pixel. ,in , each pixel The output is ,in , is the coordinate of the center point of the bounding box, w and h are the width and height of the bounding box, and the bounding box parameters on the feature map are mapped to the original image. , , , ,in , is the height and width of the original image; in the post-processing stage, the images with confidence values ​​below the threshold are filtered out. The prediction box of , for For boxes with high confidence, sort them from high to low according to their confidence, and calculate the intersection and union ratio between the candidate boxes , , for the intersection ratio Only the boxes with higher confidence are retained , for The final output is , where N is the number of targets after post-processing, and each target object contains a category index , confidence , bounding box parameters .

Citation Information

Patent Citations

  • Highway small target detection method and device based on convolutional neural network

    CN115761401A

  • Guide depth map super-resolution method based on multi-scale feedback network

    CN116433487A

  • Image processing method and device, equipment, storage medium and program product

    CN116958759A

  • Remote sensing image few-sample target detection method based on scale information enhancement

    CN117456377A

  • Dense scene target detection method and system

    CN117853854A