Wheel mixed package detection method and device, model and training method and device
By combining the YOLOv10-S network and the Triplet/Siamese network network and introducing a wheel mixed packet detection model with channel and spatial attention mechanism, the problems of low detection efficiency and low accuracy in the wheel mixed packet detection are solved, and accurate identification of wheel images and improvement of detection performance are achieved.
Patent Information
- Application Number
- CN202510187713.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The prior art has problems of low detection efficiency and low accuracy in wheel mixed bag detection, especially under the influence of factors such as occlusion, light changes and different angles.
A wheel mixed packet detection model combining YOLOv10-S network and Triplet/Siamese network network is adopted, and a channel attention mechanism and spatial attention mechanism are introduced to improve detection performance through feature extraction, global pooling and attention mechanism processing.
Accurate identification of wheel images is achieved, the accuracy and reliability of wheel mixed bag detection is improved, and the challenges in wheel mixed bag detection can be effectively solved.
Smart Images

Figure CN120107681A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of wheel detection, and in particular, relates to a wheel mix-up detection method and device, a model, and a training method and device. Background Art
[0002] Due to various reasons (such as human error, equipment failure, etc.), wheels may be mixed up in the same package, that is, wheels of different models or specifications are mistakenly mixed in the same package. This situation not only affects the efficiency of the production line, but may also cause serious quality problems and even endanger driving safety.
[0003] Traditional wheel mix-up detection methods mainly rely on manual inspection, which is not only time-consuming and laborious, but also easily affected by human factors, resulting in low accuracy and reliability of detection results. With the rapid development of computer vision technology, automated detection methods based on image recognition have gradually become an effective way to solve this problem.
[0004] Among existing image recognition technologies, the YOLO (You Only Look Once) series of models have attracted widespread attention due to their high efficiency and accuracy. As a variant of the YOLO series, YOLOv10-S has faster detection speed and higher accuracy, making it very suitable for real-time detection tasks. However, directly applying the YOLOv10-S network to mixed wheel detection may face some challenges, such as occlusion, lighting changes, different angles, and other factors may lead to poor detection results. Summary of the invention
[0005] The purpose of this application is to overcome the defects in the above-mentioned prior art and provide a wheel mix-up detection method and device, model and training method and device.
[0006] The present application provides a wheel mixed packet detection model, including: a YOLOv10-S network and a Triplet / Siam esenetwork network;
[0007] The Triplet / Siamese network is connected to the YOLOv10-S network. The Triplet / Siamese network uses the first 47 layers of the ResNet50 network as feature extraction layers, and introduces a channel attention mechanism module and a spatial attention mechanism module.
[0008] The channel attention mechanism module performs global maximum pooling and global average pooling on the feature maps extracted by the feature extraction layer, and inputs the result to the shared perceptron; the spatial attention mechanism module processes the feature maps output by the channel attention mechanism module, respectively.
[0009] The present application also provides a wheel mixed package detection model training method, which is used for training the above-mentioned wheel mixed package detection model, comprising:
[0010] Get the image training set, including two category labels: complete wheels and occluded wheels;
[0011] Input the image training set into the YOLOv10-S network to obtain bounding boxes and category labels;
[0012] According to the bounding box screenshot of the complete wheel category label, positive sample and negative sample labels are marked to generate a triple sample, wherein the triple sample includes an anchor point sample, a positive sample and a negative sample;
[0013] The feature extraction layer extracts the feature map of the triple sample, including the feature Figure 1 ,feature Figure 2 and Features Figure 3 Based on the characteristics Figure 1 ,feature Figure 2 and Features Figure 3 Perform global maximum pooling and global average pooling respectively, input into the channel attention mechanism module, and obtain features Figure 1-1 ,feature Figure 2-1 and Features Figure 3-1 ; The features are respectively Figure 1-1 ,feature Figure 2-1 and Features Figure 3-1 Input into the spatial attention mechanism module to obtain improved features Figure 1 , Improved Features Figure 2 and improved features Figure 3 ;
[0014] The improved features Figure 1 , Improved Features Figure 2 and improved features Figure 3 They are input into the flattening layer and the fully connected layer respectively, and then the loss function is calculated.
[0015] The present application also provides a wheel mix-up detection method, which performs wheel mix-up detection based on the wheel mix-up detection model trained by the wheel mix-up detection model training method, including:
[0016] Input the acquired wheel image into the YOLOv10-S network to obtain the bounding box, confidence, and category of the complete wheel category target;
[0017] Selecting the complete wheel category target whose confidence is above a preset confidence threshold, identifying the complete wheel category target closest to the center point of the wheel image as an anchor target, and setting other complete wheel category targets as comparison targets;
[0018] The anchor target and the comparison target are scaled, and a plurality of the comparison targets and the anchor target are formed into a matching queue and input into the Triplet / Siamese network to obtain a detection result.
[0019] Optionally, a plurality of the comparison targets and the anchor target are combined into a matching queue and input into the Triplet / Siamese network to obtain a detection result, further comprising:
[0020] If the predicted values in the detection results are all less than the prediction threshold, the wheels in the wheel image are the same and no processing is performed;
[0021] If there is a predicted value in the detection result that is greater than the prediction threshold, the comparison target is marked on the global map and the sound and light alarm is triggered.
[0022] Optionally, in the detection network, the channel attention weight value is directly referenced.
[0023] Optionally, the preset confidence threshold is set to 0.8.
[0024] Optionally, the prediction threshold is set to 0.1.
[0025] The present application also provides a wheel mixed bag detection model training device, comprising:
[0026] Get the module to get the image training set, including two category labels: complete wheel and occluded wheel;
[0027] The training module inputs the image training set into the YOLOv10-S network to obtain a bounding box and a category label; generates a triplet sample according to the bounding box and the category label, wherein the triplet sample includes an anchor sample, a positive sample, and a negative sample; extracts a feature map of the triplet sample through a feature extraction layer, including a feature Figure 1 ,feature Figure 2 and Features Figure 3 Based on the characteristics Figure 1 ,feature Figure 2 and Features Figure 3 Perform global maximum pooling and global average pooling respectively, input into the channel attention mechanism module, and obtain features Figure 1-1 ,feature Figure 2-1 and Features Figure 3-1 ; The features are respectively Figure 1-1 ,feature Figure 2-1 and Features Figure 3-1 Input into the spatial attention mechanism module to obtain improved features Figure 1 , Improved Features Figure 2 and improved features Figure 3 ; The improved features Figure 1 , Improved Features Figure 2 and improved features Figure 3 They are input into the flattening layer and the fully connected layer respectively, and then the loss function is calculated.
[0028] The present application also provides a wheel mix-up detection device, comprising:
[0029] The input module inputs the acquired wheel image into the YOLOv10-S network to obtain the bounding box, confidence and category of the complete wheel category target;
[0030] The recognition module selects the complete wheel category target whose confidence is above the preset confidence threshold, identifies the complete wheel category target closest to the center point of the wheel image as the anchor target, and sets other complete wheel category targets as comparison targets;
[0031] The detection module scales the anchor target and the comparison target, forms a matching queue with a plurality of the comparison targets and the anchor target, and inputs the queue into the Triplet / Siamese network to obtain a detection result.
[0032] Optionally, the detection module forms a matching queue with the plurality of comparison targets and the anchor target and inputs the queue into the Triplet / Siamese network to obtain a detection result, further comprising:
[0033] If the predicted values in the detection results are all less than the prediction threshold, the wheels in the wheel image are the same and no processing is performed;
[0034] If there is a predicted value in the detection result that is greater than the prediction threshold, the comparison target is marked on the global map and the sound and light alarm is triggered.
[0035] Optionally, in the detection network, the channel attention weight value is directly referenced.
[0036] Optionally, the preset confidence threshold is set to 0.8.
[0037] Optionally, the prediction threshold is set to 0.1.
[0038] The beneficial effects of this application are:
[0039] The present application provides a wheel mixed packet detection model, including: a YOLOv10-S network and a Triplet / Siamese network; the Triplet / Siamese network is connected to the YOLOv10-S network, the Triplet / Siamese network uses the first 47 layers of the ResNet50 network as the feature extraction layer, and introduces a channel attention mechanism module and a spatial attention mechanism module; the channel attention mechanism module performs global maximum pooling and global average pooling on the feature map extracted by the feature extraction layer, and inputs it into a shared perceptron; the spatial attention mechanism module processes the feature map output by the channel attention mechanism module respectively. The present application combines the advantages of the YOLOv10-S network and the Triplet / Siamese network, and introduces a channel attention mechanism and a spatial attention mechanism to improve the detection performance. By training this model, accurate recognition of wheel images can be achieved, thereby effectively solving the problem of wheel mixed packet detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of the wheel mixed bag detection model in this application;
[0041] Figure 2 It is a flow chart of wheel mixed bag detection model training in this application;
[0042] Figure 3 This is a schematic diagram of the Triplet / Siamese network training process in this application;
[0043] Figure 4 It is a schematic diagram of the wheel mixed bag detection process in this application;
[0044] Figure 5 This is a schematic diagram of the network process of the wheel mixed packet detection Triplet / Siamese network in this application;
[0045] Figure 6 It is a schematic diagram of a wheel mix-up detection model training device in this application;
[0046] Figure 7 It is a schematic diagram of the wheel mixed bag detection model detection device in this application. DETAILED DESCRIPTION
[0047] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, the embodiments are provided in order to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0048] Please refer to Figure 1 As shown, the present application relates to a wheel mixed packet detection model, including: a YOLOv10-S network and a Triplet / Siamese network, each performing a step to realize wheel mixed packet detection.
[0049] The YOLOv10-S network recognizes wheels from the global image captured by the camera.
[0050] The Triplet / Siamese network includes a Triplet network that inputs triplet samples during training and a Siamese network that inputs two groups of samples during the detection phase. The Triplet network and the Siamese network are the same neural network.
[0051] The wheels are placed on a pallet for movement. Considering that the loading process is dynamic, the present application selects wheels that are not blocked and cuts the selected multiple target wheels to form a new picture.
[0052] The Triplet / Siamese network anchors the target cropped image closest to the center of the global image, and uses the Triplet / Siamese network (twin neural network) to calculate the similarity of multiple wheel images one by one, and those that exceed the set threshold are judged as.
[0053] The Triplet / Siamese network is connected to the YOLOv10-S network, and the Triplet / Siamese network includes the first 47 layers of the ResNet50 network as feature extraction layers, and also includes a channel attention mechanism module and a spatial attention mechanism module;
[0054] The channel attention mechanism module performs global maximum pooling and global average pooling on the feature maps extracted by the feature extraction layer, and inputs the result to the shared perceptron; the spatial attention mechanism module processes the feature maps output by the channel attention mechanism module, respectively.
[0055] Please refer to Figure 2 , Figure 3As shown, the present application also relates to a wheel mixed package detection model training method, comprising:
[0056] S101, obtaining a picture training set, including two category labels: complete wheels and occluded wheels;
[0057] Train the YOLOv10-S network to recognize complete wheel category objects.
[0058] Because the loading video is transmitted to the model in real time, in order to distinguish the installed complete wheel category targets from the complete wheel category targets that are being loaded by the loaders and occluded, this application uses pictures with two category labels, complete wheels and occluded wheels, for training, so that in subsequent predictions, complete wheel category target pictures can be obtained.
[0059] S102, inputting the image training set into the YOLOv10-S network to obtain a bounding box and a category label;
[0060] After the YOLOv10-S network inputs the pictures of the two category labels of the complete wheel and the occluded wheel, it outputs a model weight file (.pt format) for subsequent reasoning of the real-time video stream.
[0061] The output result of the YOLOv10-S network includes a bounding box + category label, which is passed to the subsequent Triplet / Siamese network for feature extraction.
[0062] S103, generating a triplet sample according to the bounding box and the category label, wherein the triplet sample includes an anchor point sample, a positive sample, and a negative sample;
[0063] Based on the complete / occluded wheel ROI (Region of Interest) image blocks detected by YOLOv10-S, a triplet sample (Anchor-Positive-Negative) is generated, including: anchor sample Anchor, the reference image of the complete wheel; positive sample Positive, the complete wheel of the same wheel type (different perspective / lighting); negative sample Negative, the occluded wheel or the complete wheel of other wheel types.
[0064] Normalize and resize to a fixed resolution.
[0065] S104, extracting a feature map of the triple sample through a feature extraction layer, including feature Figure 1 ,feature Figure 2 and Features Figure 3 Based on the characteristics Figure 1 ,feature Figure 2 and Features Figure 3Perform global maximum pooling and global average pooling respectively, input into the channel attention mechanism module, and obtain features Figure 1-1 ,feature Figure 2-1 and Features Figure 3-1 ; The features are respectively Figure 1-1 ,feature Figure 2-1 and Features Figure 3-1 Input into the spatial attention mechanism module to obtain improved features Figure 1 , Improved Features Figure 2 and improved features Figure 3 ; The improved features Figure 1 , Improved Features Figure 2 and improved features Figure 3 They are input into the flattening layer and the fully connected layer respectively, and then the loss function is calculated.
[0066] Specifically, a group of three pictures is used as a triplet sample: anchor sample anchor, positive sample posit, and negative sample negative for training.
[0067] The reason for choosing the Triplet / Siamese network is that through the loss function, the distance between the same categories is made as small as possible, and the distance between different categories is made as large as possible, which works better for samples with small differences.
[0068] In this application, the loss function uses the triplet loss function Triplet / Triplet / Siamese loss.
[0069] Triplet / Triplet / Siamese loss distinguishes the details in the sample and better models the details. The goal of the above model is to distinguish wheel images of different wheel types. Most of the differences between wheel types are fine-grained features.
[0070] L=max(d(a,p)-d(a,n)+margin,0)
[0071] Among them, "margin" is a preset threshold, a is the anchor sample anchor, d(a,p) and d(a,n) are the distances between positive samples and positive samples, and between positive samples and negative samples in the embedding space, respectively. In this application, Euclidean distance is used.
[0072]
[0073] The Triplet / Siamese network first extracts features from the input image.
[0074] In this application, the first 47 layers of the ResNet50 convolutional neural network are selected for feature map extraction.
[0075] During the training phase, in order to enable the feature detection of the Triplet / Siamese network to focus on the differences in the characteristic structure of the wheel in key areas, the CBAM convolutional attention module was introduced to extract the feature layer of the image.
[0076] The CBAM (Convolutional Block Attention Module) module is composed of two small modules connected in series, one is the channel attention module and the other is the spatial attention module.
[0077] In this application, the channel attention module focuses on which channels (i.e., feature categories) in the feature map are more important to the final result, thereby giving these channels higher weights. The expression is as follows:
[0078]
[0079] The Spatial Attention Module focuses on which spatial locations in the feature map contain more critical information, thereby giving higher weights to these locations. The expression is as follows:
[0080]
[0081] Since all three pictures in the training stage are wheels, theoretically the features that need to be focused on are the same. Therefore, in the channel attention module, the feature maps of the three pictures are summed by global maximum pooling and global average pooling respectively, and then input into the shared multi-layer perceptron (Share MLP).
[0082] The spatial features cannot be uniformly extracted due to the position, orientation and natural light of the wheels when the picture was taken, so they are input into the spatial attention module for processing.
[0083] Please refer to Figure 4 , Figure 5 As shown, the present application also provides a wheel mixed bag detection method, comprising:
[0084] S201, input the acquired wheel image into the YOLOv10-S network to obtain the bounding box, confidence and category of the complete wheel category target;
[0085] First, input the camera image into the YOLOv10-S network to obtain the bounding box coordinates, confidence level, and category of the complete wheel category target.
[0086] Preferably, considering the actual situation of wheel placement target overlap, the IOU intersection-over-union ratio threshold of the YOLO model is most appropriately set to 0.2.
[0087] S202, selecting the complete wheel category target whose confidence level is above a preset threshold, identifying the complete wheel category target closest to the center point of the wheel image as an anchor target, and setting other complete wheel category targets as comparison targets;
[0088] Take the target with a complete wheel category confidence of more than 0.8, and identify the target closest to the center point as the anchor target. For example, the predicted bounding box (x1, y1, x2, y2) is the coordinate value of the upper left corner and the lower right corner, and (w, h) is the width and height of the global photo. Traverse the bounding box of each target, calculate d = |w / 2-(x1+x1) / 2|+|h / 2-(y1+y1) / 2)|, and take the target with the smallest d value. Crop it according to its bounding box coordinates as the anchor target image, and crop other predicted targets according to the bounding box to form n comparison target images.
[0089] S203, scaling the anchor target and the comparison target, forming a matching queue with a plurality of the comparison targets and the anchor target and inputting the queue into the Triplet / Siamese network to obtain a detection result.
[0090] All target images are scaled to a size of 512×512.
[0091] The n comparison target images and the anchor target images are respectively formed into matching pairs and input into the improved Triplet / Siamese network (the model parameters are taken from the results of the corresponding modules of the training model) for prediction.
[0092] In this application, the prediction model graph is different from the training model. Only two pictures are input for comparison. Because of the single characteristic of the recognition target and the relatively fixed characteristics of the wheel, the channel attention weight value is directly referenced for calculation.
[0093] Finally, the Euclidean distance between the transformed one-dimensional vectors of the two images is calculated.
[0094]
[0095] The prediction result is a value between 0 and 1. The closer the value is to 0, the more similar the two images are. Finally, based on the actual test results, the prediction value less than 0.1 is judged as the same wheel type.
[0096] The prediction results of the anchor target image and the comparison target image are obtained. If the prediction values are all less than 0.1, it means that the wheels on the pallet are of the same type and no processing is done. If the prediction value is greater than 0.1, the comparison target image is marked with a red border on the global map and displayed on the large-screen display, and a switch value is given to the sound and light alarm to make the sound and light alarm sound.
[0097] Please refer to Figure 6 As shown, the present application also provides a wheel mixed bag detection model training device, comprising:
[0098] An acquisition module 301 acquires a training set of images, including two category labels: complete wheels and occluded wheels;
[0099] Training module 302, inputs the image training set into the YOLOv10-S network to obtain a bounding box and a category label; generates a triplet sample according to the bounding box and the category label, the triplet sample includes an anchor sample, a positive sample and a negative sample; extracts a feature map of the triplet sample through a feature extraction layer, including a feature Figure 1 ,feature Figure 2 and Features Figure 3 Based on the characteristics Figure 1 ,feature Figure 2 and Features Figure 3 Perform global maximum pooling and global average pooling respectively, input into the channel attention mechanism module, and obtain features Figure 1-1 ,feature Figure 2-1 and Features Figure 3-1 ; The features are respectively Figure 1-1 ,feature Figure 2-1 and Features Figure 3-1 Input into the spatial attention mechanism module to obtain improved features Figure 1 , Improved Features Figure 2 and improved features Figure 3 ; The improved features Figure 1 , Improved Features Figure 2 and improved features Figure 3 They are input into the flattening layer and the fully connected layer respectively, and then the loss function is calculated.
[0100] Please refer to Figure 7 As shown, the present application also provides a wheel mixed bag detection device, comprising:
[0101] Input module 401, inputs the acquired wheel image into the YOLOv10-S network to obtain the bounding box, confidence and category of the complete wheel category target;
[0102] The identification module 402 selects the complete wheel category target whose confidence is above a preset threshold, identifies the complete wheel category target closest to the center point of the wheel image as an anchor target, and sets other complete wheel category targets as comparison targets;
[0103] The detection module 403 scales the anchor target and the comparison target, and forms a matching queue with a plurality of the comparison targets and the anchor target, and inputs the queue into the Triplet / Siamese network to obtain a detection result.
[0104] Furthermore, the detection module forms a matching queue of the plurality of comparison targets and the anchor target and inputs the matching queue into the Triplet / Siamese network to obtain a detection result, and further includes:
[0105] If the predicted values in the detection results are all less than 0.1, the wheels in the wheel images are the same and no processing is performed;
[0106] If there is a predicted value greater than 0.1 in the detection result, the comparison target is marked on the global map and the sound and light alarm is triggered.
[0107] Furthermore, in the detection network, the channel attention weight value is directly referenced.
Claims
1. A wheel mix-up detection model, characterized in that: include: YOLOv10-S network and Triple t / Siamesenetwork network; The Triplet / Siamese network is connected to the YOLOv10-S network. The Triplet / Siamese network uses the first 47 layers of the ResNet50 network as feature extraction layers, and introduces a channel attention mechanism module and a spatial attention mechanism module. The channel attention mechanism module performs global maximum pooling and global average pooling on the feature maps extracted by the feature extraction layer, and inputs the result to the shared perceptron; the spatial attention mechanism module processes the feature maps output by the channel attention mechanism module, respectively.
2. A wheel mixed bag detection model training method, characterized in that: The training method for the wheel mixed bag detection model according to claim 1 comprises: Get the image training set, including two category labels: complete wheels and occluded wheels; Input the image training set into the YOLOv10-S network to obtain bounding boxes and category labels; Generate a triplet sample according to the bounding box and the category label, the triplet sample comprising an anchor sample, a positive sample, and a negative sample; Extracting feature graphs of the triple sample through the feature extraction layer, including feature graph 1, feature graph 2 and feature graph 3; performing global maximum pooling and global average pooling based on the feature graph 1, feature graph 2 and feature graph 3, respectively, inputting them into the channel attention mechanism module, and obtaining feature graph 1-1, feature graph 2-1 and feature graph 3-1; respectively inputting the feature graph 1-1, feature graph 2-1 and feature graph 3-1 into the spatial attention mechanism module, and obtaining improved feature graph 1, improved feature graph 2 and improved feature graph 3; The improved feature map 1, the improved feature map 2 and the improved feature map 3 are respectively input into the tiled layer and the fully connected layer in sequence, and then the loss function is calculated.
3. A wheel mix-up detection method, characterized in that: Performing wheel mixed package detection based on the wheel mixed package detection model of claim 1 trained by the wheel mixed package detection model training method of claim 2 includes: Input the acquired wheel image into the YOLOv10-S network to obtain the bounding box, confidence, and category of the complete wheel category target; Select the complete wheel category target whose confidence is above a preset threshold, identify the complete wheel category target closest to the center point of the wheel image as the anchor target, and set other complete wheel category targets as comparison targets; The anchor target and the comparison target are scaled, and a plurality of the comparison targets and the anchor target are formed into a matching queue and input into the Triplet / Siamese network to obtain a detection result.
4. A wheel mix-up detection method according to claim 3, characterized in that: include: The plurality of comparison targets and the anchor target are combined into a matching queue and input into the Triplet / Siamese network to obtain a detection result, further comprising: If the predicted values in the detection results are all less than 0.1, the wheels in the wheel images are the same and no processing is performed; If there is a predicted value greater than 0.1 in the detection result, the comparison target is marked on the global map and the sound and light alarm is triggered.
5. A wheel mix-up detection method according to claim 3, characterized in that: In the Tripl et / Siamese network, the channel attention weight value is directly referenced.
6. A wheel mix-up detection method according to claim 3, characterized in that: The preset threshold is set to 0.
8.
7. A wheel mixed bag detection model training device, characterized in that: include: Get the module to get the image training set, including two category labels: complete wheel and occluded wheel; A training module, inputting the image training set into a YOLOv10-S network to obtain a bounding box and a category label; generating a triple sample according to the bounding box and the category label, wherein the triple sample includes an anchor sample, a positive sample, and a negative sample; extracting feature maps of the triple sample through a feature extraction layer, including feature map 1, feature map 2, and feature map 3; performing global maximum pooling and global average pooling based on the feature map 1, feature map 2, and feature map 3, respectively, and inputting them into a channel attention mechanism module to obtain feature map 1-1, feature map 2-1, and feature map 3-1; respectively inputting the feature map 1-1, feature map 2-1, and feature map 3-1 into a spatial attention mechanism module to obtain improved feature map 1, improved feature map 2, and improved feature map 3; The improved feature map 1, the improved feature map 2 and the improved feature map 3 are respectively input into the tiled layer and the fully connected layer in sequence, and then the loss function is calculated.
8. A wheel mix-up detection device, characterized in that: include: The input module inputs the acquired wheel image into the YOLOv10-S network to obtain the bounding box, confidence and category of the complete wheel category target; A recognition module selects the complete wheel category target whose confidence is above a preset threshold, identifies the complete wheel category target closest to the center point of the wheel image as an anchor target, and sets other complete wheel category targets as comparison targets; The detection module scales the anchor target and the comparison target, forms a matching queue with a plurality of the comparison targets and the anchor target, and inputs the queue into the Triplet / Siamese network to obtain a detection result.
9. A wheel mix-up detection device according to claim 8, characterized in that: include: The detection module forms a matching queue with the plurality of comparison targets and the anchor target and inputs the queue into the Triplet / Siamese network to obtain a detection result, and further includes: If the predicted values in the detection results are all less than 0.1, the wheels in the wheel images are the same and no processing is performed; If there is a predicted value greater than 0.1 in the detection result, the comparison target is marked on the global map and the sound and light alarm is triggered.
10. The wheel mix-up detection device according to claim 8, characterized in that: In the detection network, the channel attention weight value is directly referenced.
Citation Information
Patent Citations
Shielding object detection method and device
CN114187491A
Classroom behavior identification method based on deformable multi-scale adaptive detection network
CN119296174A
General target detection method for adaptive attention guidance mechanism
WO2021139069A1