A RV reducer pinion tooth detection method based on improved RetinaNet
By improving the RetinaNet network, combining the convolutional attention model and CIoU algorithm, the problem of low network performance in RV reducer needle teeth detection is solved, and high-precision and high-quality needle teeth detection are achieved.
Patent Information
- Application Number
- CN202210667778.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-06-14
AI Technical Summary
When the existing RetinaNet target detection model detects needle teeth in RV reducers, the network performance is low, the detection accuracy is low, and the detection effect is poor, so it cannot effectively increase the attention of important features.
Improve the RetinaNet network model, and combine it with ResNet50 to improve the importance and richness of feature extraction, and use CIoU algorithm and anchor box selection algorithm for target box regression and needle teeth classification in feature fusion network.
The performance of needle teeth detection is improved, real-time detection of the installed position and quantity of needle teeth is achieved, ensuring the quality of RV reducer products, and improving network performance and detection accuracy.
Smart Images

Figure CN115115586B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of computer vision and intelligent manufacturing, and in particular relates to a needle tooth detection method for an RV reducer based on an improved RetinaNet. Background Art
[0002] RV (Rotary Vector) reducer is the core component of industrial robots. Figure 1 The pin teeth are important parts in the cycloid pinwheel planetary transmission mechanism of the RV reducer. They are small in size and large in number. The installation position of the pin teeth is the groove between the cycloid pinwheel and the housing, and they are embedded between the two parts.
[0003] In the current RV reducer assembly process, the needle teeth are manually assembled by workers. Relying solely on the operating experience and skills of the assembly workers, it is easy to miss the needle teeth. If the problem is not discovered in time, it will affect production efficiency and product quality, and waste a lot of time and cost for the company. Therefore, during the installation of the needle teeth, the number of needle teeth installed should be detected in real time to ensure that the number of needle teeth is correct, thereby ensuring the quality of RV reducer products.
[0004] With the development of computer vision technology, more and more mechanical products have adopted visual inspection methods to ensure the accuracy of the production process. Deep learning, as an important method in the field of computer vision, has gradually been widely used in the mechanical field due to its advantages such as strong versatility and high robustness. Therefore, using the target detection model in deep learning to detect the number and position of needle teeth in the RV reducer is an effective and feasible way.
[0005] In the method of target detection based on RetinaNet network, the Chinese invention patent with the publication number of "CN113505699A" provides a ship detection method based on RetinaNet algorithm, and its technical solution is as follows: S1: Establish a ship target detection model based on RetinaNet algorithm; S2: Obtain the image of the ship target to be detected; S3: Input the image of the ship target to be detected into the ship target detection model to obtain the image of the ship target detection result. The Chinese invention patent with the publication number of "CN113159063A" provides a small target detection method based on improved RetinaNet, and its technical solution is as follows: In view of the problem of complex detection scenes, a multi-layer fusion module is added to the FPN in the RetinaNet model structure, and the multi-layer fusion can solve the problem of dilution of the top-level semantic information in the feature pyramid structure to a certain extent; in view of the problem of small targets, since the selection flexibility of small targets in the feature layer in multi-scale detection is low, it depends to a large extent on the detail information of the bottom layer of the pyramid, and the super-resolution SR technology is used to compensate the bottom feature information, so that the bottom detail information and texture information are richer.
[0006] However, the pin teeth in the RV reducer are embedded in other parts and there are many of them. The original RetinaNet target detection model does not improve the attention paid to important features, and there are technical problems such as poor network performance, low detection accuracy, and poor detection effect. Summary of the invention
[0007] The present invention provides a RV reducer pinion detection method based on improved RetinaNet, aiming to solve the problems that conventional target detection models do not improve the degree of attention paid to important features, have low network performance, low detection accuracy and poor detection effect.
[0008] In order to solve the above technical problems, the present invention makes targeted improvements to the target detection network model RetinaNet, including the following steps:
[0009] S1: Extract the required images from the needle tooth target detection dataset and input them into the improved RetinaNet target detection network model;
[0010] S2: The feature extraction network of the improved RetinaNet target detection network model extracts the needle tooth features in the image to obtain multiple feature maps of different scales;
[0011] S3: The feature fusion network of the improved RetinaNet target detection network model fuses the feature maps of different scales to improve the hierarchy and richness of features, and outputs a detection head;
[0012] S4: using the detection head to perform target frame position regression and needle tooth classification tasks respectively;
[0013] S5: Continuously iteratively executing steps S2 to S4 using the images in the needle tooth target detection dataset until a set number of training times is reached to obtain multiple models;
[0014] S6: The model is compared by the CIoU algorithm and the anchor box selection algorithm, and the optimal model is saved.
[0015] Preferably, in step S1, a camera is used to capture images of the RV reducer pin gear assembly process, and the images are used as input to step S2.
[0016] Preferably, the feature extraction network in step S2 introduces a convolutional attention model and combines it with ResNet50, and the specific steps are as follows:
[0017] S2-1: Perform a 7×7 convolution and a maximum pooling on the input image, and output the first feature map;
[0018] S2-2: After the first feature map output by S2-1 is convolved by ResNet50, a second feature map is output;
[0019] S2-3: Input the second feature map output by S2-2 into the convolutional attention model, and output a third feature map;
[0020] S2-4: The third feature map output by S2-3 is sequentially passed through four convolutional layers to output three fourth feature maps C3, C4 and C5;
[0021] S2-5: The output X4 of the last layer of the S2-4 convolutional layer is input into the convolutional attention model, and the fifth feature map C6 is obtained through 3×3 convolution.
[0022] Preferably, the convolutional attention model in step S2-3 includes a channel attention module and a spatial attention module, and step S2-3 is specifically:
[0023] S2-3-1: Input the second feature map outputted from step S2-2 into the channel attention module, and output the channel attention weight;
[0024] S2-3-2: Multiply the channel attention weight output by S2-3-1 by the feature map output by step S2-2 element by element, and output the sixth feature map;
[0025] S2-3-3: Input the sixth feature map output by S2-3-2 into the spatial attention module, and output the spatial attention weight;
[0026] S2-3-4: Multiply the spatial attention weight output by S2-3-3 by the sixth feature map output by S2-3-2 element by element, and output the seventh feature map;
[0027] S2-3-5: Pass the seventh feature map output by S2-3-4 through four convolutional layers in sequence to output the eighth feature map, and continuously iterate steps S2 to S4 until the set number of training times is reached, and output multiple eighth feature maps whose number is based on the number of iterations.
[0028] Preferably, the feature fusion network in step S3 is specifically:
[0029] S3-1: Perform ReLU activation and 3×3 convolution on the fifth feature map C6 to obtain a ninth feature map C7;
[0030] S3-2: The fifth feature map C6 and the ninth feature map C7 are directly used as output detection heads P6 and P7 of the feature fusion network;
[0031] S3-3: performing 1×1 convolution on the fourth feature maps C3, C4, and C5 to obtain tenth feature maps M3, M4, and M5 respectively;
[0032] S3-4: The tenth feature map M5 is subjected to 3×3 convolution to obtain a detection head P5;
[0033] S3-5: The tenth feature map M5 is upsampled, added to the tenth feature map M4, and then subjected to 3×3 convolution to obtain a detection head P4;
[0034] S3-6: The tenth feature map M4 is upsampled, added to the tenth feature map M3, and then subjected to 3×3 convolution to obtain the detection head P3.
[0035] Preferably, step S4 performs target frame regression and classification on the detection heads P3, P4, P5, P6 and P7 respectively.
[0036] Preferably, the target box regression task adopts a combination of a CIoU algorithm and an anchor box selection algorithm.
[0037] Preferably, the CIoU algorithm sets a minimum rectangle that can enclose the real box and the predicted box, and evaluates the distance between the two boxes, the mutual envelopment of the two boxes, and the overlap of the center points of the two boxes.
[0038] Preferably, in the anchor box selection algorithm, neatly arranged boxes are predicted boxes, and boxes arranged in an arc shape are real boxes. The higher the degree of overlap between the predicted box and the real box, the more the predicted box is considered to be a positive sample, otherwise it is a negative sample.
[0039] Compared with the prior art, the present invention has the following technical effects:
[0040] 1. The technical solution provided by the present invention applies target detection to the detection of RV reducer needle teeth. By detecting the installed position and number of needle teeth in real time, the purpose of counting the number of needle teeth and monitoring the missing installation of needle teeth can be achieved. An improved RetinaNet network is proposed to improve the expressiveness of important features, so that the performance of needle tooth detection is improved. The CIoU and anchor box selection algorithm are used to calculate the intersection of the real box and the predicted box, which effectively improves the network performance and ensures the product quality level of the RV reducer. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a schematic diagram of the installation of the pin gear in the RV reducer;
[0042] Figure 2 It is a training flow chart of a RV reducer pin tooth detection method based on improved RetinaNet provided by the present invention;
[0043] Figure 3 It is a structural schematic diagram of the training phase of a RV reducer pin-tooth detection method based on improved RetinaNet of the present invention;
[0044] Figure 4 It is a convolutional attention model schematic diagram of a RV reducer needle tooth detection method based on improved RetinaNet of the present invention;
[0045] Figure 5 It is a schematic diagram of the structure of a channel attention module of a RV reducer needle tooth detection method based on improved RetinaNet of the present invention;
[0046] Figure 6 It is a schematic diagram of the structure of a spatial attention module of a RV reducer needle tooth detection method based on improved RetinaNet of the present invention;
[0047] Figure 7 It is a schematic diagram of a CIoU algorithm of a RV reducer pin tooth detection method based on improved RetinaNet of the present invention;
[0048] Figure 8 It is a schematic diagram of an anchor frame selection algorithm for an RV reducer needle tooth detection method based on improved RetinaNet of the present invention.
[0049] In the figure: 1. Needle gear; 2. Installed needle gear; 3. RV reducer; 4. Cycloidal pinwheel; 5. Housing. DETAILED DESCRIPTION
[0050] The present invention aims to propose a method for detecting the needle teeth of an RV reducer. By detecting the installed position and number of needle teeth in real time, the purpose of counting the number of needle teeth and monitoring the missing installation of needle teeth can be achieved. To this end, the specific implementation of the present invention provides a method for detecting the needle teeth of an RV reducer based on an improved RetinaNet; provides a convolutional attention model that improves the expressiveness of important features; provides an anchor box selection algorithm; and provides a training process for detecting the needle teeth of an RV reducer based on an improved RetinaNet.
[0051] The specific process of detecting the pin teeth of the RV reducer in this embodiment includes: a data set preparation stage, a training stage and a testing stage. In the data set preparation stage, a certain number of training samples are prepared for the network to learn; Figure 2 In the training phase, the network is made to learn the needle tooth features in the training samples and save the optimal model; in the testing phase, the features of the newly input needle tooth image are directly extracted, and the threshold is set according to the optimal model saved in the training phase to obtain the needle tooth detection image.
[0052] In the dataset preparation stage, a camera is used to capture images of the RV reducer pinion gears, and then annotation software is used to annotate each pinion gear in the image and generate a label file. Finally, data enhancement operations such as rotation, mirroring, and noise addition are performed on the captured images and label files. The pinion gear target detection dataset of this embodiment contains different numbers of pinion gears and label files at different installation stages.
[0053] The training phase includes the following steps:
[0054] S1: Extract the required images from the needle tooth target detection dataset and input them into the improved RetinaNet target detection network model;
[0055] S2: Improve the feature extraction network of the RetinaNet target detection network model to extract the needle tooth features in the image and obtain feature maps of multiple different scales;
[0056] S3: Improve the feature fusion network of the RetinaNet target detection network model to fuse feature maps of different scales to improve the hierarchy and richness of features, and output the detection head;
[0057] S4: Use the detection head to perform the target frame position regression and needle teeth classification tasks respectively;
[0058] S5: Continuously iteratively execute steps S2 to S4 using the images in the needle tooth target detection dataset until the set number of training times is reached to obtain multiple models;
[0059] S6: The model is compared by CIoU algorithm and anchor box selection algorithm, and the optimal model is saved.
[0060] In step S1, a camera is used to capture images of the RV reducer pin gear assembly process, and the images are used as input to step S2.
[0061] refer to Figure 3 The training phase of the RV reducer needle tooth detection method based on improved RetinaNet shown in the present invention includes four components: feature extraction network, feature fusion network, target frame regression module and classification module.
[0062] The feature extraction network is responsible for extracting the features of the input image. This paper introduces a convolutional attention model combined with ResNet50. Compared with the feature extraction network of the original RetinaNet target detection model, it improves the expressiveness of important features and suppresses unimportant features, ultimately improving the efficiency of network model training and the detection effect of needle teeth. The specific steps are:
[0063] S2-1: Perform a 7×7 convolution and a maximum pooling on the input image, and output the first feature map;
[0064] S2-2: After the first feature map output by S2-1 is convolved by ResNet50, the second feature map is output;
[0065] S2-3: Input the second feature map output by S2-2 into the convolutional attention model and output the third feature map;
[0066] S2-4: The third feature map output by S2-3 is sequentially passed through four convolutional layers to output three fourth feature maps C3, C4 and C5;
[0067] S2-5: The output X4 of the last layer of the S2-4 convolutional layer is input into the convolutional attention model, and the fifth feature map C6 is obtained after 3×3 convolution.
[0068] Further, if Figure 4 As shown in the figure, the convolutional attention model of the feature extraction network includes a channel attention module and a spatial attention module. The convolutional attention model is combined with the ResNet50 feature extraction network when used. After passing through the convolution layer of ResNet50, the output feature map is input into the channel attention module, and then the output channel attention weight is multiplied element by element with the feature map of the convolution layer to obtain a new feature map. Then the spatial attention module is input, and the output spatial attention weight is multiplied element by element with the original feature map, and then it is sent to four convolution layers. The output X4 of the last convolution layer is convolved by 3×3 to obtain a new feature map, which is input into the channel attention module and the spatial attention module to obtain the output of the feature extraction network.
[0069] The specific steps of the convolutional attention model of the above steps S2-3 are as follows:
[0070] S2-3-1: Input the second feature map outputted in step S2-2 into the channel attention module, and output the channel attention weight;
[0071] S2-3-2: Multiply the channel attention weight output by step S2-3-1 by the feature map output by step S2-2 element by element, and output the sixth feature map;
[0072] S2-3-3: Input the sixth feature map outputted from step S2-3-2 into the spatial attention module, and output the spatial attention weight;
[0073] S2-3-4: Multiply the spatial attention weight output in step S2-3-3 by the sixth feature map output in step S2-3-2 element by element, and output the seventh feature map;
[0074] S2-3-5: Pass the seventh feature map output by step S2-3-4 through four convolutional layers in sequence to output the eighth feature map, and continuously iterate steps S2 to S4 until the set number of training times is reached, and output multiple eighth feature maps based on the number of iterations.
[0075] Further, if Figure 5 As shown, it is the channel attention module of the convolutional attention model of the feature extraction network of an embodiment of the present invention. The channel attention module first uses wide-based global maximum pooling and high-based global average pooling on the second feature map of the input in the spatial dimension, which can extract richer high-level features compared to the original RetinaNet target detection model. Then, through a shared fully connected layer composed of two fully connected layers and a Relu activation function, the correlation between the channels of the input feature map is improved. Finally, the two outputs of the shared fully connected layer are added to obtain the channel attention weight.
[0076] like Figure 6 As shown, it is the spatial attention module of the convolutional attention model of the feature extraction network of an embodiment of the present invention. The input of the spatial attention module is the feature map after the channel attention module. First, the global maximum pooling and global average pooling are used for the feature map after the channel attention module in the channel dimension, and then the two pooled output results are spliced. Then, the number of channels is reduced to 1 through a 7×7 convolution layer, and finally the spatial attention weight is obtained. The function of the spatial attention module is to find the location of important content on the feature map. Compared with the original RetinaNet target detection model, it can effectively improve the efficiency of finding the location.
[0077] In the feature fusion network, low-level features with higher resolution and fewer convolutions are efficiently fused with high-level features with very low resolution and stronger semantic information to form a feature that is more discriminative than the input image feature. The specific steps are as follows:
[0078] S3-1: Perform ReLU activation and 3×3 convolution on the fifth feature map C6 to obtain a ninth feature map C7;
[0079] S3-2: The fifth feature map C6 and the ninth feature map C7 are directly used as output detection heads P6 and P7 of the feature fusion network;
[0080] S3-3: performing 1×1 convolution on the fourth feature maps C3, C4, and C5 to obtain tenth feature maps M3, M4, and M5 respectively;
[0081] S3-4: The tenth feature map M5 is subjected to 3×3 convolution to obtain a detection head P5;
[0082] S3-5: The tenth feature map M5 is upsampled, added to the tenth feature map M4, and then subjected to 3×3 convolution to obtain a detection head P4;
[0083] S3-6: The tenth feature map M4 is upsampled, added to the tenth feature map M3, and then subjected to 3×3 convolution to obtain the detection head P3.
[0084] The target box regression module of the embodiment of the present invention adopts the CIoU algorithm to replace the standard IoU algorithm, and combines it with the anchor box selection algorithm to jointly perform the target box regression task.
[0085] like Figure 7 As shown in FIG. 1 , it is a schematic diagram of the CIoU target frame position regression method according to an embodiment of the present invention. CIoU sets a minimum rectangle that can wrap the real frame and the predicted frame to evaluate the distance between the two frames. The diagonal distance of this rectangle is c. The center point distance d of the real frame and the predicted frame is introduced to evaluate the mutual wrapping of the two frames. The aspect ratio of the real frame and the predicted frame is introduced to evaluate the overlap of the center points of the two frames. The formula of CIoU is as follows:
[0086]
[0087]
[0088]
[0089] The present invention takes into account the dense distribution of pin teeth and uses the CIoU algorithm to calculate the intersection over union (IoU) ratio of the real box and the predicted box to measure the degree of overlap between the real box and the predicted box, effectively improving the network performance.
[0090] like Figure 8As shown, it is a schematic diagram of the anchor frame selection algorithm of an embodiment of the present invention. The neatly arranged frames are prediction frames, and the frames arranged in an arc shape are real frames. The higher the degree of overlap between the prediction frame and the real frame, the more the prediction frame is considered to be a positive sample, and vice versa. In the RV reducer, the distribution of the needle teeth has a certain regularity. A total of 40 needle teeth need to be installed in the RV reducer assembly. The installation positions are distributed on the edge of a circle, and the needle teeth are generally installed in a clockwise or counterclockwise order. At the moment of installation of a certain needle tooth, the position of the installed needle tooth is the thick frame in the figure. The closer to the thick frame, the greater the possibility of the positive sample of the prediction frame, and the farther from the thick frame, the smaller the possibility of the positive sample of the prediction frame.
[0091] When filtering the prediction box through CIoU, it is necessary to consider whether the prediction box is distributed in the arc area formed by the installed needle teeth. Therefore, the center points of all real boxes in the current image are connected. When the distance between the center point of the prediction box and any point on this line is less than a certain value, the prediction box is considered to be a high-quality positive sample. This method is the anchor box selection algorithm. The positive sample set obtained by the anchor box selection algorithm is combined with the positive sample set filtered by CIoU to obtain the final positive sample set. When the original RetinaNet target detection model performs the target box regression task, it does not filter the positive and negative samples of the prediction box by the distribution law of the detected target. The embodiment of the present invention proposes a combination of the anchor box selection algorithm and the CIoU algorithm to jointly perform the target box regression task. This method can reduce the probability of low-quality samples being classified as positive samples while increasing the number of positive samples.
[0092] This method inputs the detection heads P3, P4, P5, P6, and P7 into the box regression module, and combines the CIoU algorithm and the anchor box selection algorithm to regress the target box position.
[0093] During the testing phase, an image of a certain moment in the RV reducer pinion installation process is input, and the optimal model saved during the training phase is used to set a fixed threshold to directly output the pinion detection image.
[0094] In order to verify the effectiveness of the RV reducer pinion detection method based on improved RetinaNet proposed in this paper, the original RetinaNet target detection model and the improved RetinaNet target detection model are compared in terms of pinion detection performance. The data set uses the pinion target detection data set established in the above data set production stage, and the evaluation indicators use precision and recall. The results are shown in Table 1:
[0095]
[0096]
[0097] Table 1
[0098] The precision rate is the proportion of samples predicted as positive examples to samples that are actually positive examples, and the recall rate is the proportion of samples predicted as positive examples to samples that are actually positive examples. It can be seen from Table 1 that the precision rate and recall rate of the method proposed in the present invention are higher than those of the original RetinaNet model, and the detection performance is higher.
[0099] The above is only a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several modifications and improvements without departing from the inventive concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A RV reducer pinion detection method based on improved RetinaNet, characterized in that: The following steps are involved: S1: Extract the required images from the needle tooth target detection dataset and input them into the improved RetinaNet target detection network model; S2: The feature extraction network of the improved RetinaNet target detection network model extracts the needle tooth features in the image to obtain multiple feature maps of different scales, wherein the feature extraction network introduces a convolutional attention model and is combined with ResNet50; S3: The feature fusion network of the improved RetinaNet target detection network model fuses the feature maps of different scales to improve the hierarchy and richness of features, and outputs a detection head; S4: using the detection head to perform target frame position regression and needle tooth classification tasks respectively; S5: Continuously iteratively executing steps S2 to S4 using the images in the needle tooth target detection dataset until a set number of training times is reached to obtain multiple models; S6: The model is compared by the CIoU algorithm and the anchor box selection algorithm, and the optimal model is saved. The anchor box selection algorithm is: neatly arranged boxes are predicted boxes, and boxes arranged in an arc shape are real boxes. The higher the degree of overlap between the predicted box and the real box, the more the predicted box is considered to be a positive sample, otherwise it is a negative sample.
2. The RV reducer pinion detection method based on improved RetinaNet according to claim 1 is characterized in that: In step S1, a camera is used to capture images of the RV reducer pin gear assembly process, and the images are used as input to step S2.
3. The RV reducer pinion detection method based on improved RetinaNet according to claim 1 is characterized in that: The feature extraction network introduces the convolutional attention model and combines it with ResNet50. The specific steps are as follows: S2-1: Perform a 7×7 convolution and a maximum pooling on the input image, and output the first feature map; S2-2: After the first feature map output by S2-1 is convolved by ResNet50, a second feature map is output; S2-3: Input the second feature map output by S2-2 into the convolutional attention model, and output a third feature map; S2-4: The third feature map output by S2-3 is sequentially passed through four convolutional layers to output three fourth feature maps C3, C4 and C5; S2-5: The output X4 of the last layer of the S2-4 convolutional layer is input into the convolutional attention model, and the fifth feature map C6 is obtained through 3×3 convolution.
4. The RV reducer pinion detection method based on improved RetinaNet according to claim 3 is characterized in that: The convolutional attention model in step S2-3 includes a channel attention module and a spatial attention module, and step S2-3 is specifically as follows: S2-3-1: Input the second feature map outputted from step S2-2 into the channel attention module, and output the channel attention weight; S2-3-2: Multiply the channel attention weight output by S2-3-1 by the feature map output by step S2-2 element by element, and output the sixth feature map; S2-3-3: Input the sixth feature map output by S2-3-2 into the spatial attention module, and output the spatial attention weight; S2-3-4: Multiply the spatial attention weight output by S2-3-3 by the sixth feature map output by S2-3-2 element by element, and output the seventh feature map; S2-3-5: Pass the seventh feature map output by S2-3-4 through four convolutional layers in sequence to output the eighth feature map, and continuously iterate steps S2 to S4 until the set number of training times is reached, and output multiple eighth feature maps whose number is based on the number of iterations.
5. The RV reducer pinion tooth detection method based on improved RetinaNet according to claim 3 is characterized in that: The feature fusion network in step S3 is specifically: S3-1: Perform ReLU activation and 3×3 convolution on the fifth feature map C6 to obtain a ninth feature map C7; S3-2: The fifth feature map C6 and the ninth feature map C7 are directly used as output detection heads P6 and P7 of the feature fusion network; S3-3: performing 1×1 convolution on the fourth feature maps C3, C4, and C5 to obtain tenth feature maps M3, M4, and M5 respectively; S3-4: The tenth feature map M5 is subjected to 3×3 convolution to obtain a detection head P5; S3-5: The tenth feature map M5 is upsampled, added to the tenth feature map M4, and then subjected to 3×3 convolution to obtain a detection head P4; S3-6: The tenth feature map M4 is upsampled, added to the tenth feature map M3, and then subjected to 3×3 convolution to obtain the detection head P3.
6. The RV reducer pinion tooth detection method based on improved RetinaNet according to claim 5 is characterized in that: Step S4 performs target frame regression and classification on the detection heads P3, P4, P5, P6 and P7 respectively.
7. The RV reducer pinion tooth detection method based on improved RetinaNet according to claim 6 is characterized in that: The target box regression task adopts a combination of the CIoU algorithm and the anchor box selection algorithm.
8. The RV reducer pinion tooth detection method based on improved RetinaNet according to claim 7 is characterized in that: The CIoU algorithm sets a minimum rectangle that can enclose the real box and the predicted box, and evaluates the distance between the two boxes, the mutual envelopment of the two boxes, and the overlap of the center points of the two boxes.
Citation Information
Patent Citations
Small target detection method based on improved RetinaNet
CN113159063A
Ship detection method based on RetinaNet algorithm
CN113505699A
Rapid detection device and method for combined errors of pin gear of RV reducer
CN108225188A
Training method and device of target detection model, electronic equipment and storage medium
CN109961107A