Gear tooth surface defect online detection system and method based on machine vision

The online gear tooth surface defect detection system, which combines machine vision with a rotary encoder and an industrial camera, solves the problem of image acquisition system design and the difficulty in balancing detection accuracy and efficiency in existing technologies, and realizes efficient and automated gear defect detection and quantitative analysis.

CN120807413APending Publication Date: 2025-10-17BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510852770.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The existing gear tooth surface defect detection technology is difficult to meet the online detection requirements in the design of image acquisition systems. It is difficult to balance defect detection accuracy and efficiency. In addition, the system integration and operation convenience are insufficient, and there is a lack of defect quantitative analysis capabilities.

Method used

A machine vision-based online gear tooth surface defect detection system is adopted, including an image acquisition module, a defect detection module and an information output module. A rotary encoder is used in conjunction with a low-speed industrial camera to realize high-speed rotating tooth surface positioning and capture. ICNet and the improved DDRNet model are combined for effective tooth surface segmentation and defect detection. An automatic defect parameter quantification module is developed, and a visual detection platform is developed.

Benefits of technology

It realizes the automated detection of gear defects, improves detection accuracy and efficiency, reduces costs, overcomes the defect size measurement error caused by field of view differences, simplifies the operation process, and provides an efficient and reliable intelligent solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807413A_ABST
    Figure CN120807413A_ABST
Patent Text Reader

Abstract

The invention provides a gear tooth surface defect online detection system and method based on machine vision, and the system comprises a rotating gear tooth surface image in-situ collection module which is used for positioning and capturing a specific tooth surface image, and a defect detection module which is used for achieving effective tooth surface region segmentation and defect precise detection. The information output module is used for providing visual display of a detection result; according to the invention, both detection precision and efficiency are considered, automatic detection of gear defects can be realized, and the method has the advantages of high defect detection precision and high efficiency; meanwhile, the defect size measurement error caused by the view field difference of the tooth surface is effectively overcome, the consistency of the image quality is ensured, and the automation level and the detection precision of the system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of defect detection, and particularly relates to a gear tooth surface defect online detection system and method based on machine vision. BACKGROUND

[0002] Gear is a core rotating part in advanced aerospace vehicles, wind turbines and other high-end motion equipment, and is a key basic part of motion and power transmission. The quality of the tooth surface directly affects the transmission efficiency and the service life of the equipment. Wear and damage of the tooth surface can cause performance degradation and even trigger the failure of the whole machine. Therefore, tooth surface quality monitoring and fault prediction are important research directions for modern industrial equipment maintenance and quality control.

[0003] Existing detection methods mainly include manual detection and traditional visual analysis. Manual detection has problems such as low efficiency, high cost, unstable precision, etc., and is usually offline detection. Although the traditional visual method improves the efficiency, it is very dependent on manual feature design (such as filtering, threshold segmentation), and often has problems such as incomplete feature extraction and poor anti-interference ability when dealing with complex background and small defects, and it is difficult to meet the industrial demand in terms of high precision, automation and real-time performance.

[0004] In recent years, deep learning technology has been introduced into gear tooth surface defect detection, but existing methods have not designed an efficient rotating gear tooth surface image in-situ acquisition system, and cannot realize the functions of gear positioning and image snapshot. Only offline detection can be achieved, or high-speed cameras and other equipment are used, which not only has high cost, but also the effective information in the image is not complete, and the image data is redundant, which wastes computing resources and slows down the defect analysis speed, and is not suitable for online long-term measurement.

[0005] The existing deep learning-based gear surface defect detection methods usually adopt target detection networks (such as the YOLO series) or semantic segmentation networks (such as U-Net, FPN, Deeplabv3+). Although the target detection network performs excellently in speed and lightweight, such as the Chinese patent with publication number CN 117455836 A which utilizes an improved YOLO model for gear end surface defect detection, its output is usually a bounding box and class probability, which cannot accurately quantify the actual size and area ratio of the defect, making it difficult to meet the demand for quantitative analysis of defects in industrial detection. On the other hand, existing semantic segmentation networks, such as the Chinese patent with publication number CN 119229106 A which utilizes an improved FPN network for defect detection of industrial appearance, can achieve defect quantification through pixel-level segmentation, but high-resolution image processing results in large model parameter quantity and high computational complexity, making it difficult to be deployed in real time on resource-constrained edge devices. In addition, although existing lightweight segmentation networks reduce the computational demand, they often sacrifice segmentation accuracy, especially the ability to identify small defects, resulting in unreliable detection results. This contradiction between precision and efficiency, as well as the limitations of edge device deployment, seriously restricts the practical application of existing methods in industrial online detection.

[0006] The existing gear surface defect detection technology has significant deficiencies in key links. First, the image acquisition system design cannot meet the online detection requirements: existing solutions rely on offline detection or expensive high-speed cameras, which cannot accurately position the gear surface and capture high-quality images when the gear is rotating at high speed, resulting in data redundancy, missing of effective information, and high cost. Second, the precision and efficiency of defect detection are difficult to balance: methods based on target detection networks (such as the YOLO series) can only output bounding boxes, which cannot quantify the actual size of defects; while existing semantic segmentation networks (such as improved U-Net, Deeplabv3+) have pixel-level segmentation capabilities, but the model parameter quantity is large and the calculation is complex, and after lightweight improvement, the precision is greatly reduced, especially for small defects, the edge segmentation is fuzzy, and it is difficult to meet the real-time requirements of edge devices. In addition, existing technologies lack quantitative analysis of defects: mask image post-processing relies on manual intervention, cannot automatically extract defect area, number, and damage ratio, and does not establish a mapping relationship between pixels and actual size, resulting in significant quantitative error. Finally, the system integration and operation convenience are insufficient: the hardware control, model deployment, and result analysis processes are fragmented, lack of visual interactive interface, high operation threshold, and low efficiency. SUMMARY

[0007] To solve the problems of low efficiency, insufficient precision, and low automation in traditional gear detection, the present application provides a gear surface defect online detection system and method based on machine vision, which can improve the detection precision of small defects without increasing the computational complexity.

[0008] A machine vision-based online gear tooth surface defect detection system includes an image acquisition module, a defect detection module, and an information output module; wherein the defect detection module includes a preprocessing unit, an effective tooth surface segmentation network, and a defect detection network;

[0009] The image acquisition module is used to acquire the image of the tooth surface to be measured;

[0010] The preprocessing module is used to perform noise removal and contrast enhancement processing on the tooth surface image to be measured to obtain a preprocessed tooth surface image;

[0011] The effective tooth surface segmentation network is used to extract the position of the gear in the preprocessed tooth surface image, and segment the gear according to the position to obtain the effective tooth surface image;

[0012] The defect detection network is used to extract global features and local features of the effective tooth surface image, and generate a category label corresponding to each pixel in the effective tooth surface image based on the fused global features and local features;

[0013] The information output module is used to obtain the area of ​​each category of defects based on the category label corresponding to each pixel in the effective tooth surface image.

[0014] Furthermore, the image acquisition module includes a rotary encoder, a data acquisition card, an industrial camera, a host computer and a dedicated light source;

[0015] The rotary encoder is installed on the gear spindle to measure the gear rotation angle in real time and send the rotation angle to the host computer through the data acquisition card;

[0016] The industrial camera is installed in a position facing the gear in the gear wear testing machine through a rigid support structure;

[0017] The dedicated light source is used to provide lighting for the gear wear testing machine;

[0018] The host computer is used to trigger the industrial camera to collect the image of the tooth surface to be measured when the gear rotates to a set angle.

[0019] Furthermore, the host computer determines whether the gear rotates to the set angle by:

[0020] The host computer reads the angle value θ output by the rotary encoder at the current moment in real time during the continuous rotation of the gear. k , and the angle value θ is filtered by an exponential weighted average filter k Perform smoothing to obtain the smoothed angle value at the current moment

[0021]

[0022] in, is a smoothed angle value of the previous time, and k is a smoothing coefficient of the current time for adjusting the response sensitivity.

[0023] determining whether the smoothed angle value of the current time satisfies the following trigger condition, if yes, indicating that the gear is rotated to the set angle, if no, indicating that the gear is not rotated to the set angle:

[0024]

[0025] wherein, Δθ is a theoretical angle interval corresponding to a single gear tooth, and ∈ is an allowable error tolerance.

[0026] Further, the setting method of the smoothing coefficient of the current time k is as follows:

[0027]

[0028] wherein, τ is a set threshold.

[0029] Further, the effective tooth surface segmentation network adopts a three-level resolution cascade network of an ICNet model to extract the position of the gear in the pretreated tooth surface image, and segments the gear according to the position to obtain an effective tooth surface image.

[0030] Further, the defect detection network comprises a first EMAM network, a second EMAM network, a context branch network, a spatial branch network, a dot product unit, a first fusion unit, a second fusion unit, and a segmentation head.

[0031] The first EMAM network is configured to capture features of the effective tooth surface image for the first time through multi-scale convolution and attention mechanism to obtain first feature maps of different scales.

[0032] The second EMAM network is configured to capture features of the first feature maps for the second time through multi-scale convolution and attention mechanism to obtain second feature maps of different scales.

[0033] The context branch network is configured to extract global information of the second feature maps through multi-layer residual basic blocks and residual bottleneck blocks to obtain global features.

[0034] The spatial branch network is configured to extract local information of the second feature maps through multi-layer residual basic blocks and residual bottleneck blocks to obtain local features.

[0035] The dot product unit is configured to fuse the global features and the local features through pixel point multiplication to obtain first fusion features.

[0036] The first fusion unit is configured to fuse the first fusion feature and the second feature map through a skip connection and obtain a second fusion feature through transposed convolution upsampling;

[0037] The second fusion unit is configured to fuse the second fusion feature and the first feature map through a skip connection and obtain a third fusion feature through transposed convolution upsampling;

[0038] The segmentation head is configured to generate a class label corresponding to each pixel in the effective tooth surface image based on the third fusion feature.

[0039] Further, the method for any one EMAM network to capture any feature map through multi-scale convolution and attention mechanism is as follows:

[0040] The feature image sequentially passes through 3x3 convolution, batch normalization and ReLU activation, and then enters the Efficient Channel Attention module. The Efficient Channel Attention module uses adaptive average pooling, 1x1 convolution and Sigmoid activation to enhance the channel attention of the feature image, and obtains an initial enhanced feature map;

[0041] Three dilated convolutions with different dilation rates are respectively used to operate on the initial enhanced feature map, wherein batch normalization and ReLU activation are followed after each dilated convolution;

[0042] 1x1 convolution is used to fuse the initial enhanced feature map and the outputs of the three dilated convolutions, and the fusion result is subjected to spatial dimension reduction through down-sampling convolution to obtain a first branch reduced feature map;

[0043] Cascade 3x3 convolution and 1x1 convolution are used to operate on the initial enhanced feature map to obtain a second branch reduced feature map;

[0044] The product of the first branch reduced feature map and the second branch reduced feature map is taken as the feature map of different scales output by the EMAM network.

[0045] Further, the following branch network includes a first residual basic block RB, a second residual basic block RB, a third residual basic block RB, a first residual bottleneck block RBB and a DAPPM network in sequence; and the spatial branch network includes a first residual basic block RB, a fourth residual basic block RB, a fifth residual basic block RB and a second residual bottleneck block RBB in sequence.

[0046] Among them, the following branch network and the spatial branch network share the first residual basic block RB, and the input of the third residual basic block RB in the following branch network is the output of the second residual basic block RB and the output of the fourth residual basic block RB in the spatial branch network. The input of the first residual bottleneck block RBB is the output of the third residual basic block RB and the output of the fifth residual basic block RB in the spatial branch network; the output of the DAPPM network is the final global feature;

[0047] The input of the fifth residual basic block RB in the spatial branch network is the output of the fourth residual basic block RB and the output of the second residual basic block RB in the following branch network. The input of the second residual bottleneck block RBB is the output of the fifth residual basic block RB and the output of the third residual basic block RB; the output of the second residual bottleneck block RBB is the final local feature.

[0048] Furthermore, the loss function L used in training the defect detection module is t as follows:

[0049] L t =L c +λL s +βL consistency

[0050] Among them, λ is the spatial branch loss weight, β is the consistency regularization weight, L consistency is the consistency loss term, L c is the context branch loss term, L s is the spatial branch loss term;

[0051] Among them, L consistency The calculation method is:

[0052]

[0053] Among them, N is the total number of pixels, p c,i is the probability that the context branch network predicts the i-th pixel as the true category, p s,i is the probability that the spatial branch network predicts the i-th pixel as the true category;

[0054] Spatial branch loss term L s The calculation method is:

[0055]

[0056] in, For the spatial branch network, the i s pixels are predicted as the true class The probability of For the i sThe uncertainty weight of the spatial branch loss term corresponding to pixels, K s is a set value, indicating the top K with the largest spatial branch loss s pixels;

[0057] Context branch loss term L c The calculation method is:

[0058]

[0059] in, The context branch network will be i c pixels are predicted as the true class The probability of For the i c The uncertainty weight of the context branch loss term corresponding to pixels, K c is a set value, indicating the top K with the largest context branch loss c pixels;

[0060] Uncertainty weight and The method for determining is:

[0061] w i =1+εH i

[0062]

[0063] Among them, ε is the uncertainty weight adjustment coefficient, w i for or H i is the predicted entropy corresponding to the i-th pixel, C is the total number of defect categories, is the probability that the i-th pixel is predicted to be the c-th defect category.

[0064] Furthermore, a method for online detection of gear tooth surface defects based on machine vision comprises the following steps:

[0065] Collecting images of the tooth surface to be measured;

[0066] Perform noise removal and contrast enhancement processing on the tooth surface image to be measured to obtain a pre-processed tooth surface image;

[0067] Extracting the position of the gear in the preprocessed tooth surface image, and segmenting the gear according to the position to obtain a valid tooth surface image;

[0068] Extract the global features and local features of the effective tooth surface image, and generate the category label corresponding to each pixel in the effective tooth surface image based on the fused global features and local features;

[0069] According to the category label corresponding to each pixel in the effective tooth surface image, the area of each category defect is obtained.

[0070] Advantages:

[0071] 1. The application provides a gear tooth surface defect online detection system based on machine vision, which comprises a rotating gear tooth surface image in-situ acquisition module for positioning and specific tooth surface image snapshot, a defect detection module for realizing effective tooth surface area segmentation and accurate defect detection, and an information output module for providing visual display of detection results.

[0072] 2. The application provides a gear tooth surface defect online detection system based on machine vision, which realizes automatic positioning, focusing and snapshot of high-speed rotating tooth surface through cooperation of a rotary encoder and a low-speed industrial camera, eliminates field difference and improves image consistency, and has the advantages of low cost and low redundancy.

[0073] 3. The application provides a gear tooth surface defect online detection system based on machine vision, which constructs a high-precision lightweight semantic segmentation model, segments the effective tooth surface area through a three-stage resolution cascade network of the first-stage ICNet model, and detects defects in the segmented effective tooth surface through the improved DDRNet model of the second stage.

[0074] 4. The application provides a gear tooth surface defect online detection system based on machine vision, which develops a defect parameter automatic quantization module based on OpenCV connected domain analysis and pixel-physical scale mapping to realize accurate calculation of defect number, area and damage ratio. BRIEF DESCRIPTION OF DRAWINGS

[0075] Figure 1 The flow chart of the working principle of the system of the application;

[0076] Figure 2 It is a schematic diagram of the image acquisition module of the system of the application;

[0077] Figure 3 It is a schematic diagram of the two-stage model of the system of the application;

[0078] Figure 4 It is a schematic diagram of the EMAM module of the system of the application;

[0079] Figure 5 It is a schematic diagram of the DAMMP module of the system of the application;

[0080] Figure 6 It is a schematic diagram of the defect detection result of the system of the application;

[0081] Figure 7 It is a schematic diagram of the user interface designed by the system of the application. DETAILED DESCRIPTION

[0082] In order for those skilled in the art to better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0083] The application proposes a gear tooth surface defect online detection system and method based on machine vision, which adopts a combination of deep learning technology and machine vision, covering the whole process from image acquisition, defect detection to information output, as shown in Figure 1 Specifically, it is:

[0084] A gear tooth surface defect online detection system based on machine vision, comprising an image acquisition module, a defect detection module and an information output module; wherein the defect detection module comprises a preprocessing unit, an effective tooth surface segmentation network and a defect detection network;

[0085] The image acquisition module is used to acquire the tooth surface image to be detected;

[0086] The preprocessing module is used to remove noise and enhance contrast in the tooth surface image to be detected, to obtain a preprocessed tooth surface image;

[0087] The effective tooth surface segmentation network is used to extract the position of the gear in the preprocessed tooth surface image, and segment the gear according to the position, to obtain an effective tooth surface image;

[0088] The defect detection network is used to extract the global features and local features of the effective tooth surface image, and generate the class labels corresponding to each pixel in the effective tooth surface image according to the fused global features and local features;

[0089] The information output module is used to obtain the area of each type of defect according to the class labels corresponding to each pixel in the effective tooth surface image.

[0090] The composition and processing flow of each module are described in detail below.

[0091] 1. Image acquisition module

[0092] As shown in Figure 2 , the image acquisition module includes a rotary encoder, a data acquisition card, a computer, an industrial camera, and a dedicated light source. The rotary encoder is used to measure the rotation angle of the gear in real time, ensuring the timing and angle of each image acquisition are accurate. The data acquisition card is connected to the computer through a USB interface and is responsible for real-time transmission and processing of the angle signal from the encoder.

[0093] The image acquisition device installs the camera in the position directly opposite the gear in the gear wear testing machine through a rigid support structure. To enhance the visibility of minor defects, the brightness of the light source is adjusted to ensure moderate contrast and brightness of the image, and a diffuser plate is used to reduce surface reflection. The rotary encoder is installed on the gear spindle and must ensure that there is no relative sliding between its housing and the gear shaft. It is isolated from the reducer by an insulating film to reduce electrostatic interference. According to Figure 2 and equipment identification, the connections between the rotary encoder, data acquisition card, industrial camera, and computer are completed. It is necessary to ensure normal signal transmission between each hardware. The camera focal length and field of view range are precisely calibrated to ensure that the captured images are clear and complete, especially the gear surface details.

[0094] The system synchronizes with the angular motion of the gear through the rotary encoder and triggers the industrial camera to shoot when the gear rotates to the specified angle. Each tooth image data corresponds to a fixed rotation angle, ensuring comprehensive coverage of the gear surface information. The specific acquisition logic is as follows:

[0095] A fixed interval triggering mechanism based on absolute angle matching is adopted to avoid the acquisition drift problem caused by error accumulation in traditional relative angle difference triggering. At the beginning of program startup, the system first waits for the rotation angle to return to the vicinity of 0°, with a tolerance range of ±1°, which serves as the synchronization starting point. Once the angle is identified to fall within this range, it is considered as triggering the system to align with the initial angle, indicating successful angle synchronization, and the system starts the image acquisition control process.

[0096] Subsequently, the system reads the angle value θ k output by the rotary encoder in real time during the continuous rotation of the gear and judges whether it meets the following triggering conditions:

[0097] |θ k modΔθ|≤∈

[0098] Wherein, Δθ represents the theoretical angle interval corresponding to a single gear tooth (for example, 27.5°), and ∈ is the allowable error tolerance (such as ±0.5°). Whenever the current angle is an integer multiple of Δθ (that is: 0°, 27.5°, 55.0°, 82.5°, …), that is, the effective acquisition time, the system immediately sends a digital trigger pulse signal to the industrial camera to complete the acquisition of a frame of image.

[0099] To enhance the anti-interference ability of the system, the encoder signal is first smoothed in real time by an exponential weighted average filter (EMA) before being sent to the control logic, and the formula is:

[0100]

[0101] Wherein, α k is a smoothing coefficient for adjusting the response sensitivity, and the typical value range is 0.3-0.8. Further, to resist false triggering caused by peak interference, the present application designs an adaptive weight adjustment mechanism based on EMA. When the angle sampling value jumps by more than a preset threshold τ (such as 10°) compared to the historical estimate, the current frame weight coefficient is automatically adjusted lower to achieve robust suppression of outliers, and the weight formula is as follows:

[0102]

[0103] The triggered image is converted to RGB format through the callback interface of the industrial camera and saved in the specified path in the format of "group number-group number.jpg", and at the same time, the detection model is sent for detection. Every 16 frames of image is a group, and the image naming is automatically rolled over, which is convenient for subsequent tracking and defect positioning analysis.

[0104] In summary, through the acquisition trigger mechanism based on fixed angle multiples, combined with the robust angle filtering algorithm, the present application effectively avoids the shooting misalignment problem caused by error accumulation, significantly improves the consistency and positioning accuracy of image acquisition, and provides a stable and reliable image basis for the semantic segmentation detection model of gear defects.

[0105] 2. Defect detection module

[0106] As shown in Figure 3 , the defect detection module includes a preprocessing unit, an effective tooth surface segmentation network, and a defect detection network.

[0107] Image preprocessing unit; image preprocessing is a key step in the defect detection process, which improves the recognizability of defects by improving image quality and optimizing visual effects. First, the collected tooth surface image to be detected is subjected to Gaussian filtering processing. Gaussian filtering is a smoothing processing of an image through convolution operation to remove noise and details. It uses a Gaussian function to weight each pixel point in the image, and the formula is as follows:

[0108]

[0109] Where I′(x,y) is the pixel value in the original image. G(i,j) is the weight value in the Gaussian kernel, which is usually calculated based on the Gaussian function. k is the radius of the Gaussian kernel, which is usually determined by the standard deviation of the Gaussian function.

[0110] On this basis, the image is further contrast-enhanced to optimize the brightness distribution, making the tiny defects on the tooth surface more prominent and clear. This preprocessing step not only improves the overall image quality but also enhances the algorithm's adaptability to different lighting conditions, laying a solid foundation for accurate defect segmentation and identification.

[0111] Effective tooth surface segmentation network; a comparison revealed that ICNet exhibits strong timeliness and accuracy, making it suitable for effective tooth surface segmentation tasks. The network's design leverages the advantages of multi-scale feature extraction and cascaded feature fusion to ensure semantic accuracy and spatial consistency in the segmentation results. The neck network hierarchically fuses features of varying resolutions, effectively combining global information with local details, while the decoding network further enhances the classification accuracy of the segmentation results. The overall network maintains high efficiency while also balancing segmentation performance, making it ideally suited for high-resolution semantic segmentation tasks in real-time scenarios.

[0112] Defect detection network; After obtaining the results of the effective tooth surface segmentation network, the most important thing is to detect defects such as pitting. DDRNet was selected as the basic network, and a series of improvements were made on its basis. Its network structure is as follows Figure 3 shown in the lower half of the .

[0113] The defect detection network includes a first EMAM network, a second EMAM network, a context branch network, a space branch network, a dot product unit, a first fusion unit, a second fusion unit and a segmentation head;

[0114] The first EMAM network is used to capture the features of the effective tooth surface image for the first time through multi-scale convolution and attention mechanism to obtain first feature maps of different scales; the second EMAM network is used to capture the features of the first feature maps for the second time through multi-scale convolution and attention mechanism to obtain second feature maps of different scales; the context branch network is used to extract global information of the second feature maps through multi-layer residual basic blocks and residual bottleneck blocks to obtain global features; the spatial branch network is used to extract local information of the second feature maps through multi-layer residual basic blocks and residual bottleneck blocks to obtain local features; the dot product unit is used to fuse the global features and the local features through pixel point multiplication to obtain first fusion features; the first fusion unit is used to fuse the first fusion features and the second feature maps through jump connection, and the second fusion features are obtained through transposed convolution upsampling; the second fusion unit is used to fuse the second fusion features and the first feature maps through jump connection, and the third fusion features are obtained through transposed convolution upsampling; and the segmentation head is used to generate the category labels corresponding to each pixel in the effective tooth surface image based on the third fusion features.

[0115] That is, the defect detection network starts processing from an effective tooth surface image, and first extracts features through an initial convolution layer. The role of the convolution layer is to process the 3 channels of the image through convolution operation and convert them into a feature map of 16 channels. In this stage, the spatial resolution of the image remains unchanged and is still the resolution of the original image. The EMAM module captures features of different scales through multi-scale convolution and attention mechanism, and the residual basic block is the basic building block of the residual network, which contains two 3x3 convolution layers, each of which is followed by batch normalization. The input is added to the output of the second convolution layer through a shortcut and then applied to the ReLU activation. This structure alleviates the gradient vanishing problem of deep networks by allowing gradients to flow directly through the shortcut, thereby facilitating training. As shown in FIG. 2, the EMAM module contains a 3x3 convolution layer and an ECA (Efficient Channel Attention) module, and the ECA module can adaptively adjust the importance of different channels. The EMAM module also combines three scalable convolutions with different expansion rates to extract features of different scales, and adds residual links and spatial attention calculation to enhance the effect of feature extraction. Figure 4

[0116] ​Specifically, first, the image passes through EMAM 1. The input 512x512x3 image is processed by 3x3 convolution with 16 output channels, a stride of 1, padding of 1, and the size is kept unchanged. Then, batch normalization and ReLU activation are applied. Next, an Efficient Channel Attention (ECA) module is applied to enhance channel attention using adaptive average pooling (output 1x1), 1x1 convolution, and Sigmoid activation. After that, three consecutive 5x5 dilated convolutions with dilation rates of 1, 2, and 3, respectively, are applied using group convolution (groups = 16). Each convolution is followed by batch normalization and ReLU, keeping the number of channels at 16. A channel mixer refines the features using 1x1 convolution and batch normalization, and a final down-sampling convolution (stride 2, padding 1) reduces the spatial dimension to 256x256 with 16 channels, resulting in an output of 256x256x16. The process is similar for EMAM 2, but the output channel increases from 16 to 32. After down-sampling, the output is 128x128x32. Subsequent residual building blocks further process the features. Residual building blocks 1 and 2 maintain 128x128x32, residual building block 3 reduces the size to 64x64 with 64 channels through a stride 2 convolution, and residual building block 4 maintains 64x64x64. The final output of the initial layer is 64x64x64.

[0117] It should be noted that after further processing by the second EMAM module, the number of channels of the feature map is expanded from 16 to 32. Starting from the original image, multiple down-sampling operations are involved in the backbone network, gradually reducing the resolution of the feature map. First, after two basic convolution blocks, the number of channels of the feature map is gradually expanded from 32 to 64 and 128, and the resolution is reduced to 1 / 2, 1 / 4, and 1 / 8 of the original image each time the down-sampling is performed. Down-sampling is achieved through 3x3 convolution and a stride of 2. After each down-sampling, the network can capture more macro context information.

[0118] The core role of these down-sampling processes is to gradually increase the number of channels while reducing the spatial resolution, so that the model can pay more attention to global context features at a lower resolution. The gradual compression of the feature map also helps the model optimize computing resources, enabling it to process more abstract feature information in subsequent stages.

[0119] It should be noted that the context branch network is mainly responsible for capturing global information of the image. Through multiple layers of convolution and down-sampling operations, the network can gradually extract features at different levels and construct high-level context information. This context information is crucial for understanding the global structure and spatial relationships of the image. In the Figure 5In the DDRNet, the design of the context branch is carefully constructed to ensure that useful features can be extracted at multiple scales. The input to the context branch network is a 64x64x64 feature map, which is progressively downsampled through multiple layers of residual basic blocks and residual bottleneck blocks to capture more extensive scene information. The residual bottleneck block is a variant of the residual basic block, which reduces dimensionality through a 1x1 convolution, processes through a 3x3 convolution, and increases dimensionality through a 1x1 convolution, combined with a shortcut connection to improve efficiency. The first layer includes two residual basic blocks, which reduce dimensionality to 32x32x128 through a 3x3 convolution with a stride of 2, and the second layer maintains the size. The second layer further reduces dimensionality to 16x16x256, and the third layer uses a residual bottleneck block to reduce to 8x8x256. The final output is 8x8x256, rich in global context information.

[0120] The spatial branch focuses on preserving local spatial details at high resolution, with an input of 64x64x64, aiming to retain precise information at the pixel level. The spatial branch processes these feature maps through a series of basic residual blocks and residual bottleneck block modules, but its core feature is to maintain the spatial dimension unchanged, i.e., always maintaining a resolution of 64x64. Specifically, the first and second layers each contain two basic residual blocks, each consisting of two 3x3 convolution layers with a stride of 1 and a padding of 1, so the input and output sizes remain unchanged (64x64x64). The basic residual block is the basic building block of the residual network, where the input is added to the output of the second layer convolution through a shortcut connection, and then activated by ReLU. This design not only helps to alleviate the problem of gradient vanishing, but also ensures the effective propagation of features. The third layer uses a residual bottleneck block module, also maintaining the spatial dimension unchanged (64x64x64). The residual bottleneck block module further refines spatial features through a 1x1 convolution, a 3x3 convolution, and a 1x1 convolution, combined with a shortcut connection. Finally, the output of the spatial branch is a 64x64x64 feature map, rich in local spatial details. These details are crucial in the semantic segmentation task, especially for scenarios that require accurate boundaries and small target segmentation.

[0121] In the intermediate feature extraction process, DDRNet also uses a bilateral fusion operation to strengthen the information extracted from the context branch and the spatial branch. Bilateral fusion fuses the information of the context branch and the spatial branch, ensuring the effective combination of global context and local details. The process includes compression and downsampling, such as compressing the 64x64x64 features of the spatial branch through 1x1 convolution and using 3x3 convolution (step 2) to reduce dimension to match the resolution of the context branch (such as 32x32 or lower). When the context branch is fused into the spatial branch, the low-resolution features of the context branch need to be upsampled to the high-resolution of the spatial branch. This upsampling process uses bilinear interpolation to achieve by smoothly increasing the size of the feature map. Fusion combines compressed spatial features and context features through addition or concatenation, performed multiple times to enhance feature interaction. That is, the features in the context and spatial branches of the present invention are fused at this stage to form a more comprehensive feature representation. Through this multi-scale, multi-level feature fusion, the network can more accurately capture different information in the image and effectively reduce errors.

[0122] Meanwhile, the DAPPM module plays a crucial role. Its goal is to enhance the model's ability to process multi-scale information, especially combining global context information and local detail information. In deep learning models, tasks like semantic segmentation tasks often need to pay attention to both the general location of objects and the information of detailed parts. The DAPPM module further enhances the model's understanding and processing of different scale information through pyramid pooling. DAPPM is applied to low-resolution context features (8x8x256), capturing multi-scale information through pooling operations with different kernel sizes (3, 5, 9, 17) and strides (2, 4, 8). These multi-scale features are then combined, outputting 8x8x128, enhancing multi-scale context representation. That is, while the context branch extracts global features, the DAPPM module provides multi-scale information enhancement for these features, making the final feature representation more rich and diverse.

[0123] The core task of the decoding head is to map the feature maps extracted from the backbone network back to the original image resolution and output the final segmentation result. The decoding head contains multiple upsampling and feature fusion modules, which gradually restore the spatial resolution of the image and generate classification labels for each pixel.

[0124] It is to be noted that the decoding head upsamples the fused features back to the original resolution to generate the segmentation map. The segmentation head refines the features through convolution, providing additional supervision during training, and the SubPixel Convolution restores the features to 512x512 through three times of factor-2 upsampling. After each upsampling layer, the network fuses the features from different branches in the backbone network. First, the features from the spatial branch are merged with the upsampling results in the decoding head, and then further processed through convolution. This feature fusion operation ensures that the decoding head can integrate the feature information from different branches to achieve more accurate semantic segmentation. After convolution processing of these fused feature maps, a softmax layer is finally used to output the class labels of each pixel.

[0125] The improved DDRNet shows high performance in handling complex image segmentation tasks through its innovative multi-branch structure, multi-scale pooling and attention mechanism of the DAPPM module, and fine upsampling and feature fusion of the decoding head. By gradually extracting global and local features, layer-by-layer compression and recovery of spatial resolution, and combining the attention mechanism, the improved DDRNet can effectively capture and fuse important information in the image, provide accurate semantic segmentation results, and especially perform excellently in detail processing and complex scene segmentation tasks.

[0126] Further, in order to improve the discrimination ability of the deep learning segmentation model in the gear tooth surface defect detection system to edge details, small defects and complex backgrounds, the present application proposes an uncertainty-aware multi-branch cooperative online hard example mining loss function (Uncertainty-aware Cooperative Online Hard Example Mining Loss, referred to as UCOhemLoss). Based on the traditional cross-entropy and online hard example mining mechanism, the loss function introduces an uncertainty weight factor and a multi-branch consistency regularization term, which has stronger discrimination ability, stability and generalization performance. In addition, in order to enhance the cooperativity between the multiple branches of the model, the present application also introduces an output consistency regularization term, which encourages the consistency of the outputs of the two branches, thereby improving the overall stability and discrimination accuracy of the network.

[0127] Specifically, the loss function L t is as follows:

[0128] L t = L c + λL s + βL consistency

[0129] wherein λ is a spatial branch loss weight, preferably set to 0.4, β is a consistency regularization term weight, preferably set to 0.2, and Lconsistency is the consistency loss term, L c is the context branch loss term, L s is the spatial branch loss term;

[0130] Among them, L consistency The calculation method is:

[0131]

[0132] Among them, N is the total number of pixels, p c,i is the probability that the context branch network predicts the i-th pixel as the true category, p s,i is the probability that the spatial branch network predicts the i-th pixel as the true category;

[0133] Spatial branch loss term L s The calculation method is:

[0134]

[0135] in, For the spatial branch network, the i s pixels are predicted as the true class The probability of For the i s The uncertainty weight of the spatial branch loss term corresponding to pixels, K s is a set value, indicating the top K with the largest spatial branch loss s pixels;

[0136] Context branch loss term L c The calculation method is:

[0137]

[0138] in, The context branch network will be i c pixels are predicted as the true class The probability of For the i c The uncertainty weight of the context branch loss term corresponding to pixels, K c is a set value, indicating the top K with the largest context branch loss c pixels;

[0139] It should be noted that, in order to solve the problem that traditional online hard example mining only selects pixel samples based on the loss size and cannot effectively focus on the prediction of uncertain areas, this paper proposes to introduce an uncertainty weight mechanism based on the softmax output entropy value. and The method for determining is:

[0140] w i =1+εH i

[0141]

[0142] wherein, ε∈[0,1] is an uncertainty weight adjustment coefficient, the application preferably sets to 0.5, so that the pixels with high prediction uncertainty have higher gradient contribution in back propagation, w i is or H i is the prediction entropy corresponding to the i-th pixel, C is the total number of defect categories, is the probability of the i-th pixel being predicted as the c-th defect category.

[0143] It should be noted that the uncertainty-aware multi-branch collaborative OHEM loss function proposed in the application effectively improves the perception ability of the model to the edge blur area and the tiny defect under the premise of maintaining the training stability and the inference efficiency, and enhances the robustness to the complex background and the easily confused samples. By introducing the uncertainty weight mechanism and the multi-branch consistency regularization term, fine feature learning and branch collaborative optimization are realized, and the deficiencies of the existing OHEM method in detail capture and output consistency are significantly improved. At the same time, the loss function has good structure compatibility and can be widely applied to multi-branch image segmentation networks, and has strong engineering adaptability and popularization value.

[0144] It should be noted that when obtaining the samples of the training defect detection module, the data set can be expanded first; data set expansion is an important link to improve the performance of the deep learning model. In view of the problem that only 300 images of tooth surface with defects such as pitting are available under actual conditions, the StyleGAN3 generative adversarial network is used for data expansion in the present example.

[0145] StyleGAN3 generates high-quality synthetic images similar to real images through the cooperation of the generator and the discriminator. The generator uses periodic consistency design and multi-layer style mapping network to accurately control the style of the generated image, while fusing multi-scale features to enhance local and global consistency. The discriminator compares the generated image with the real image to optimize the performance of the generator.

[0146] Based on 300 real images of tooth surface defects, the present example successfully expands about 1200 synthetic images, and the generated images retain the details and distribution rules of the defects, greatly enhancing the diversity of the data set. This data expansion provides rich training data for the deep learning model, significantly improving the precision and robustness of the model in defect detection such as pitting.

[0147] 3. Information output module

[0148] Quantitative indicators; after the segmentation model successfully identifies and segments the tooth surface defect area, the present application uses the Connected Component Analysis algorithm to quantitatively analyze the defect area. This algorithm processes binary images and labels connected pixel regions to accurately calculate the number, area and loss ratio of defects. Specifically, the algorithm traverses all pixels through depth-first search (DFS) or breadth-first search (BFS), labels each connected region, and calculates the pixel area to obtain the actual size of each defect.

[0149] User interface; to simplify hardware device invocation and model deployment, the present application develops a graphical user interface based on the Unity engine. The camera acquisition program is written in C++ language and integrated into the Unity platform after compiling to generate a dynamic link library. In the user interface, interactive buttons are set to control image acquisition and processing:

[0150] Real-time acquisition; by starting the button, the camera acquisition program is called to acquire real-time tooth surface images and store them in the specified path, while displaying the acquired images in the interface.

[0151] Real-time detection; the system processes the real-time acquired tooth surface images by calling the deployed segmentation model and generates the corresponding detection mask. The detection result is displayed on the image in an overlay manner, which is convenient for users to observe and compare in real time.

[0152] Three-dimensional model; the system also integrates the quantitative information display function of the defect detection result. Each time a new defect is detected, the system will intuitively display its number, area and loss ratio in the interface. By mapping the detection results to the three-dimensional gear model, users can freely rotate and zoom in on the three-dimensional model in the Unity scene to intuitively view the defect location and characteristics of each tooth surface. The three-dimensional model is created by Blender and presents the defect detection results through mapping.

[0153] Detailed information sub-interface; the interface has a data analysis submodule that lists the detection results of each tooth surface, including defect number, area and loss ratio. This function aims to help users quickly grasp the gear wear condition and provide accurate basis for subsequent maintenance and repair.

[0154] The system realizes accurate quantification of defect characteristics through connected component analysis, and combines the visualization interface of Unity with the three-dimensional gear model display, enabling users to efficiently obtain detection results and intuitively understand the tooth surface wear condition. This design simplifies the hardware invocation and model deployment process, significantly improves the real-time performance and ease of use of defect detection, and provides comprehensive technical support for gear maintenance.

[0155] Further, a machine vision-based gear tooth surface defect online detection method comprises the following steps:

[0156] Collecting a tooth surface image to be measured; performing noise removal processing and contrast enhancement processing on the tooth surface image to be measured to obtain a pretreated tooth surface image; extracting the position of the gear in the pretreated tooth surface image and segmenting the gear according to the position to obtain an effective tooth surface image; extracting global features and local features of the effective tooth surface image, and generating a class label corresponding to each pixel in the effective tooth surface image according to the fused global features and local features; and obtaining the area of each type of defect according to the class label corresponding to each pixel in the effective tooth surface image.

[0157] In summary, the present application provides a machine vision-based gear tooth surface defect online detection system, which has the following advantages compared with the prior art:

[0158] 1. The image acquisition system designed in the present application integrates a rotary encoder, a data acquisition card, a low-speed or medium-speed industrial camera, and a light source, etc. devices, which can realize the accurate positioning and image capture function of one or several tooth surfaces to be measured during the high-speed rotation of the gear, i.e. realize the in-situ acquisition, automatic focusing and accurate angle adjustment of the gear tooth surface under different rotations. Through this system, the defect size measurement error caused by the difference in the field of view is effectively overcome, the consistency of the image quality is ensured, and the automation level and detection accuracy of the system are significantly improved.

[0159] 2. Compared with the traditional high-speed camera acquisition method, the system of the present application reduces the demand for computing resources and storage space. By reducing image redundancy and optimizing data acquisition, the system effectively reduces the cost and improves the processing efficiency while ensuring the detection accuracy.

[0160] 3. In view of the problem of limited number of real images of gear pitting and other defects, StyleGAN3 generative adversarial network is used to expand the data set, and the generated images are highly similar to real images and rich in details, so that the deep learning model can better learn the gear surface defect features, break through the data bottleneck, enhance the accuracy and robustness of detection, and effectively solve the hindrance caused by the shortage of data set to model training.

[0161] 4. A two-stage semantic segmentation model is proposed, the first stage uses ICNet to quickly segment the effective tooth surface area, and the second stage uses an improved DDRNet to accurately detect the tooth surface defects. Through this design, when processing high-resolution images, both detection accuracy and efficiency are taken into account, the model inference speed and segmentation accuracy are significantly improved, and the problem that a single model cannot simultaneously satisfy high precision and high efficiency detection is effectively solved.

[0162] 5. The improvement of DDRNet, on the basis of the original network, the EMAM module designed containing attention, scalable convolution, multi-scale convolution, etc. is adopted to strengthen the feature extraction of high-resolution feature map. In addition, the feature fusion of high-resolution feature map in the encoding process and decoding process is increased, which realizes more accurate segmentation. The ability to capture defect edge features is enhanced, especially the precise positioning of the edge of a small defect is significantly improved, effectively solving the problem of fuzzy defect edge segmentation and poor detection of small defects of DDRNet.

[0163] 6. The OpenCV connected component analysis algorithm is introduced to process the mask image, which can automatically identify the number of defects, area and damage ratio, etc. information, and calculate the pixel to real object scale according to the physical size of the image acquisition system, realize the accurate quantification of the actual size of the defect, effectively solve the problem of difficulty in analyzing the mask image output by the deep learning model and inaccuracy in quantifying the size of the defect.

[0164] 7. A graphical user interface based on Unity is developed, which integrates image acquisition, two-stage defect detection model and quantitative detection result display functions, realizes the visualization of the detection process, simplifies the hardware device calling and model deployment process, and the user can complete the whole detection process through simple operation, improves the operation efficiency and use convenience of the system, effectively solves the problems of complex operation and inconvenient management of the detection system.

[0165] Of course, the present application can have other various embodiments, and those skilled in the art can certainly make various corresponding changes and modifications according to the present application without departing from the spirit and essence of the present application. However, these corresponding changes and modifications should all belong to the protection scope of the claims attached to the present application.

Claims

1. A gear tooth surface defect online detection system based on machine vision, characterized in that: It includes an image acquisition module, a defect detection module, and an information output module; wherein the defect detection module includes a preprocessing unit, an effective tooth surface segmentation network, and a defect detection network; The image acquisition module is used to acquire the image of the tooth surface to be measured; The preprocessing module is used to perform noise removal and contrast enhancement processing on the tooth surface image to be measured to obtain a preprocessed tooth surface image; The effective tooth surface segmentation network is used to extract the position of the gear in the preprocessed tooth surface image, and segment the gear according to the position to obtain the effective tooth surface image; The defect detection network is used to extract global features and local features of the effective tooth surface image, and generate a category label corresponding to each pixel in the effective tooth surface image based on the fused global features and local features; The information output module is used to obtain the area of ​​each category of defects based on the category label corresponding to each pixel in the effective tooth surface image.

2. The machine vision-based online detection system for gear tooth surface defects according to claim 1, characterized in that: The image acquisition module includes a rotary encoder, a data acquisition card, an industrial camera, a host computer and a dedicated light source; The rotary encoder is installed on the gear spindle to measure the gear rotation angle in real time and send the rotation angle to the host computer through the data acquisition card; The industrial camera is installed in a position facing the gear in the gear wear testing machine through a rigid support structure; The dedicated light source is used to provide lighting for the gear wear testing machine; The host computer is used to trigger the industrial camera to collect the image of the tooth surface to be measured when the gear rotates to a set angle.

3. The machine vision-based online detection system for gear tooth surface defects according to claim 2, characterized in that: The method for the host computer to determine whether the gear has rotated to the set angle is: The host computer reads the angle value θ output by the rotary encoder at the current moment in real time during the continuous rotation of the gear. k , and the angle value θ is filtered by an exponential weighted average filter k Perform smoothing to obtain the smoothed angle value at the current moment in, is the smoothed angle value of the previous moment, α k The smoothing coefficient for adjusting the response sensitivity at the current moment; Determine the angle value after smoothing at the current moment Check if the following trigger conditions are met. If yes, it means the gear has rotated to the set angle. If no, it means the gear has not rotated to the set angle: Where Δθ is the theoretical angular interval corresponding to a single gear tooth, and ∈ is the allowable error tolerance.

4. The machine vision-based online detection system for gear tooth surface defects according to claim 3, characterized in that: The smoothing coefficient α at the current moment k The setting method is: Among them, τ is the set threshold.

5. The on-line detection system for gear tooth surface defects based on machine vision according to claim 1, characterized in that: The effective tooth surface segmentation network uses a three-level resolution cascade network of the ICNet model to extract the position of the gear in the preprocessed tooth surface image, and segments the gear according to the position to obtain an effective tooth surface image.

6. The online detection system for gear tooth surface defects based on machine vision according to claim 1, characterized in that: The defect detection network includes a first EMAM network, a second EMAM network, a context branch network, a space branch network, a dot product unit, a first fusion unit, a second fusion unit and a segmentation head; The first EMAM network is used to capture the features of the effective tooth surface image for the first time through multi-scale convolution and attention mechanism to obtain first feature maps of different scales; The second EMAM network is used to capture the features of the first feature map for a second time through multi-scale convolution and attention mechanism to obtain second feature maps of different scales; The context branch network is used to extract global information of the second feature map through multiple layers of residual basic blocks and residual bottleneck blocks to obtain global features; The spatial branch network is used to extract local information of the second feature map through multiple layers of residual basic blocks and residual bottleneck blocks to obtain local features; The dot product unit is used to fuse the global feature and the local feature by multiplying the corresponding pixels to obtain a first fused feature; The first fusion unit is used to fuse the first fusion feature and the second feature map through a jump connection, and obtain the second fusion feature through transposed convolution upsampling; The second fusion unit is used to fuse the second fusion feature and the first feature map through a jump connection, and obtain a third fusion feature through transposed convolution upsampling; The segmentation head is used to generate a category label corresponding to each pixel in the effective tooth surface image based on the third fusion feature.

7. The machine vision-based online detection system for gear tooth surface defects according to claim 6, characterized in that: The method for any EMAM network to capture any feature map through multi-scale convolution and attention mechanism is: The feature image passes through 3x3 convolution, batch normalization, and ReLU activation in sequence, and then enters the Efficient Channel Attention module. The Efficient Channel Attention module uses adaptive average pooling, 1x1 convolution, and Sigmoid activation to enhance the channel attention of the feature image and obtain the initial enhanced feature map. Three dilated convolutions with different dilation rates are used to operate on the initial enhanced feature map, where each dilated convolution is followed by batch normalization and ReLU activation; The 1x1 convolution is used to fuse the initial enhanced feature map and the output of the three dilated convolutions. The fusion result is then subjected to spatial dimensionality reduction through downsampling convolution to obtain the first branch dimensionality reduction feature map. The initial enhanced feature map is operated by cascaded 3x3 convolution and 1x1 convolution to obtain the second branch dimensionality reduction feature map; The product of the first branch dimensionality reduction feature map and the second branch dimensionality reduction feature map is used as the feature map of different scales finally output by the EMAM network.

8. The machine vision-based online detection system for gear tooth surface defects according to claim 6, characterized in that: The following branch network includes a first residual basic block RB, a second residual basic block RB, a third residual basic block RB, a first residual bottleneck block RBB and a DAPPM network that are sequentially cascaded; The spatial branch network includes a first residual basic block RB, a fourth residual basic block RB, a fifth residual basic block RB and a second residual bottleneck block RBB that are sequentially cascaded; Among them, the following branch network and the spatial branch network share the first residual basic block RB, and the input of the third residual basic block RB in the following branch network is the output of the second residual basic block RB and the output of the fourth residual basic block RB in the spatial branch network. The input of the first residual bottleneck block RBB is the output of the third residual basic block RB and the output of the fifth residual basic block RB in the spatial branch network; the output of the DAPPM network is the final global feature; The input of the fifth residual basic block RB in the spatial branch network is the output of the fourth residual basic block RB and the output of the second residual basic block RB in the following branch network. The input of the second residual bottleneck block RBB is the output of the fifth residual basic block RB and the output of the third residual basic block RB; the output of the second residual bottleneck block RBB is the final local feature.

9. The machine vision-based online detection system for gear tooth surface defects according to claim 1, characterized in that: The loss function L used to train the defect detection module t as follows: L t =L c +λL s +βL consistency Among them, λ is the spatial branch loss weight, β is the consistency regularization weight, L consistency is the consistency loss term, L c is the context branch loss term, L s is the spatial branch loss term; Among them, L consistency The calculation method is: Among them, N is the total number of pixels, p c,i is the probability that the context branch network predicts the i-th pixel as the true category, p s,i is the probability that the spatial branch network predicts the i-th pixel as the true category; Spatial branch loss term L s The calculation method is: in, For the spatial branch network, the i s pixels are predicted as the true class The probability of For the i s The uncertainty weight of the spatial branch loss term corresponding to pixels, K s is a set value, indicating the top K with the largest spatial branch loss s pixels; Context branch loss term L c The calculation method is: in, The context branch network will be i c pixels are predicted as the true class The probability of For the i c The uncertainty weight of the context branch loss term corresponding to pixels, K c is a set value, indicating the top K with the largest context branch loss c pixels; Uncertainty weight and The method to determine is: In i =1+εH i Among them, ε is the uncertainty weight adjustment coefficient, w i for or H i is the predicted entropy corresponding to the i-th pixel, C is the total number of defect categories, is the probability that the i-th pixel is predicted to be the c-th defect category.

10. A method for online detection of gear tooth surface defects based on machine vision, characterized in that: The following steps are involved: Collecting images of the tooth surface to be measured; Perform noise removal and contrast enhancement processing on the tooth surface image to be measured to obtain a pre-processed tooth surface image; Extracting the position of the gear in the preprocessed tooth surface image, and segmenting the gear according to the position to obtain a valid tooth surface image; Extract the global features and local features of the effective tooth surface image, and generate the category label corresponding to each pixel in the effective tooth surface image based on the fused global features and local features; The area of ​​each category of defects is obtained according to the category label corresponding to each pixel in the effective tooth surface image.

Citation Information

Patent Citations

  • SF-YOLONet metal gear end face defect detection method and system

    CN117455836A

  • Industrial product appearance defect semantic segmentation method and system

    CN119229106A