Hub defect detection method based on 2.5 D photometric stereo image

Through the combination of 2.5D photometric stereo camera and deep learning, the problem of manual detection and labor-consuming and missed detection in wheel hub defect detection is solved, and efficient and accurate defect identification and positioning is achieved.

CN120259189APending Publication Date: 2025-07-04HANGZHOU YINGSHIMAI INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510256301.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the detection of wheel hub defects relies on manual detection to be time-consuming and labor-intensive, and there are problems of mis-checking and missed detection. It is difficult to accurately identify defects on the surface of painted wheel hubs with high reflectance and complex shapes.

Method used

Images are collected using a 2.5D photometric stereo camera, combined with exposure fusion algorithm and multi-scale fusion technology for image preprocessing, unsupervised and supervised deep learning models are used to identify defects, eliminate reflection effects through exposure fusion algorithm, use multi-channel data to display defects, and combine CSPDarkNet53 network layer and FPN+PAN module for feature extraction and fusion.

Benefits of technology

It realizes accurate identification and positioning of defects in high-reflective coating wheel hubs, reduces the dependence of manual detection, improves detection efficiency and accuracy, and reduces false detection and missed detection rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259189A_ABST
    Figure CN120259189A_ABST
Patent Text Reader

Abstract

The invention discloses a hub defect detection method based on a 2.5 D photometric stereo image. The method comprises the following steps: sending an acquired preprocessed hub image into a pre-trained deep learning model for identifying and positioning coating hub defects so as to output a detection result; wherein the preprocessing of the hub image comprises the following steps: acquiring images of a coated hub under various exposures through a camera, and fusing the images under the various exposures by adopting an exposure fusion algorithm to obtain the preprocessed hub image. The 2.5 D photometric stereo image technology is utilized, and supervised learning and unsupervised learning methods are combined, so that the detection of the defects of the high-reflective coating hubs with different colors is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition and detection in deep learning, and particularly to a method for detecting wheel hub defects based on 2.5D photometric stereo images. Background Art

[0002] Among the numerous components of an automobile, the wheel hub is particularly crucial. However, in the actual production process, due to the influence of various factors, the manufactured wheel hubs may have various defects. These defects may cause safety problems during daily driving. Therefore, it is particularly urgent to be able to efficiently and accurately identify wheel hub defects. Currently, wheel hub manufacturers mainly rely on manual inspection for defect detection, but this method is time-consuming and labor-intensive. In addition, long-term high-intensity work is likely to cause physical and mental fatigue of employees, resulting in misdetection and missed detection. Manual inspection is also subjective, which makes the stability and reliability of the detection results questioned.

[0003] With the development of machine vision technology, the use of image recognition technology in deep learning to automatically identify and detect defects in the collected images has been widely applied and achieved good results. For example, the patent application No. 202410189363.8, a method, device, electronic device and storage medium for detecting wheel hub surface defects by integrating an attention mechanism and deformable convolution, discloses collecting image data of a wheel hub to be detected according to position information; obtaining an original data set of the wheel hub to be detected according to the image data; processing the original data set to obtain an enhanced data set; and inputting the enhanced data set into a pre-trained neural network to obtain the detection result of the wheel hub to be detected.

[0004] The above patent identifies the input wheel hub image through a pre-trained neural network to output the corresponding defect detection result. However, the accuracy of its recognition result depends on the clarity of the image. Due to the rich variety of coating colors, ranging from classic silver-white to modern black, there are various choices for the color of the wheel hub. The surface of the painted wheel hub usually has strong reflective characteristics, which poses a challenge to shooting with an optical camera. The presence of a highly reflective surface makes it easy to have a reflection phenomenon during shooting, which not only reduces the quality of the picture but also may cause the details and textures of the wheel hub to not be clearly captured, thus affecting the accurate identification and positioning of defects on the painted wheel hub. In addition, the shape of the wheel hub surface is complex and variable, and the diversity of the curved surface further increases the difficulty of shooting and defect recognition. Moreover, factors such as the clarity and temperature of the captured image will affect the detection output result of the pre-trained neural network model. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for detecting wheel hub defects based on 2.5D photometric stereo images, which improves the clarity of the image by preprocessing the collected image, thereby enhancing the accuracy of defect detection and recognition.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: A wheel hub defect detection method based on 2.5D photometric stereo images. The preprocessed wheel hub images collected are sent into a pre-trained deep learning model for the identification and positioning of defects on the painted wheel hubs to output detection results. The preprocessing of the wheel hub images includes: collecting images of the painted wheel hubs under multiple exposures through a camera, and using an exposure fusion algorithm to fuse the images under multiple exposures to obtain the preprocessed wheel hub images.

[0007] The wheel hub images under multiple exposures include images under overexposure, normal exposure, and underexposure conditions. The exposure fusion algorithm performs fusion calculations based on the contrast, saturation, and exposure metrics of each pixel in the image to obtain the finally fused image as the preprocessed wheel hub image.

[0008] The fusion process of the exposure fusion algorithm for the wheel hub images includes:

[0009] Calculating the contrast metric, saturation metric, and exposure metric of each pixel;

[0010] Calculating the final weight of each pixel based on the contrast metric, saturation metric, and exposure metric of each pixel;

[0011] The weight map refers to the calculated weight values. The combination of the final weight values of each pixel forms the weight map.

[0012] The preprocessing further includes: using a multi-scale fusion method to process the weight map image. First, use the Laplacian pyramid to decompose the input original image to obtain image representations at different scales; then perform Gaussian pyramid processing on the calculated weight map to obtain weight representations corresponding to the scales of the Laplacian pyramid; then multiply the Laplacian pyramid and the Gaussian pyramid at the corresponding scales to obtain the fused Laplacian pyramid; finally, through the inverse Laplacian process, represent the fused Laplacian pyramid as a fused image, and use the fused image as the input of the pre-trained deep learning model.

[0013] The pre-trained deep learning model includes an unsupervised learning model and a supervised learning model;

[0014] Among them, the fused image is first sent into the unsupervised learning model for defect identification. If a defect is identified at this time, the detected defect result is output; otherwise, the fused image is sent into the supervised learning model again for re-identification, and the corresponding defect detection result is output based on the supervised learning model.

[0015] The images of the painted wheels are collected by a 2.5D photometric stereo camera. After the images of each channel are fused using the exposure fusion algorithm, the images of each channel of the 2.5D photometric stereo camera are divided into sub-regions of m*n, where the size of m*n is the same as the number of channels corresponding to the 2.5D photometric stereo camera; sub-images at the same position in each channel are selected for stitching to obtain the complete stitched image of the channel, and the stitched complete image is used as the input to the pre-trained deep learning model.

[0016] The unsupervised learning model is trained through a pre-built training set. Among them, defect-free images are selected from the collected wheel images, and then the defect-free images are pre-processed and sent into the unsupervised learning model for training.

[0017] The supervised learning model is trained through a pre-built training set. Among them, defective images are selected from the collected wheel images and the images are labeled. After the labeling is completed, a painted wheel defect data set is formed. The supervised learning model is trained with the images in the painted wheel defect data set. The supervised learning model is improved based on the YOLO network model. The images in the painted wheel defect data set are pre-processed and then sent into the supervised learning model for training. The supervised learning model improved based on the YOLO network model includes:

[0018] Multi-channel input layer: The images of each channel of the 2.5D photometric stereo camera are divided into sub-regions of m*n, where the size of m*n is the same as the number of channels corresponding to the 2.5D photometric stereo camera; sub-images at the same position in each channel are selected for stitching, so as to be able to construct a new complete image. The stitching order is preset. For example, the first sub-image of channel 1 will be located in the first area of the new image, the first sub-image of channel 2 will be located in the second area, and so on. Such a method helps to ensure that any defect detected subsequently can be accurately traced back to its original channel image. Here, not only the pre-processed (exposure fusion) images are input, but first the images of each channel processed by the exposure fusion algorithm are segmented into multiple sub-blocks, and then the sub-blocks at the same position are separately taken out from each channel and stitched into a new image as the input of the network.

[0019] Feature Extraction and Fusion Layer: The spliced images are subjected to feature extraction using the CSPDarkNet53 network layer; then the results after feature extraction are used as the input of the Neck layer. The structure of the Neck layer consists of the CBL layer + SPP module and the FPN + PAN module. The CBL layer is used to further extract features, and the SPP module is used to extract features of different scales to enhance the model's detection ability for multi-scale targets; in the FPN + PAN module, FPN adopts a top-down structure, and the high-level feature information is passed to the low-level through upsampling and feature fusion is performed during the transfer process; PAN is a bottom-up structure, which adds a bottom-up path on the basis of FPN to transfer the low-level localization features to the high-level, thus helping the low-level features obtain high-level information; the output end of the Neck module is connected to the Prediction layer to output the image detection results.

[0020] When the detection result of the deep learning model is defective, calculate the width or area parameter of the circumscribed rectangle of the defect in the image, compare the width or area parameter with the threshold, and perform a secondary judgment on the defective image based on the threshold and output the secondary judgment result.

[0021] The advantages of the present invention are as follows: By using the 2.5D photometric stereo image technology and combining supervised learning and unsupervised learning methods, the detection of defects on high-reflective painted wheels of different colors is realized. This technology can accurately identify and locate defects, effectively solving the problems of time-consuming, laborious, and possible missed and false detections in the manual detection process. Through the multi-channel images obtained by the 2.5D photometric stereo camera, the problem of reflection on the surface of high-reflective painted wheels is successfully solved, making the defects clearly visible in the images. In addition, the application of deep learning realizes automatic defect detection, reduces the dependence on manual operations, and thus reduces labor costs. Especially in an environment of long-term and high-intensity labor, the defect detection technology based on 2.5D photometric stereo images significantly reduces the labor intensity of workers and improves work efficiency. Compared with traditional machine learning methods, deep learning can identify more types of defects in defect detection and is applicable to a wider range of scenarios. Brief Description of the Drawings

[0022] The following briefly describes the content expressed in each drawing of the present invention specification and the marks in the drawings:

[0023] Figure 1 It is a schematic diagram of the defect detection process of the painted wheel of the present invention;

[0024] Figure 2 It is a schematic diagram of the improved YOLOv4 network structure of the present invention;

[0025] Figure 3 It is a structural diagram of the FPN + PAN module of the present invention;

[0026] Figure 4 This is a schematic diagram of image segmentation for the present invention. Detailed implementation manners

[0027] The following further describes in detail the specific implementation manners of the present invention by describing the optimal embodiments with reference to the accompanying drawings.

[0028] A method for detecting wheel hub defects based on 2.5D photometric stereo images provided by this solution has at least four main improvements:

[0029] 1. Use a 2.5D photometric stereo camera to collect images. The collected images are clearer, and different types of defects can be shown through multi-channel data, thus ensuring the accuracy of subsequent detection of painted wheel hub defects.

[0030] 2. Design an image preprocessing method for the collected images (each channel image of the 2.5D photometric stereo camera). For wheel hub images with multiple exposures, use an exposure fusion algorithm to eliminate overexposed or underexposed areas in the images, improve image clarity, and thus provide clearer images for subsequent model recognition.

[0031] 3. Use two pre-trained models to separately identify defects in painted wheel hubs. The pre-trained deep learning models include: an unsupervised learning model and a supervised learning model. The preprocessed painted wheel hub images are first sent to the unsupervised learning model to identify large defects. If defects are identified, the defect results are directly output; if the unsupervised learning model does not identify large-scale defects, the preprocessed wheel hub images are further sent to the supervised learning model for further learning to identify smaller defects. If defects are identified, the defect results are directly output. By identifying in these two modes, the accuracy of defect identification is improved.

[0032] 4. For the supervised learning model, design the model structure, and improve the YOLOv4 network model to obtain the supervised learning model applied to this solution. Improve the YOLOv4 network to make it more suitable for the images collected by the 2.5D camera in this solution and more reliably meet image defect recognition.

[0033] As Figures 1-4 shown, a method for detecting wheel hub defects based on 2.5D photometric stereo images in this solution includes sending the preprocessed wheel hub images collected into a pre-trained deep learning model for identifying and locating defects in painted wheel hubs to output detection results; wherein the preprocessing of the wheel hub images includes: collecting images of painted wheel hubs with multiple exposures through a camera, and using an exposure fusion algorithm to fuse the images with multiple exposures to obtain the preprocessed wheel hub images.

[0034] The hub images under multiple exposures include images under overexposure, normal exposure, and underexposure conditions. The exposure fusion algorithm performs fusion calculations based on the contrast, saturation, and exposure metrics of each pixel in the image to obtain the finally fused image as the preprocessed hub image.

[0035] The fusion process of the exposure fusion algorithm for the hub image includes: calculating the contrast metric, saturation metric, and exposure metric of each pixel; calculating the final weight of each pixel based on the contrast metric, saturation metric, and exposure metric of each pixel; the weight map refers to the weight values finally calculated, and the weight map is obtained based on the final weight of each pixel.

[0036] The preprocessing also includes: processing the weight map image using a multi-scale fusion method. First, the input original image is decomposed using the Laplacian pyramid to obtain image representations at different scales; then, the calculated weight map is processed using the Gaussian pyramid to obtain the weight representation corresponding to the scale of the Laplacian pyramid; then, the Laplacian pyramid and the Gaussian pyramid at the corresponding scale are multiplied to obtain the fused Laplacian pyramid; finally, through the inverse Laplacian process, the fused Laplacian pyramid is represented as a fused image, and the fused image is used as the input to the pre-trained deep learning model.

[0037] The pre-trained deep learning model includes an unsupervised learning model and a supervised learning model; among them, the fused image is first fed into the unsupervised learning model for defect recognition. If a defect is recognized at this time, the detected defect result is output; otherwise, the fused image is fed into the supervised learning model again for re-recognition, and the corresponding defect detection result is output based on the supervised learning model.

[0038] The present invention proposes a method for using a 2.5D photometric stereo camera to capture the surface data of a painted hub, and combines deep learning techniques with multi-channel input to achieve precise recognition and positioning of defects on the painted hub.

[0039] For the surface of a painted hub with extremely complex structure and high reflectivity, the present invention proposes a set of data shooting solutions, adopting advanced 2.5D photometric stereo camera technology to solve the difficulties that are difficult to overcome by traditional optical cameras. The 2.5D photometric stereo camera is realized based on the photometric stereo method. Its core principle is to use light from different angles to have different effects on the object surface. By observing the brightness changes of each pixel point under different lighting conditions, the normal direction (the direction perpendicular to the object surface) of the pixel point can be calculated. Finally, using the normal information on the entire image, the three-dimensional shape of the object surface can be reconstructed. Therefore, compared with ordinary cameras, it has the following advantages:

[0040] (1) The 2.5D photometric stereo camera irradiates the product to be detected from different angles by using a multi-spectral light source or a standard light source, etc., so as to capture multiple images. By performing corresponding processing on these images, such as taking the minimum value or the median value pixel by pixel, new images can be formed, and the reflective parts in these images are effectively suppressed, thereby improving the image quality;

[0041] (2) An ordinary camera mainly captures the planar image of an object, that is, a two-dimensional image, and cannot obtain the three-dimensional information of the object. While the 2.5D photometric stereo camera can capture the three-dimensional information of the object, and according to the reconstructed three-dimensional information and normal information, the normal information can be projected along the x, y, and z directions respectively. This projection method can highlight the defects extending along different directions, so as to clearly display the defects in the image.

[0042] Based on the above analysis, the 2.5D photometric stereo camera can adopt different processing methods to obtain multi-channel data, so as to be able to more clearly display various types of defects on the surface of the hub, and the data of each channel can display different types of defects. For example: the channel data obtained by calculating the reflectivity can represent the light and dark degree of the object surface, and is suitable for detecting defects such as surface foreign objects and dirt; the data obtained by calculating the normal vector, projected in the x, y, and z directions respectively, can detect unevenness, pits, scratches, etc. extending along these directions; the channel data of selecting the median gray value of multiple images pixel by pixel can effectively eliminate reflection and can detect defects on a highly reflective surface, such as knife marks on the surface of a painted hub; by fusing the normal vector channel data calculated along the x, y, and z directions, the surface texture information can be highlighted, which is used to detect defects such as scratches, scratches, protrusions or depressions on the surface of the object to be detected; by fusing the data of the R, G, and B channels, the color information can be displayed, so as to be able to detect defects such as the presence or absence of clear coat on the surface. Therefore, compared with traditional optical cameras, it can display different types of defects through multi-channel data, thus ensuring the accuracy of subsequent detection of defects on painted hubs.

[0043] Aiming at the fact that there are many types of defects and the defects are small in painted hubs, the present invention proposes a set of defect detection solutions, and uses a deep learning-based method to realize the surface defect detection of painted hubs, which mainly consists of two parts: image preprocessing and deep learning. Among them, the image preprocessing part is used to further eliminate overexposed and underexposed areas in the image and improve the image quality; the deep learning part is based on a method combining supervised learning and unsupervised learning to complete the recognition and positioning of defects on painted hubs;

[0044] Image preprocessing part:

[0045] 1. For the same position, images of painted automotive wheels are obtained with different exposures;

[0046] 2. The exposure fusion algorithm is adopted to eliminate overexposed and underexposed areas in the image and improve the image clarity;

[0047] The exposure fusion algorithm is as follows:

[0048] Input multiple painted wheel images with different exposures, including overexposed, normal, and underexposed situations, and calculate the weights based on three indicators: contrast, saturation, and exposure. The specific calculation formula is as follows.

[0049] W(x,y) = (C(x,y)) Wc ×(S(x,y)) Ws ×(E(x,y)) We

[0050] Among them, W(x,y) represents the final weight of pixel (x,y), C(x,y) represents the contrast index of pixel (x,y), S(x,y) represents the saturation index of pixel (x,y), E(x,y) represents the exposure index, and W c 、W s and W e respectively represent the weights of the contrast, saturation, and exposure indicators.

[0051] The contrast index is used to measure the intensity of important information such as edges and textures in the image. By performing Laplacian filtering on the grayscale image and taking the absolute value as the contrast index, the calculation formula is as follows.

[0052] C(x,y) = |△gray(I(x,y))|

[0053] Among them, △gray represents the Laplacian operator of the grayscale image, and I(x,y) represents the grayscale value of pixel (x,y).

[0054] The saturation index is used to measure the vividness of colors in the image. By calculating the standard deviation between the three RGB channels as the saturation index, its calculation formula is as follows.

[0055] S(x,y) = σ(R(x,y), G(x,y), B(x,y))

[0056] Among them, σ represents the standard deviation, and R(x,y), G(x,y), and B(x,y) respectively represent the red, green, and blue channel values of pixel (x,y).

[0057] The exposure index is used to measure the exposure degree of pixels in the image. For pixels whose values are closer to 0 or 255, they are likely to be in overexposed or underexposed areas, while pixels with grayscale values around 128 are often considered to be in well-exposed areas. The exposure index can be calculated by the method of the Gaussian weight function, as follows.

[0058]

[0059] Where I(x, y) represents the pixel value corresponding to the pixel (x, y).

[0060] The final weight value of each pixel is calculated through the above formula, and these weight values form a weight map.

[0061] For the obtained weight map, a multi-scale fusion method is further used to process the input weight map image. First, the Laplacian pyramid is used to decompose the original image captured by the input camera to obtain image representations at different scales;

[0062] Then, the calculated weight map is processed by the Gaussian pyramid to obtain the weight representation corresponding to the scale of the Laplacian pyramid; then the Laplacian pyramid and the Gaussian pyramid at the corresponding scale are multiplied to obtain the fused Laplacian pyramid; finally, through the inverse Laplacian process, the fused Laplacian pyramid is represented as a fused image.

[0063] The process of constructing the Gaussian pyramid is as follows: First, the input original weight map image is used as the bottom layer image of the pyramid. Then, Gaussian filtering is performed on this bottom layer image; subsequently, the filtered image is downsampled to reduce the image size. Downsampling is usually done by removing the even rows and columns of the image, resulting in the image size being reduced to half of the original (i.e., both the height and width are halved). The above Gaussian filtering and downsampling steps are repeated multiple times, each time on the previous layer image. In each iteration, a new image layer is generated, whose size is half of the previous layer. These consecutive image layers together form the complete Gaussian pyramid.

[0064] The process of constructing the Laplacian pyramid is as follows: First, a Gaussian pyramid is constructed for the original input image (the image captured by the camera), and its top layer is selected as the starting layer of the Laplacian pyramid; then, starting from the second layer of the Gaussian pyramid, upsampling is performed to enlarge the image size; subsequently, the upsampled image is differenced from the image of the next layer of the Gaussian pyramid to generate the Laplacian image of the current layer, and this differencing operation is usually done by subtracting pixel by pixel; the above upsampling and differencing operations are repeated until all Gaussian pyramid layers are processed, finally forming the complete Laplacian pyramid.

[0065] 3. During the training phase, the pre - processed images need to be manually divided into two major categories: one category contains defective painted wheel hub defect data, and the other is defect - free painted wheel hub data. The data of the two categories are respectively fed into the supervised learning model and the unsupervised learning model in the deep learning part to train the models. When the models are deployed for use after training, the pre - processed images are directly fed into the unsupervised learning model first, and then into the supervised learning model, so as to achieve the purpose of defect recognition based on painted wheel hub images.

[0066] Deep learning part:

[0067] 1. For the data set of painted wheel hubs classified as defect - free, unsupervised learning techniques are used for training, aiming to detect potential large - area defect regions in the painted wheel hubs. At present, the vast majority of existing network models are used to process single - image inputs. This invention performs the painted wheel hub defect detection task based on multi - channel images captured by a 2.5D photometric camera, where each channel represents an image obtained after different processing steps, so the image features of each channel are different. If the images of these channels are directly used as network inputs, the expected recognition effect cannot be achieved. Therefore, the input images must be appropriately improved to ensure that each input image has similar features. The specific improvement measures are as follows:

[0068] (1) Divide each channel image of the 2.5D photometric stereo camera into sub - regions of m*n, where the size of m*n is the same as the number of channels of the 2.5D photometric stereo camera;

[0069] (2) Select sub - images at the same position in each channel for splicing, so as to be able to construct a new complete image. The splicing order is preset. For example, the first sub - image of channel 1 will be in the first area of the new image, the first sub - image of channel 2 will be in the second area, and so on. Such a method helps to ensure that any detected defect can be accurately traced back to its original channel image. The benefit of splicing is to ensure that each image input into the network has similar features, thus ensuring the detection effect.

[0070] 2. Use the spliced image as the input for unsupervised learning for training to obtain an unsupervised training model;

[0071] 3. For those images classified as having minor defects, the present invention uses a supervised learning method for training, specifically based on the Yolov4 architecture. Considering that the data output by the 2.5D photometric stereo camera is multi-channel, the present invention makes appropriate adjustments to the Yolov4 network structure. The designed network aims to simultaneously process the multi-channel input data of the 2.5D photometric stereo camera and be able to output the detection results for each channel. The detailed layout of the network structure is shown in the figure.

[0072] The specific structure is as follows:

[0073] (1) Input layer design: Each channel image of the 2.5D photometric stereo camera is divided into sub-regions of m*n, where the size of m*n is the same as the number of channels corresponding to the 2.5D photometric stereo camera; sub-images at the same position in each channel are selected for splicing, so as to be able to construct a new complete image. The splicing order is preset. For example, the first sub-image of channel 1 will be located in the first region of the new image, the first sub-image of channel 2 will be located in the second region, and so on. Such a method helps to ensure that any detected defect can be accurately traced back to its original channel image.

[0074] (2) Feature extraction: In the feature extraction stage, the CSPDarkNet53 network layer is used for feature extraction. The CSPDarkNet53 network layer introduces the CSP structure, reduces the computational redundancy of the feature map, enables the model to converge faster during training, and thus reduces the consumption of training time and computational resources.

[0075] (3) Neck layer: The feature map after the above-mentioned feature fusion processing becomes the input of the Neck layer. The Neck layer structure combines the designs of the CBL layer + SPP module and the FPN + PAN module. The CBL layer is responsible for further extracting features, and the SPP module uses pooling kernels of different sizes to perform pooling operations on the input feature map to obtain feature maps of different scales. Subsequently, the pooling results of these different scales are spliced to achieve feature fusion. The structure of the FPN + PAN module is shown in the figure, where the FPN adopts a top-down structure, passes the high-level feature information to the low-level through upsampling, and performs feature fusion during the transmission process; the PAN is a bottom-up structure, adding a bottom-up path on the basis of the FPN, and passing the low-level localization features to the high-level. The Neck layer can perform operations such as dimensionality reduction, adjustment, or fusion on the features; the final result output is the prediction layer (Prediction), which is used for the final prediction, including information such as the type of defect, the center coordinates, length, and width of the defect circumscribed rectangle.

[0076] 4. Use the dataset with labels created as the input for the improved YOLO v4 network for training to obtain the YOLOv4 training model;

[0077] 5. Output design: The network output will display the image detection results after splicing, and further reverse-derive the corresponding channel images and defect positions. The specific derivation process is as follows.

[0078] 1) During the original image segmentation process, the width and height of the sub-images are set to w and h respectively, so as to calculate the upper-left coordinates of each sub-image in the spliced image as: (x (i,j) , y (i,j) ) = (i * w, j * h), where i represents the horizontal position coordinate of the region block and j represents the vertical position coordinate of the region block; as Figure 4 shown, m and n represent the number of horizontal and vertical segments respectively.

[0079] 2) The results of the network output include the center point coordinates of the defect's rectangular box in the spliced image, denoted as (x0, y0). If (x (i,j) , y (i,j) ) ≤ (x0, y0) ≤ (x (i,j) + w, y (i,j) + h), it means that the current defect falls within the (i, j) region block, and further the channel number in the original image where the currently detected defect is located can be calculated. The calculation formula is:

[0080] index = i * m + j

[0081] where index represents the channel number corresponding to the 2.5D photometric stereo camera, i represents the horizontal position coordinate, and j represents the vertical position coordinate;

[0082] And calculate the horizontal and vertical distances of the center point coordinates of the currently detected rectangular box relative to (x (i,j) , y (i,j) ) as:

[0083] △x = x0 - x (i,j) △y = y0 - y (i,j)

[0084] 3) Given that the currently detected spliced image is the kth spliced image, the corresponding region block position in the original channels of the 2.5D photometric stereo camera can be reverse-derived as (u, v) by establishing the relationship between the two, as shown in the table:

[0085]

[0086] 4) Further calculate the upper-left coordinates of the current region block on the original channel image as

[0087] (x (u,v) , y (u,v) ) = (u * w, v * h)

[0088] 5) Furthermore, calculate that the center coordinates of the detected defect rectangle in the center of the original channel image are:

[0089] (x, y) = (x (u,v) + Δx, y (u,v) + Δy)

[0090] 6) Finally, draw the defect position on the image of the corresponding channel according to the calculated center point coordinates and the length and width of the rectangle.

[0091] 6. After the model training is completed, use the trained unsupervised model and Yolo v4 model to detect unknown painted wheels. The detection process is as follows: First, use the unsupervised model for preliminary detection. Once a defect is found, it is recognized as a real defect and a prompt is sent to the on-site employees. If no defect is detected in the detection, the Yolo v4 model will be used for secondary detection. If the Yolo v4 model also fails to detect a defect, it is considered that the current painted wheel has no defect. On the contrary, if a defect is detected, it is necessary to further evaluate according to parameters such as the width or area of the circumscribed rectangle of the defect. If these parameters exceed the preset standard threshold, the defect is determined to be real and a prompt needs to be sent to the on-site personnel; if not, the defect is considered unimportant and can be ignored. After the model detection, the category of the defect and the parameters related to the circumscribed rectangle of the defect (center point coordinates, length, width, etc.) of the rectangle can be obtained, so that the width or area parameters of the circumscribed rectangle can be obtained.

[0092] The present invention utilizes 2.5D photometric stereo image technology, combines supervised learning and unsupervised learning methods, and realizes the detection of defects on high-reflective painted wheels of different colors. This technology can accurately identify and locate defects, effectively solving the problems of time-consuming, laborious, and possible missed and misdetected in the manual detection process. Through the multi-channel images obtained by the 2.5D photometric stereo camera, the reflection problem on the surface of the high-reflective painted wheel is successfully solved, making the defects clearly visible in the image. In addition, the application of deep learning realizes automatic defect detection, reduces the dependence on manual operations, and thus reduces the labor cost. Especially in an environment of long-term and high-intensity labor, the defect detection technology based on 2.5D photometric stereo images significantly reduces the labor intensity of workers and improves work efficiency. Compared with traditional machine learning methods, deep learning can identify more types of defects in defect detection and is applicable to a wider range of scenarios.

[0093] The present invention acquires the surface data of the painted wheel hub through a 2.5D photometric stereo camera for defect detection, which can solve the problem of easy reflection on the surface of the painted wheel hub when collecting images. Compared with traditional optical cameras, it can provide multi-channel images and clearly display various defect types. By combining supervised and unsupervised learning methods to complete the defect detection of the painted wheel hub, compared with a single supervised learning method, unsupervised learning can significantly reduce the workload of defect annotation and overcome the problems of low detection accuracy and high false detection rate. The present invention redesigned the input and output layers of the unsupervised learning network and the YOLO v4 network for multi-channel input data, which can make full use of the multi-channel pictures collected by the 2.5D photometric stereo camera, thus ensuring the accuracy of defect detection. The present invention uses deep learning methods for defect detection of painted wheel hubs, which can automatically identify and locate defects, solve problems such as many false detections and missed detections in traditional manual detection, and can detect a variety of defect types, solving the problem of singularity in defect detection by traditional machine learning.

[0094] Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made by adopting the method concept and technical solution of the present invention, they are within the protection scope of the present invention.

Claims

1. A wheel hub defect detection method based on 2.5D photometric stereo images, characterized in that: The preprocessed hub images collected are sent into a pre-trained deep learning model for the identification and location of defects on the painted hubs to output the detection results; Among them, the preprocessing of the hub images includes: collecting images of the painted hubs under multiple exposures through a camera, and using an exposure fusion algorithm to fuse the images under multiple exposures to obtain the preprocessed hub images.

2. The hub defect detection method based on 2.5D photometric stereo images according to claim 1, wherein: The hub images under multiple exposures include images under overexposure, normal exposure, and underexposure conditions. The exposure fusion algorithm performs fusion calculations based on the contrast, saturation, and exposure metrics of each pixel in the image to obtain the finally fused image as the preprocessed hub image.

3. The hub defect detection method based on 2.5D photometric stereo images according to claim 2, wherein: The fusion process of the exposure fusion algorithm for the hub images includes: Calculating the contrast metric, saturation metric, and exposure metric of each pixel; Calculating the final weight of each pixel based on the contrast metric, saturation metric, and exposure metric of each pixel; The final weight values of each pixel are used to construct a weight map.

4. A hub defect detection method based on 2.5D photometric stereo images according to claim 3, characterized in that: The preprocessing also includes: using a multi-scale fusion method to process the weight map image. First, use the Laplacian pyramid to decompose the input original image to obtain image representations at different scales; then perform Gaussian pyramid processing on the calculated weight map to obtain the weight representation corresponding to the scale of the Laplacian pyramid; then multiply the Laplacian pyramid and the Gaussian pyramid at the corresponding scale to obtain the fused Laplacian pyramid; finally, through the inverse Laplacian process, represent the fused Laplacian pyramid as a fused image, and use the fused image as the input of the pre-trained deep learning model.

5. A hub defect detection method based on 2.5D photometric stereo images according to any one of claims 1-4, characterized in that: The pre-trained deep learning model includes an unsupervised learning model and a supervised learning model; Among them, the fused image is first sent into the unsupervised learning model for defect identification. If a defect is identified at this time, the detection defect result is output; otherwise, the fused image is sent into the supervised learning model again for re-identification, and the corresponding defect detection result is output based on the supervised learning model.

6. A hub defect detection method based on 2.5D photometric stereo images according to any one of claims 1-4, characterized in that: Use a 2.5D photometric stereo camera to collect images of the painted hubs. After each channel image is fused using the exposure fusion algorithm, then divide each channel image of the 2.5D photometric stereo camera into m*n sub-regions, where the size of m*n is the same as the number of channels corresponding to the 2.5D photometric stereo camera; select the sub-images at the same position in each channel for stitching to obtain the complete stitched image of the channel, and use the stitched complete image as the input of the pre-trained deep learning model; Among them, the 2.5D photometric stereo camera irradiates the product to be detected from multiple different angles by using a multi-spectral light source or a standard light source, thereby capturing multiple images, and forming new images by performing corresponding processing on the images to suppress the reflective parts in the images; At the same time, the 2.5D photometric stereo camera is used to capture the three-dimensional information of the object, and according to the reconstructed three-dimensional information and normal information, project the normal information along the x, y, and z directions respectively, thereby highlighting the defects extending along different directions, and realizing the clear display of the defects in the image.

7. The hub defect detection method based on 2.5D photometric stereo images according to claim 5, wherein: The unsupervised learning model is trained through a pre-built training set. Among them, defect-free images are selected from the collected hub images, and then the defect-free images are preprocessed and sent into the unsupervised learning model for training.

8. The hub defect detection method based on 2.5D photometric stereo images according to claim 5, characterized in that: The supervised learning model is trained through a pre-built training set. Among them, defective images are selected from the collected hub images and the images are labeled. After the labeling is completed, a painted hub defect data set is formed, and the supervised learning model is trained through the images in the painted hub defect data set.

9. A hub defect detection method based on 2.5D photometric stereo images according to claim 8, characterized in that: The supervised learning model is improved based on the YOLO network model. The images in the painted hub defect data set are preprocessed and then sent into the supervised learning model for training. The supervised learning model improved based on the YOLO network model includes: Multi-channel input layer: Each channel image of the 2.5D photometric stereo camera is divided into sub-regions of m*n, where the size of m*n is the same as the number of channels corresponding to the 2.5D photometric stereo camera; sub-images at the same position in each channel are selected for splicing, so as to be able to construct a new complete image; Feature extraction and fusion layer: The spliced image is subjected to feature extraction using the CSPDarkNet53 network layer; then the result after feature extraction is used as the input of the Neck layer. The structure of the Neck layer is composed of a CBL layer + SPP module and an FPN + PAN module. The CBL layer is used to further extract features, and the SPP module is used to extract features of different scales to enhance the model's detection ability for multi-scale targets; in the FPN + PAN module, FPN adopts a top-down structure, and the high-level feature information is transmitted to the low-level through upsampling and feature fusion is performed during the transmission process; PAN is a bottom-up structure, which adds a bottom-up path on the basis of FPN to transmit the low-level localization features to the high-level, which helps the low-level features to obtain high-level information; the output end of the Neck part is connected to the Prediction layer to output the image detection result.

10. A method for detecting hub defects based on 2.5D photometric stereo images according to any one of claims 1-4, characterized in that: When the detection result of the deep learning model is defective, calculate the width or area parameter of the circumscribed rectangle of the defect in the image, compare the width or area parameter with a threshold, and perform a secondary judgment on the defective image based on the threshold and output the secondary judgment result.

Citation Information

Patent Citations

  • Attention mechanism and deformable convolution fused hub surface defect detection method and device, electronic equipment and storage medium

    CN118172308A

Cited By

  • Visual inspection system for hub baking varnish quality

    CN121347533A

  • Aluminum alloy hub surface defect identification method and system based on visual inspection

    CN122385619A