Hub defect detection method and system based on multiple visual features

Through the visual multi-feature detection method, combined with RGB images and 3D point cloud data, deep and lightweight convolutional neural networks are used to extract features, build a feature library and perform defect judgment, which solves the problems of low wheel hub detection accuracy and high cost and achieves efficient defect recognition.

CN120765639AActive Publication Date: 2025-10-10UNIV OF JINAN

Patent Information

Application Number
CN202511245473.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-10-10
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing wheel hub defect detection methods rely on single feature recognition, which has low detection accuracy and high cost. In addition, it is difficult to determine the feature splicing weight when fusing multiple features, resulting in inaccurate detection.

Method used

A detection method based on visual multi-features is adopted. RGB images and 3D point clouds are synchronously collected through a color camera and a 3D profilometer. Deep convolutional neural networks, residual neural networks and lightweight convolutional neural networks are combined to extract features, build a feature library and perform defect judgment. The fast Bayesian core set is used to control the size of the feature library.

Benefits of technology

It improves the accuracy of wheel hub defect detection, solves the problem of inaccurate detection from a single data source, avoids the problem of difficulty in determining feature splicing weights, and achieves efficient defect identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765639A_ABST
    Figure CN120765639A_ABST
Patent Text Reader

Abstract

The invention provides a hub defect detection method and system based on multiple visual features, and relates to the technical field of hub visual detection, and the method comprises the steps: synchronously collecting an RGB image and a three-dimensional point cloud of a hub based on a color camera and a three-dimensional contourgraph; projecting the three-dimensional point cloud to obtain a multi-view projection, and performing down-sampling on the acquired RGB image and the three-dimensional point cloud; extracting three-dimensional point cloud features, RGB image features and multi-view projection features; respectively storing the three features contained in each training sample into corresponding feature libraries; and obtaining three features of the test sample according to the same steps, calculating feature scores based on a feature library, and performing defect judgment based on a defect threshold value. According to the method, a plurality of feature libraries are established, the problems of inaccurate detection and positioning of a single data source and insufficient abnormal data are solved, classification and recognition are carried out through the overall feature difference, the method does not depend on local high-gradient features, various features are independently recognized, and fusion of different features is not needed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of wheel hub visual inspection technology, and more specifically, to a wheel hub defect detection method and system based on visual multi-features. Background Art

[0002] As a core component of the vehicle's driving system, the quality of the wheel hub directly affects driving safety and comfort. Defects such as wheel hub cracks, air holes, and wear are the main causes of wheel hub breakage during high-speed driving, and traditional manual visual inspection has a high rate of missed detection. As automobile manufacturing transitions to "intelligent and automated" manufacturing, production line beats are increasing. Traditional inspection methods are unable to match high-speed assembly line operations, resulting in low inspection accuracy and high inspection costs. In order to identify wheel hub defects through machine vision and improve the efficiency of wheel hub defect detection, there are several typical technical improvement directions in the existing technology: (1) Single feature comparison - comparison with the sample under test: The Chinese invention patent with publication number CN114972325A discloses a method for detecting defects in automobile wheels based on image processing, including collecting an X-ray image of the wheel; obtaining difference pixels; obtaining non-structural pixels of the difference pixels; obtaining defective and non-defective periodic regions based on the frequency of non-structural pixels in each periodic region; calculating the defect probability of each pixel in the defective periodic region based on the noise coincidence rate of each pixel in the defective periodic region and the noise description weight of the non-defective periodic region; and obtaining the defective region based on the defect probability of each pixel. The Chinese invention patent with publication number CN119359639A also adopts a similar technical solution. However, this solution is highly dependent on the structural regularity and symmetry of the wheel, as well as on specific defect types that exhibit high gradients in image recognition, and the defects identified by a single feature are limited.

[0003] (2) Single feature comparison - compared with the original design: The Chinese invention patent with publication number CN118037661A discloses a wheel hub surface defect detection method, device, equipment, storage medium and product. The method aligns the template wheel hub image corresponding to the template wheel hub and the test wheel hub image corresponding to the test wheel hub to obtain the image to be inspected; the candidate defect area is determined in the overall difference image between the image to be inspected and the template wheel hub image. This solution also has the problem of relying on specific defect types that show high gradients in image recognition and the defects that can be identified by a single feature are limited.

[0004] (3) Mutual comparison of multiple features – same-position matching: Chinese invention patent publication number CN114419038A discloses a method and device, storage medium, and electronic equipment for identifying wheel hub surface defects. The method collects two-dimensional images of the sample being tested to obtain a first candidate defect list, and collects three-dimensional information to obtain a second candidate defect list. Based on the first and second candidate defect lists, the same location is identified to determine whether the target wheel hub has defects. Chinese invention patent publication number CN117095002A also adopts a similar solution, which requires different features to have similar recognition capabilities for the same defect, resulting in a fundamental flaw in the underlying logic.

[0005] (4) Multi-feature fusion: The Chinese invention patent with publication number CN118172308A discloses a wheel hub surface defect detection method, device, electronic device and storage medium that integrates attention mechanism and deformable convolution, including inputting an enhanced data set into a pre-trained neural network to obtain the detection result of the wheel hub to be inspected; wherein, the pre-trained neural network includes a feature extraction module that integrates a multimodal attention mechanism and a multi-channel feature extraction module that integrates deformable convolution.

[0006] Although this solution enhances features from the same source, it actually has the ability to jointly identify features from multiple sources. The difficulty of multi-feature fusion recognition lies in determining the weights of multiple features when splicing them together. This solution does not address this issue, but instead directly splices them together.

[0007] Based on the above-mentioned deficiencies of the existing wheel hub defect identification solutions, an effective wheel hub defect detection method is still needed. Summary of the Invention

[0008] To solve the above problems, the technical solution adopted in this application is a wheel hub defect detection method based on visual multi-features, which includes the following steps: Data acquisition: The RGB image and 3D point cloud of the wheel hub are collected synchronously using a color camera and a 3D profilometer; Data preprocessing: Project the 3D point cloud onto multiple views to obtain multi-view projections, and downsample the collected RGB images and 3D point clouds; Feature extraction: Deep convolutional neural networks are used to extract 3D point cloud features; residual neural networks are used to extract RGB image features; lightweight convolutional neural networks are used to extract multi-view projection features, and then feature aggregation is performed on the multi-view projection features; Feature library construction: The 3D point cloud features, RGB images, and multi-view projection features contained in each training sample are stored in the corresponding feature library respectively; Defect identification: According to the above data collection, data preprocessing and feature extraction steps, three features of the test sample are obtained, the feature score is calculated based on the feature library, and the defect is determined based on the defect threshold; the test sample identified as defective is segmented and located, and the features of the defective test sample are stored in the defect library; the features of the test sample identified as normal are stored in the feature library, and the fast Bayesian core set sampling method is used to control the size of the feature library.

[0009] Optionally, projecting the three-dimensional point cloud onto multiple views to obtain multi-view projections includes: projecting the three-dimensional point cloud onto the same image physical coordinate system as the two-dimensional RGB image according to an external parameter matrix and an internal parameter matrix of the color camera to obtain a reference projection view; using a plane where the reference view is located as a reference plane, performing a rotation transformation on the reference view using a rotation function to obtain a rotated projection view, generating rotated projection views of an even number of point clouds on different planes, where the absolute value of the rotation angle does not exceed pi / 12; Data preprocessing also includes: for multi-view projection, the baseline view and the projection view obtained by rotation transformation are normalized and contrast enhanced using CLAHE, and bilateral filtering is used to reduce the noise of the image while retaining edge information; then the segmentation module divides each view into non-overlapping patches of fixed size.

[0010] Optionally, data preprocessing of the RGB image includes applying a nearest neighbor interpolation algorithm to the RGB image, downsampling the R, G, and B channels respectively, sequentially applying height threshold filtering, normal vector consistency filtering, and connected domain analysis to remove low workbench background, workbench surface background, and isolated noise points, and then applying a bicubic interpolation algorithm to the downsampled image.

[0011] Optionally, data preprocessing of the three-dimensional point cloud includes downsampling using a body centroid method, and then sequentially using height threshold filtering, normal vector consistency filtering, and connected domain analysis to remove low workbench background, workbench surface background, and isolated noise points.

[0012] Optionally, extracting RGB image features using a residual neural network includes: Perform two convolution operations on the pre-processed RGB image to increase the number of input channels; Then it enters the first level containing two multi-scale branches. After feature extraction from the two branches, high-resolution features are output after feature splicing, channel compression, and residual connection. Then it enters the second level, uses the convolution layer for downsampling, and then goes through three multi-scale branches to obtain medium-resolution features; Then enter the third level, use the convolution layer for downsampling, and then go through four multi-scale branches to obtain low-resolution features; The feature maps of the second and third levels are upsampled to the same size as the first level, and are concatenated together to form a feature map, and then the number of channels is compressed through the fusion module; Finally, the patch-level feature map is obtained after downsampling and reshaping.

[0013] Optionally, using a deep convolutional neural network to extract 3D point cloud features includes: The downsampled 3D point cloud is divided into The features of each voxel grid are aggregated, and the entire point cloud is meshed and output as , is the number of non-empty voxels; Convert the voxel grid into a dense tensor, use 3D convolution to extract its overall geometric features, and output the geometric feature tensor; by non-empty voxels as the center and the radius is , the neighborhood voxels within the search radius are formed Patch; Take the features of each patch as input and enter the spatial attention mechanism network to apply spatial attention to each patch; Perform maximum pooling aggregation to obtain patch features of specified dimensions.

[0014] Optionally, extracting multi-view projection features using a lightweight convolutional neural network includes: sequentially passing the segmented patch through a first convolution layer, a maximum pooling layer, a second convolution layer, a maximum pooling layer, a third convolution layer, an average pooling layer, and an L2 normalization layer for feature extraction; After feature extraction, the feature map is reorganized through view alignment. The rotation function is used to align all rotated projection view features to the reference view coordinate system. After the alignment is completed, the position code of each position is added to the feature. The weight is calculated according to the feature variance of each view, and then weighted averaging is performed to achieve cross-view fusion. The fused feature vector is reorganized into a feature map to complete the feature aggregation of multi-view projection features.

[0015] Optionally, calculating the feature score based on the feature library includes: setting, is the point cloud feature library, is the RGB image feature library, is a multi-view image feature library, For any feature library, is a sample data of the test sample, For The feature set extracted from the feature set is calculated according to the following formula: ; Where, is the feature score of a certain patch of the sample, is a feature in any feature library, for For a single Patch feature in Each Patch feature in traverses the feature library M , and the feature library The feature with the smallest difference Recorded as ,exist Among all the Patch features in The most different features Recorded as .

[0016] Optionally, defect determination based on a defect threshold includes: determining whether there is a defect based on the defect threshold for the three-dimensional point cloud features, RGB image features, and multi-view projection features of the test sample; if two or more features indicate that the sample has a defect, it is directly determined to have a defect; if only one feature indicates that the sample has a defect, manual intervention is introduced to identify and determine the final result; if it is determined to be a defect, the features of the test sample are stored in a defect library; if it is determined not to be a defect, the features of the test sample are stored in three feature libraries; The defect threshold is the sum of the mean and 2 times the standard deviation of the feature score.

[0017] The present application also provides a wheel hub defect detection system based on visual multi-features, which is suitable for implementing any of the aforementioned wheel hub defect detection methods based on visual multi-features, including: Data acquisition module: used to synchronously acquire the RGB image and 3D point cloud of the wheel hub based on a color camera and a 3D profilometer; Data preprocessing module: used to project the 3D point cloud into multiple views to obtain multi-view projections, and downsample the collected RGB images and 3D point clouds; Feature extraction module: used to extract 3D point cloud features using a deep convolutional neural network; extract RGB image features using a residual neural network; extract multi-view projection features using a lightweight convolutional neural network, and then perform feature aggregation on the multi-view projection features; Defect Identification Module: This module is used to call the data acquisition module, data preprocessing module, and feature extraction module to obtain three features of the test sample, calculate the feature score based on the feature library module, and make defect determinations based on the defect threshold; it segments and locates the test sample identified as defective, and stores the features of the defective test sample in the defect library module; it stores the features of the test sample identified as normal in the feature library module; Feature library module: used to store the three features contained in the training samples and the test samples identified as normal, and use the fast Bayesian core set sampling method to control the size of the feature library; Defect library module: used to store three types of features contained in test samples identified as defects.

[0018] The beneficial effects of the wheel hub defect detection method and system based on multi-feature visual features provided by this application are: This application establishes a feature library based on RGB image features, 3D point cloud features and multi-view projection features, calculates feature scores based on the feature library, and makes defect judgments based on defect thresholds, thereby solving the problem of inaccurate detection and positioning from a single data source. The feature library is constructed by normal samples, which solves the problem of insufficient abnormal data. Classification and identification are performed through overall feature differences, without relying on local high-gradient features. Various features are independently identified, and there is no problem of different features having different recognition capabilities for the same defect. There is no need to fuse different features, and there is no problem of difficulty in determining feature splicing weights. Various data sources are effectively utilized to improve detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art.

[0020] Figure 1 This is a general flow chart of the wheel hub defect detection method based on multi-feature vision provided in an embodiment of the present application; Figure 2 This is a diagram of the RGB image feature extraction network structure provided by the embodiment of the present application; Figure 3 This is a diagram of the point cloud feature extraction network structure provided by an embodiment of the present application; Figure 4 This is a diagram of the multi-view image feature extraction network structure provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to make the technical problems, technical solutions and beneficial effects to be solved by this application more clearly understood, this application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0022] like Figure 1 As shown, this application provides a wheel hub defect detection method based on visual multi-features, which is performed during the training phase: Data acquisition: The RGB image and 3D point cloud of the wheel hub are collected synchronously using a color camera and a 3D profilometer; Data preprocessing: Project the 3D point cloud onto multiple views to obtain multi-view projections, and downsample the collected RGB images and 3D point clouds; Feature extraction: Deep convolutional neural networks are used to extract 3D point cloud features; residual neural networks are used to extract RGB image features; lightweight convolutional neural networks are used to extract multi-view projection features, and then feature aggregation is performed on the multi-view projection features; In order to distinguish between neural networks for extracting three-dimensional point cloud features and neural networks for extracting multi-view projection features, this application calls the neural network for extracting three-dimensional point cloud features that includes multi-stage feature transformation and has a deep architecture a "deep" convolutional neural network; and the neural network for extracting multi-view projection features that only includes a 7-layer network architecture is called a "lightweight" convolutional neural network.

[0023] Feature library construction: The 3D point cloud features, RGB images and multi-view projection features contained in each training sample are stored in the corresponding feature library respectively; after the features of the above three types of data are extracted, they are stored in three different feature libraries, namely the point cloud feature library , RGB image feature library , multi-view image feature library When the feature library is initially constructed, all sample features are stored in it. As the feature library is regularly supplemented, its size will continue to grow. To improve the inference efficiency of the feature library, a fast Bayesian core set sampling method is used to maintain the size of the feature library. The size of each feature library is limited to 3,000,000 features.

[0024] During the training phase, multiple sample data sets must be collected, with at least 1,000 samples. Data collection uses a color camera to acquire RGB images, requiring a resolution of at least 4,000 × 3,000 pixels and a frame rate of at least 30 frames per second. A high-precision 3D profilometer is used to acquire wheel hub profile data, stored as a point cloud. The profilometer's measurement accuracy must reach ±0.01 mm, and its scanning speed must be at least 5,000 points per second. A synchronous trigger mechanism is used to capture RGB images and 3D point clouds, ensuring time synchronization. All samples collected during the training phase are normal.

[0025] During the testing phase, perform: Defect identification: According to the above data collection, data preprocessing and feature extraction steps, three features of the test sample are obtained, the feature score is calculated based on the feature library, and the defect is determined based on the defect threshold; the test sample identified as defective is segmented and located, and the features of the defective test sample are stored in the defect library; the features of the test sample identified as normal are stored in the feature library, and the fast Bayesian core set sampling method is used to control the size of the feature library.

[0026] Due to the relatively small amount of wheel hub defect data, all training samples were normal. When collecting RGB training samples, the color camera's height and shooting angle were pre-set. In this embodiment, the camera was placed perpendicular to the wheel hub surface, with the lens facing the hub and 30 cm from the hub surface. Point cloud data was acquired using a high-precision 3D profilometer and stored in PCD format. Triggering was used between the color camera and the high-precision 3D profilometer to synchronize sampling times. This embodiment used 1,568 samples, all of which were local wheel hub data.

[0027] RGB image data preprocessing involves applying the nearest neighbor interpolation algorithm to the RGB image, downsampling the R, G, and B channels separately. Height threshold filtering, normal consistency filtering, and connected domain analysis are then used to remove low workbench background, workbench surface background, and isolated noise points. The downsampled image is then subjected to bicubic interpolation. In this embodiment, the nearest neighbor interpolation algorithm is applied to a 4000×3000 RGB image, downsampling the R, G, and B channels separately to obtain a 512×512 image. The background is then removed. The bicubic interpolation algorithm is then applied to the 512×512 RGB image to obtain a 224×224 RGB image.

[0028] like Figure 2 As shown in the figure, the use of residual neural network to extract RGB image features includes first using convolution to extract and embed patch features of the image, then using convolution kernels with different expansion rates in the wide residual neural network to extract features of different scales, and finally performing channel compression on features of different scales to fuse multi-scale features. Specifically: Perform two convolution operations on the pre-processed RGB image to increase the number of input channels; Then it enters the first level containing two multi-scale branches. After feature extraction from the two branches, high-resolution features are output after feature splicing, channel compression, and residual connection. Then it enters the second level, uses the convolution layer for downsampling, and then goes through three multi-scale branches to obtain medium-resolution features; Then enter the third level, use the convolution layer for downsampling, and then go through four multi-scale branches to obtain low-resolution features; The feature maps of the second and third levels are upsampled to the same size as the first level, and are concatenated together to form a feature map, and then the number of channels is compressed through the fusion module; Finally, the patch-level feature map is obtained after downsampling and reshaping.

[0029] In this embodiment, the acquired 224×224 RGB image is convolved twice to increase the number of input channels from 3 to 64. Then it enters the first level, which contains two multi-scale branches with expansion rates of 1 and 2 respectively. After feature extraction by the two branches, high-resolution features are output after feature splicing, channel compression, and residual connection, with a size of 224×224×64. Then it enters the second level, and uses a convolution layer with a stride of 2 and a convolution kernel of 3×3 for downsampling. The number of channels is 128, and then passes through three multi-scale branches with expansion rates of 1, 2, and 4 to obtain medium-resolution features. Then it enters the third level, and uses a convolution layer with a stride of 2 and a convolution kernel of 3×3 for downsampling. The number of channels is 56, and then passes through four multi-scale branches with expansion rates of 1, 2, 4, and 6 to obtain low-resolution features. The feature maps of the second and third levels are upsampled to the same size as the first level and concatenated to form a 224×224×192 feature map. This map is then compressed to 64 channels using a fusion module. The fused features are then enhanced for detail, downsampled to 56×56×64 using a 4×4 convolution kernel, and reshaped to 3136×64 to produce a patch-level feature map.

[0030] Perform data preprocessing and feature extraction on all training samples in batches to obtain the RGB image color features of all samples and store them in the RGB image feature library .

[0031] For 3D point clouds, data preprocessing involves downsampling using the volumetric centroid method, followed by height threshold filtering, normal vector consistency filtering, and connected domain analysis to remove low workbench background, workbench surface background, and isolated noise points. In this example, a 2mm voxel downsampling scheme was used, retaining approximately 50,000 points. The normal vector for each point was then calculated, and height threshold filtering, normal vector consistency filtering, and connected domain analysis were used to remove low workbench background, workbench surface background, and isolated noise points.

[0032] like Figure 3 As shown in Figure 2, the deep convolutional neural network is used to extract 3D point cloud features including: The downsampled 3D point cloud is divided into The features of each voxel grid are aggregated, and the entire point cloud is meshed and output as , is the number of non-empty voxels; Convert the voxel grid into a dense tensor, use 3D convolution to extract its overall geometric features, and output the geometric feature tensor; by non-empty voxels as the center and the radius is , the neighborhood voxels within the search radius are formed Patch; Take the features of each patch as input and enter the spatial attention mechanism network to apply spatial attention to each patch; Perform maximum pooling aggregation to obtain patch features of specified dimensions.

[0033] In this embodiment, the downsampled point cloud is divided into The features of each voxel grid are aggregated, and each voxel grid outputs a feature vector with a dimension of 3. The entire point cloud is gridded and output as , The voxel grid is converted into a dense tensor, and the overall geometric features are extracted using 3D convolution, and the geometric feature tensor is output. ,in , and Respectively represent the number of voxel grids in the directions of the three coordinate axes x, y, and z. non-empty voxels as the center and the radius is , search for neighboring voxels within the radius, each patch contains up to 32 voxels, forming Patches, then we know that the feature dimension of each patch is Taking the features of each patch as input, it enters the spatial attention mechanism network, applies spatial attention to each patch, and the output dimension is Perform maximum pooling to compress the features of each patch and obtain the patch features of the specified dimension. The output is .

[0034] Perform data preprocessing and feature extraction on all training samples in batches to obtain the point cloud features of all samples and store them in the point cloud feature library. .

[0035] Projecting the 3D point cloud onto multiple views to obtain multi-view projections includes: projecting the 3D point cloud onto the same image physical coordinate system as the 2D RGB image based on the extrinsic parameter matrix and the intrinsic parameter matrix of the color camera to obtain a reference projection view; using the plane where the reference view is located as the reference plane, rotating the reference view using a rotation function to obtain a rotated projection view, generating rotated projection views of an even number of point clouds on different planes, with the absolute value of the rotation angle not exceeding pi / 12; make is the color camera intrinsic parameter matrix, is the external parameter matrix, For the The pixel plane coordinates of the point, is the first point cloud after downsampling The spatial coordinates of a point are as follows: ; in is a normalization function that converts the depth value in the 224×224 grid into a grayscale value. The view obtained according to the above formula is the reference view. The coordinate system used is the right-hand coordinate system. are the Euler angles of rotation around the xyz axis, respectively. When both are 0, that is, the base view The plane where is located is the same as the plane where the RGB two-dimensional image is located. , ,right , Evenly take 4 pairs of values ​​to generate the rotation function : ; use To base view Perform the rotation operation to obtain the other 4 rotated views.

[0036] like Figure 4 As shown in the figure, for multi-view projection, the base view and the projection view obtained by rotation transformation are normalized and contrast enhanced by CLAHE, and bilateral filtering is used to reduce the noise of the image while retaining the edge information; then the segmentation module divides each view into non-overlapping patches of fixed size.

[0037] The above five views are input into the preprocessing module for normalization, CLAHE contrast enhancement, and bilateral filtering to reduce the noise of the image while retaining the edge information. The output shape of this module is The segmentation module divides each view into non-overlapping patches of size 8×8. Each view is divided into 784 patches, and the 5 views have a total of 3920 patches. The output shape is Tensor of .

[0038] like Figure 4 As shown in FIG, a lightweight convolutional neural network is used to extract multi-view projection features, including: passing the segmented patch through the first convolution layer, the maximum pooling layer, the second convolution layer, the maximum pooling layer, the third convolution layer, the average pooling layer and the L2 normalization layer for feature extraction; After feature extraction, the feature map is reorganized through view alignment. The rotation function is used to align all rotated projection view features to the reference view coordinate system. After the alignment is completed, the position code of each position is added to the feature. The weight is calculated according to the feature variance of each view, and then weighted averaging is performed to achieve cross-view fusion. The fused feature vector is reorganized into a feature map to complete the feature aggregation of multi-view projection features.

[0039] The feature extraction module is a lightweight CNN module with only 7 layers and outputs The view alignment module reorganizes the feature map into , using the known rotation matrix to align all features to the reference view coordinate system. The feature aggregation module adds the position code of each position to the feature, calculates the weight according to the feature variance of each view, and then performs weighted averaging to achieve cross-view fusion. The fused feature vector is reorganized into Characteristic map of size.

[0040] Perform data preprocessing and feature extraction on all training samples in batches to obtain multi-view image features of all samples and store them in the multi-view image feature library .

[0041] The above-mentioned RGB image feature library is established , point cloud feature library , multi-view image feature library The steps can be performed in parallel or in series, which can be selected according to the computing power and parallelism of the device.

[0042] When the feature library is initially constructed, all sample features are stored in it. As the feature library is supplemented, its size increases. To improve the inference efficiency of the feature library, a fast Bayesian core set sampling method is used to maintain the size of the feature library. The size of each feature library is limited to 3,000,000 features.

[0043] As a feasible method to control the size of the feature library, when the number of features in the feature library reaches the upper limit of the feature library, it is reduced to less than 1 / 3 of the upper limit through the fast Bayesian core set sampling method.

[0044] In the defect detection phase, the same hardware equipment is still used to collect RGB image data and 3D point cloud data of the wheel to be inspected, and generate multi-view data. The three types of data are subjected to feature extraction in the above steps respectively. The extracted features are retrieved and compared with the features in the feature library, and defect scoring is performed. The scoring scheme for all features is consistent. Let, is the point cloud feature library, is the RGB image feature library, is a multi-view image feature library, For any feature library, is a sample data of the test sample, For The feature set extracted from the feature set is calculated according to the following formula: ; Where, is the feature score of a certain patch of the sample, is a feature in any feature library, for For a single Patch feature in Each Patch feature in traverses the feature library M , and the feature library The feature with the smallest difference Recorded as ,exist Among all the Patch features in The most different features Recorded as .

[0045] After obtaining the score of the sample Patch feature, according to the predefined threshold Determine whether a defect exists. If two or more features indicate a defect, the sample is directly judged as defective. If only one feature indicates a defect, manual intervention (such as visual judgment) is introduced to determine the final result. If it is determined to be a defect, the sample feature is stored in the defect database; if it is not determined to be a defect, the feature is stored in three feature databases.

[0046] Threshold According to the characteristic distribution of normal samples, in this embodiment, the sum of the mean and 2 times the standard deviation is used as the threshold. Patch features with characteristic values ​​less than or equal to the sum of the mean and 2 times the standard deviation of the characteristic score are considered normal samples, and patch features with characteristic values ​​greater than the sum of the mean and 2 times the standard deviation of the characteristic score are considered defective samples.

[0047] For patches identified as defects, upsampling is used to resize the patch features to the original image size, thereby locating the defects. For patches requiring manual intervention, the same method is used to locate the defects, improving the efficiency of manual identification.

[0048] The above examples are only used to illustrate the technical solutions of the present application, but not limit the same; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that the technical solutions recorded in the foregoing examples can be modified, or some technical features can be replaced by equivalent ones; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A wheel hub defect detection method based on visual multi-features, characterized in that: The following steps are involved: Data acquisition: The RGB image and 3D point cloud of the wheel hub are collected synchronously using a color camera and a 3D profilometer; Data preprocessing: Project the 3D point cloud onto multiple views to obtain multi-view projections, and downsample the collected RGB images and 3D point clouds; Feature extraction: Deep convolutional neural networks are used to extract 3D point cloud features; residual neural networks are used to extract RGB image features; lightweight convolutional neural networks are used to extract multi-view projection features, and then feature aggregation is performed on the multi-view projection features; Feature library construction: The 3D point cloud features, RGB images, and multi-view projection features contained in each training sample are stored in the corresponding feature library respectively; Defect identification: According to the above data collection, data preprocessing and feature extraction steps, the three features of the test sample are obtained, the feature scores are calculated based on the feature library, and the defect is determined based on the defect threshold; Segment and locate the defects of the test samples identified as defects, and store the features of the defective test samples in the defect database; The features of the test samples identified as normal are stored in the feature library, and the size of the feature library is controlled using the fast Bayesian core set sampling method.

2. The wheel hub defect detection method based on visual multi-features according to claim 1, characterized in that: The projecting of the three-dimensional point cloud onto multiple views to obtain multi-view projections includes: projecting the three-dimensional point cloud onto the same image physical coordinate system as the two-dimensional RGB image based on the external parameter matrix and the internal parameter matrix of the color camera to obtain a reference projection view; using the plane where the reference view is located as the reference plane, performing a rotation transformation on the reference view using a rotation function to obtain a rotated projection view, generating rotated projection views of an even number of point clouds on different planes, wherein the absolute value of the rotation angle does not exceed pi / 12; The data preprocessing also includes: for multi-view projection, normalizing and contrast enhancing the base view and the projection view obtained by rotation transformation, and using bilateral filtering to reduce noise on the image while retaining edge information; and then using a segmentation module to segment each view into non-overlapping patches of fixed size.

3. The wheel hub defect detection method based on visual multi-features according to claim 1, characterized in that: Data preprocessing of RGB images includes using the nearest neighbor interpolation algorithm to downsample the R, G, and B channels respectively, and then using height threshold filtering, normal vector consistency filtering, and connected domain analysis in sequence to remove the low workbench background, workbench surface background, and isolated noise points. The downsampled image is then subjected to the bicubic interpolation algorithm.

4. The wheel hub defect detection method based on visual multi-features according to claim 1, characterized in that: The data preprocessing of the 3D point cloud includes downsampling using the body centroid method, followed by height threshold filtering, normal vector consistency filtering, and connected domain analysis to remove low workbench background, workbench surface background, and isolated noise points.

5. The wheel hub defect detection method based on visual multi-features according to claim 3 is characterized in that: The extraction of RGB image features using a residual neural network includes: Perform two convolution operations on the pre-processed RGB image to increase the number of input channels; Then it enters the first level containing two multi-scale branches. After feature extraction from the two branches, high-resolution features are output after feature splicing, channel compression, and residual connection. Then it enters the second level, uses the convolution layer for downsampling, and then goes through three multi-scale branches to obtain medium-resolution features; Then enter the third level, use the convolution layer for downsampling, and then go through four multi-scale branches to obtain low-resolution features; The feature maps of the second and third levels are upsampled to the same size as the first level, and are concatenated together to form a feature map, and then the number of channels is compressed through the fusion module; Finally, the patch-level feature map is obtained after downsampling and reshaping.

6. The wheel hub defect detection method based on visual multi-features according to claim 4, characterized in that: The method of extracting three-dimensional point cloud features using a deep convolutional neural network includes: The downsampled 3D point cloud is divided into The features of each voxel grid are aggregated, and the entire point cloud is meshed and output as , is the number of non-empty voxels; Convert the voxel grid into a dense tensor, use 3D convolution to extract its overall geometric features, and output the geometric feature tensor; by non-empty voxels as the center and the radius is , the neighborhood voxels within the search radius are formed Patch; Take the features of each patch as input and enter the spatial attention mechanism network to apply spatial attention to each patch; Perform maximum pooling aggregation to obtain patch features of specified dimensions.

7. The wheel hub defect detection method based on visual multi-features according to claim 2, characterized in that: The lightweight convolutional neural network is used to extract multi-view projection features, which includes: sequentially passing the segmented patch through a first convolutional layer, a maximum pooling layer, a second convolutional layer, a maximum pooling layer, a third convolutional layer, an average pooling layer, and an L2 normalization layer for feature extraction; After feature extraction, the feature map is reorganized through view alignment. The rotation function is used to align all rotated projection view features to the reference view coordinate system. After the alignment is completed, the position code of each position is added to the feature. The weight is calculated according to the feature variance of each view, and then weighted averaging is performed to achieve cross-view fusion. The fused feature vector is reorganized into a feature map to complete the feature aggregation of multi-view projection features.

8. The wheel hub defect detection method based on visual multi-features according to claim 1, characterized in that: Calculating feature scores based on the feature library includes: setting, is the point cloud feature library, is the RGB image feature library, is a multi-view image feature library, For any feature library, is a sample data of the test sample, For The feature set extracted from the feature set is calculated according to the following formula: ; Where, is the feature score of a certain patch of the sample, is a feature in any feature library, for For a single Patch feature in Each Patch feature in traverses the feature library M , and the feature library The feature with the smallest difference Recorded as ,exist Among all the Patch features in The most different features Recorded as .

9. The wheel hub defect detection method based on visual multi-features according to claim 1, characterized in that: The defect determination based on the defect threshold includes: determining whether there is a defect based on the three-dimensional point cloud features, RGB image features, and multi-view projection features of the test sample according to the defect threshold; if two or more features indicate that the sample has a defect, it is directly determined to have a defect; if only one feature indicates that the sample has a defect, manual intervention is introduced to identify and determine the final result; if it is determined to be a defect, the features of the test sample are stored in a defect library; if it is determined not to be a defect, the features of the test sample are stored in three feature libraries; The defect threshold is the sum of the mean value and 2 times the standard deviation of the feature score.

10. A wheel hub defect detection system based on visual multi-features, characterized in that: Suitable for implementing the wheel hub defect detection method based on visual multi-features according to any one of claims 1 to 9, comprising: Data acquisition module: used to synchronously acquire the RGB image and 3D point cloud of the wheel hub based on a color camera and a 3D profilometer; Data preprocessing module: used to project the 3D point cloud into multiple views to obtain multi-view projections, and downsample the collected RGB images and 3D point clouds; Feature extraction module: used to extract 3D point cloud features using a deep convolutional neural network; extract RGB image features using a residual neural network; extract multi-view projection features using a lightweight convolutional neural network, and then perform feature aggregation on the multi-view projection features; Defect Identification Module: This module is used to call the data acquisition module, data preprocessing module, and feature extraction module to obtain three features of the test sample, calculate the feature score based on the feature library module, and make defect determinations based on the defect threshold; it segments and locates the test sample identified as defective, and stores the features of the defective test sample in the defect library module; it stores the features of the test sample identified as normal in the feature library module; Feature library module: used to store the three features contained in the training samples and the test samples identified as normal, and use the fast Bayesian core set sampling method to control the size of the feature library; Defect library module: used to store three types of features contained in test samples identified as defects.

Citation Information

Patent Citations

  • Automobile hub defect detection method based on image processing

    CN114972325A

  • Hub defect detection method and device and storage medium

    CN117095002A

  • Hub apparent defect detection method, device and equipment, storage medium and product

    CN118037661A

  • Attention mechanism and deformable convolution fused hub surface defect detection method and device, electronic equipment and storage medium

    CN118172308A

  • Defect category identification method of hub image, medium and equipment

    CN119359639A

Cited By

  • Three-dimensional defect detection method, system and device based on typical geometric alignment and multi-view adaptive rendering

    CN122415608A