Automobile part defect detection method

By fusing two-dimensional images and three-dimensional point cloud data and utilizing feature extraction networks and topological association graph processing, the problem of light and shadow interference in automotive parts inspection is solved, accurate identification and quantitative evaluation of complex defects are achieved, and the stability and accuracy of inspection are improved.

CN120707560AActive Publication Date: 2025-09-26SHAANXI SANYUAN YANGYIHAO AUTOMOBILE CO LTD

Patent Information

Application Number
CN202511140494.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-26
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

In the existing technology, automobile parts defect detection is easily affected by lighting conditions and shadows, making it difficult to accurately identify three-dimensional spatial geometric features, and the accuracy of the detection results is poor.

Method used

By acquiring two-dimensional image data and three-dimensional point cloud data, a depth map is generated and spliced ​​into a multi-channel input tensor. The feature extraction network is used to generate feature saliency maps and deformable convolutions. The anchor-free box detection head is combined for multi-task prediction, and a topological association map is constructed for post-processing.

Benefits of technology

It significantly improves the comprehensive detection capability and detection stability of various defects, can accurately capture the complex defect characteristics with irregular shapes and variable sizes, and output continuous defect severity level scores to ensure the integrity and accuracy of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707560A_ABST
    Figure CN120707560A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses an automobile part defect detection method, which comprises the following steps of: acquiring two-dimensional image data and three-dimensional point cloud data of an automobile part, generating a depth mapping graph from the three-dimensional point cloud data, and splicing the depth mapping graph with a color channel of the two-dimensional image data to form a multi-channel input tensor; the multi-channel input tensor is fed into the feature extraction network to generate a feature saliency map, and the feature saliency map is utilized to guide a deformable convolution module to perform adaptive sampling and convolution on a feature map; performing multi-task prediction on the defect candidate area through an anchor-frame-free detection head; and constructing a topological association graph of the defect candidate regions, calculating a connection weight between nodes, pruning the graph by using a preset weight threshold, and merging the defect candidate regions into a final defect detection result. According to the method, the detection limitation of a single data source under complex illumination and diversified defect forms is overcome, and the comprehensive detection capability and detection stability of various defects are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a method for detecting defects in automobile parts. Background Art

[0002] In the automotive manufacturing industry, the surface quality of parts directly impacts the performance, safety, and aesthetics of the entire vehicle. Therefore, automated, high-precision defect detection is a core component of the quality control system. Currently, machine vision-based inspection technology is the mainstream means of achieving automation. Traditional two-dimensional image inspection methods use industrial cameras to capture images of the part surface and then analyze them using image processing algorithms or classic machine learning models. However, these methods have significant bottlenecks. First, the inspection results are easily affected by environmental factors such as on-site lighting conditions, surface reflections, and shadows, resulting in poor detection stability and repeatability. Second, for defects with three-dimensional geometric characteristics, such as dents, protrusions, and scratch depth, two-dimensional images lack depth information, making accurate identification and quantitative assessment difficult, and often confusing surface stains with geometric deformation defects.

[0003] To overcome the limitations of two-dimensional inspection, some solutions have introduced three-dimensional measurement technologies, such as laser scanning or structured light imaging, to detect geometric deformation defects by acquiring three-dimensional point cloud data of the part surface. Although three-dimensional data can accurately describe the contour and morphology of the part, it lacks key surface information such as color and texture, and its ability to detect non-geometric defects such as rust, color difference, printing errors, and fine cracks is limited. In addition, after deep learning algorithms are applied to defect detection, existing models still face challenges. For example, standard convolutional neural networks (CNNs) use a fixed receptive field and sampling method, which makes feature extraction less adaptable and targeted for irregular defects of varying shapes and sizes in industrial scenarios. At the same time, in the output and integration of detection results, most methods only stop at locating and classifying defects and cannot provide quantitative severity assessments. Moreover, when dealing with dense or banded defect groups, simple post-processing logic (such as non-maximum suppression) can easily lead to accidental deletion or incomplete merging, and cannot truly reflect the overall morphology and spatial correlation of defects. This limits the depth of application of detection results in subsequent intelligent production links such as quality grading and maintenance decision-making. Summary of the Invention

[0004] The present invention provides a method for detecting defects in automobile parts to solve the problem in the above-mentioned prior art that the detection results are easily affected by environmental factors such as on-site lighting conditions, surface reflections and shadows of parts, and it is difficult to accurately identify defects in three-dimensional spatial geometric features, resulting in poor accuracy of detection results.

[0005] The automobile parts defect detection method of the present invention comprises the following steps: Acquire two-dimensional image data and three-dimensional point cloud data of the automotive part to be inspected, project the three-dimensional point cloud data onto a two-dimensional plane to generate a depth map, and concatenate it with the color channel of the two-dimensional image data to form a multi-channel input tensor; Feeding the multi-channel input tensor into a preset feature extraction network, the feature extraction network generates a feature saliency map through a spatial attention module, and uses the feature saliency map to guide the deformable convolution module to adaptively sample and convolve high-saliency areas in the feature map to extract deep contextual features of the part; Based on the deep context features, a multi-task prediction is performed on the defect candidate area through an anchor-free detection head to obtain the center point coordinates, size, category probability and a corresponding continuous defect severity score of the defect candidate area; A topological association graph of the defect candidate areas is constructed, where each defect candidate area is a graph node. The connection weights between the nodes are calculated based on the normalized center distance and intersection-over-union ratio between the defect candidate areas. The graph is pruned using a preset weight threshold, and the defect candidate areas within each connected component are merged into the final defect detection result.

[0006] The automotive parts defect detection method of the present invention constructs a multi-channel input by fusing the color texture information of a two-dimensional image with the geometric depth information of a three-dimensional point cloud, effectively overcoming the detection limitations of a single data source under complex lighting and diverse defect morphologies, and significantly improving the comprehensive detection capability and detection stability of various defects. At the feature extraction level, the spatial attention mechanism is used to guide the deformable convolution, so that the network can focus on potential defect areas and perform feature sampling based on the actual contours of the defects, thereby achieving accurate capture of complex defect features with irregular shapes and variable sizes. In addition, the present invention not only locates and classifies defects, but also outputs a continuous defect severity grade score, providing a precise quantitative basis for quality assessment. The post-processing strategy based on the topological association graph can reasonably merge dense or banded defect clusters, avoiding the accidental deletion or incomplete merging of traditional methods, and ensuring the integrity and accuracy of the final detection results.

[0007] Preferably, the three-dimensional point cloud data is projected onto a two-dimensional plane to generate a depth map, and is spliced ​​with the color channels of the two-dimensional image data to form a multi-channel input tensor, including: using a Z-buffer algorithm to orthogonally project the three-dimensional point cloud data along the Z axis onto the XY plane, obtaining the original depth value of each pixel and normalizing it to the range of [0, 255] to generate a depth map; and splicing the depth map with the three color channels R, G, and B of the two-dimensional image data to form a 4-channel input tensor.

[0008] The Z-buffer algorithm can solve the visual occlusion problem of "near points occluding far points" in three-dimensional space. Ultimately, at each pixel position in the two-dimensional depth map, only the depth value of the three-dimensional point that is actually visible from that perspective is retained.

[0009] Preferably, the feature extraction network generates a feature saliency map through a spatial attention module, including: applying maximum pooling and average pooling operations to the input feature map along the channel dimension to generate two two-dimensional feature maps; splicing the two two-dimensional feature maps to form a dual-channel feature description; and convolution of the dual-channel feature description through a convolution kernel with a size of 7. The convolution layer 7 is convolved, and the output value thereof is mapped to the range of 0 to 1 using the Sigmoid activation function to generate the feature saliency map.

[0010] By performing maximum pooling and average pooling on the input feature map along the channel dimension, key spatial information can be captured from different angles, thereby enhancing effective features and suppressing redundant information.

[0011] Preferably, the method of using the feature saliency map to guide the deformable convolution module to perform adaptive sampling and convolution on the high-saliency area in the feature map includes: performing element-wise multiplication of the feature saliency map and the input feature map to obtain a weighted feature map; feeding the weighted feature map into a 3D 3, the bypass convolution layer specifically learns the two-dimensional offset of the sampling grid; the learned two-dimensional offset is applied to the regular grid sampling points of the deformable convolution module, so that the regular grid sampling points are aggregated to the area with stronger feature response, completing adaptive sampling and convolution.

[0012] Preferably, the multi-task prediction of the defect candidate area by the anchor-free frame detection head includes: the anchor-free frame detection head contains four parallel prediction branches, each prediction branch consists of two 3×3 convolution layers and one 1×1 convolution layer; the four prediction branches are the first branch, the second branch, the third branch and the fourth branch, the first branch is used to predict the defect center point heat map; the second branch is used to predict the width and height of the defect area; the third branch is used to predict the category probability of the defect; the fourth branch is used to regress a continuous value between 0 and 1 as the defect severity level score.

[0013] The four parallel prediction branches predict defects from different angles, which can capture defect characteristics more comprehensively. By predicting the heat map, width and height of the defect center point, the defect position and range can be located more accurately; the predicted defect category probability can accurately determine the defect type; the predicted severity level score can further provide a quantitative basis for defect assessment, thereby improving overall detection accuracy.

[0014] Preferably, the connection weights between nodes are calculated based on the normalized center distance and intersection-over-union ratio between the defect candidate regions, including: for any two nodes i and j, the connection weights Calculated by the following formula: ; in, is the intersection-over-union ratio of the defect candidate regions corresponding to nodes i and j, is the Euclidean distance between the center points of two defect candidate regions, σ is a preset hyperparameter used to control the decay rate of the Gaussian function, and exp() is an exponential function.

[0015] Preferably, the method of pruning the graph using a preset weight threshold and merging the defect candidate areas within each connected component into the final defect detection result includes: removing all edges in the topological association graph whose connection weights are less than the preset weight threshold to complete the graph pruning; in the pruned graph, merging all defect candidate areas belonging to the same connected component into a final defect, the bounding box of the final defect being determined by the minimum circumscribed rectangle of all defect candidate areas within the connected component, and its defect severity level score being the weighted average of the corresponding scores of all defect candidate areas within the connected component, with the weight being the category probability score of each defect candidate area.

[0016] Preferably, the original depth value is normalized to the range of [0, 255] using a maximum-minimum normalization method.

[0017] The maximum and minimum normalization method can only linearly scale the original depth values ​​without changing the relative size relationship and overall distribution trend of the data, ensuring that the normalized depth map can still reflect the real spatial geometric relationship.

[0018] Preferably, the acquisition of two-dimensional image data and three-dimensional point cloud data of the automobile part to be inspected includes: acquiring an RGB three-channel two-dimensional image with a resolution of 2048×2048 pixels through an industrial line array camera, and acquiring three-dimensional point cloud data of the surface of the automobile part using a structured light three-dimensional scanner.

[0019] Preferably, feeding the multi-channel input tensor into a preset feature extraction network includes: feeding the multi-channel input tensor into a feature extraction network with ResNet-50 as the backbone.

[0020] The beneficial effects of the present invention are as follows: by fusing the color texture information of the two-dimensional image with the geometric depth information of the three-dimensional point cloud, a multi-channel input is constructed, which effectively overcomes the detection limitations of a single data source under complex lighting and diverse defect forms, and significantly improves the comprehensive detection capability and detection stability of various defects. At the feature extraction level, the spatial attention mechanism is used to guide the deformable convolution, so that the network can focus on potential defect areas and perform feature sampling based on the actual contours of the defects, thereby achieving accurate capture of complex defect features with irregular shapes and variable sizes. In addition, the present invention not only locates and classifies defects, but also outputs a continuous defect severity grade score, providing a precise quantitative basis for quality assessment. The post-processing strategy based on the topological association graph can reasonably merge dense or banded defect clusters, avoiding the accidental deletion or incomplete merging of traditional methods, and ensuring the integrity and accuracy of the final detection results. By predicting defects from different angles through four parallel prediction branches, the defect characteristics can be captured more comprehensively. By predicting the heat map, width and height of the defect center point, the defect position and range can be located more accurately. The predicted defect category probability can accurately determine the defect type. The predicted severity level score can further provide a quantitative basis for defect assessment, thereby improving overall detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 A schematic flow chart of a method for detecting defects in automotive parts provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to be used to explain the present invention, but should not be understood as limiting the present invention.

[0023] like Figure 1 As shown, the automobile parts defect detection method provided by the embodiment of the present invention specifically includes the following steps: S1. Acquire two-dimensional image data and three-dimensional point cloud data of the automotive part to be inspected, project the three-dimensional point cloud data onto a two-dimensional plane to generate a depth map, and splice it with the color channel of the two-dimensional image data to form a multi-channel input tensor.

[0024] Specifically, an industrial line scan camera acquires a 2048×2048 pixel RGB three-channel 2D image, while a structured light 3D scanner simultaneously captures 3D point cloud data of the automotive part surface. Using orthogonal projection, the Z coordinate of each point in the 3D point cloud is normalized to a range of 0 to 255, generating a single-channel depth map. Finally, this single-channel depth map is used as the fourth channel and concatenated with the R, G, and B color channels of the 2D image to form a 4×2048×2048 four-channel input tensor.

[0025] S2. Feed the multi-channel input tensor into a preset feature extraction network. The feature extraction network generates a feature saliency map through a spatial attention module, and uses the feature saliency map to guide the deformable convolution module to adaptively sample and convolve the high-significance areas in the feature map to extract deep contextual features of the part.

[0026] Specifically, the four-channel input tensor is fed into a feature extraction network based on ResNet-50. After the third and fourth residual stages of the feature extraction network, a spatial attention module in a convolutional block attention module is inserted. The spatial attention module performs global average pooling and maximum pooling operations on the input feature map, and then concatenates the results through a convolutional layer to generate a two-dimensional feature saliency map. The value of each pixel in the feature saliency map reflects the importance of the corresponding position. Subsequently, the feature saliency map and the original feature map are input into the deformable convolution module DCNv2. The value of the feature saliency map is used to guide the deformable convolution module DCNv2 to perform non-grid sampling on the feature map, so that the convolution operation focuses on the defect-related areas identified by the feature saliency map, thereby extracting more discriminative deep features.

[0027] S3. Based on the deep context features, a multi-task prediction is performed on the defect candidate area through an anchor-free frame detection head to obtain the center point coordinates, size, category probability and a corresponding continuous defect severity level score of the defect candidate area.

[0028] Specifically, an anchor-free detection head similar to FCOS is adopted. The anchor-free detection head contains three parallel prediction branches, which act on the multi-scale feature map output by the feature extraction network. The first branch is the classification branch, which uses a convolutional layer with a Sigmoid activation function to predict the probability of each pixel belonging to each type of defect. The second branch is the regression branch, which predicts the four distances from the current pixel to the upper, lower, left and right distances of the bounding box of the corresponding defect candidate area, thereby determining the coordinates and size of the center point of the defect candidate area. The third branch is the severity level prediction branch, which maps the feature into a continuous numerical value through a fully connected layer, and then normalizes it through the Sigmoid function and multiplies it by a preset maximum level value, such as 10, to obtain a continuous defect severity level score ranging from 0 to 10.

[0029] S4. Construct a topological association graph of the defect candidate areas, where each defect candidate area is a graph node. Calculate the connection weights between nodes based on the normalized center distance and intersection-over-union ratio between the defect candidate areas. Prune the graph using a preset weight threshold, and merge the defect candidate areas within each connected component into the final defect detection result.

[0030] Specifically, all defect candidate regions output in the previous step with confidence levels higher than a preset threshold are used as the initial nodes of the graph. For any two nodes, the Euclidean distance between the center points of their bounding boxes is calculated and normalized by dividing it by the length of the image diagonal to obtain the normalized center distance. At the same time, the intersection-union ratio of the two nodes is calculated. . The connection weight is a weighted combination of these two indicators, for example, the weight is equal to 0.5 multiplied by the inverse of the normalized center distance plus 0.5 multiplied by the intersection-union ratio. Set a weight threshold, such as 0.7, and remove all connection edges with weights lower than the weight threshold from the graph to complete the graph pruning. Finally, use the breadth-first search algorithm to traverse the pruned graph and find all connected components. For each connected component, merge the bounding boxes of all defect candidate areas contained in it to form a minimum enclosing rectangle that can completely enclose all defects in the defect candidate area as the final defect detection result.

[0031] In an optional embodiment, the 3D point cloud data is projected onto a 2D plane to generate a depth map, which is then concatenated with the color channels of the 2D image data to form a multi-channel input tensor. This process includes: using a Z-buffer algorithm to orthogonally project the 3D point cloud data onto the XY plane along the Z axis, obtaining the original depth value of each pixel and normalizing it to the range [0, 255] to generate a depth map; and concatenating the depth map with the R, G, and B color channels of the 2D image data to form a 4-channel input tensor. These steps are intended to fuse 3D geometric information with 2D texture information. First, the collected 3D point cloud data is processed using a Z-buffer algorithm. This process can be imagined as looking down at the entire point cloud from directly above the Z axis and projecting each 3D point (X, Y, Z) onto the XY 2D plane. When multiple points are projected onto the same pixel coordinate, the Z-buffer algorithm retains the point with the smallest Z value, i.e., the point closest to the projection plane. This Z value is recorded as the original depth value of the pixel.

[0032] Specifically, after obtaining the original depth values ​​of all pixels, they need to be normalized in order to be combined with the image data. For example, assuming that the Z coordinate range of all points in the scene is 1000 mm to 1500 mm, then the normalized depth value of a point with a Z value of 1250 mm is , approximately equal to 128. This converts the entire point cloud into a grayscale image, or depth map. Each pixel in this map has a grayscale value between 0 and 255, representing the relative distance of the object. This depth map is stacked with the original R, G, and B three-channel color image to form a 1920×1080×4 four-channel tensor, which serves as the input for the subsequent neural network, enabling it to simultaneously perceive color texture and spatial depth.

[0033] In an optional embodiment, the feature extraction network generates a feature saliency map through a spatial attention module, including: applying maximum pooling and average pooling operations to the input feature map along the channel dimension to generate two two-dimensional feature maps; splicing the two two-dimensional feature maps to form a dual-channel feature description; convolving the dual-channel feature description through a convolution layer with a convolution kernel size of 7×7, and using a Sigmoid activation function to map its output value to a range of 0 to 1 to generate the feature saliency map.

[0034] The purpose of the spatial attention module is to allow the feature extraction network to focus on the most informative areas in the feature map. Assume that the input is a feature map of size H, W, and C, for example, 64×64×256. First, the spatial attention module performs two pooling operations along the channel dimension. For each spatial position (h, w) on the feature map, max pooling selects the maximum value from its 256 channel values, while average pooling calculates the average of these 256 values. After the operation is completed, two two-dimensional feature maps of size H by W are generated, which capture the most prominent feature response and overall feature information at each position, respectively.

[0035] Two H by W two-dimensional feature maps are stacked together to form a two-channel feature description of size H by W by 2. The descriptor is fed into a convolutional layer using a larger 7×7 convolution kernel, which helps integrate contextual information from a wider spatial range to determine the importance of a region. The output of the convolution operation is passed through a sigmoid activation function, which compresses the output value at each location to between 0 and 1. The resulting H by W single-channel map is the feature saliency map. Pixels with values ​​close to 1 correspond to areas in the original feature map that are more likely to contain key information such as defects.

[0036] In an optional embodiment, the feature saliency map is used to guide the deformable convolution module to adaptively sample and convolve high-significance areas in the feature map, including: element-wise multiplication of the feature saliency map and the input feature map to obtain a weighted feature map; feeding the weighted feature map into a 3×3 bypass convolution layer, which specifically learns the two-dimensional offset of the sampling grid; applying the learned two-dimensional offset to the regular grid sampling points of the deformable convolution module, so that the regular grid sampling points are aggregated to areas with stronger feature responses, thereby completing adaptive sampling and convolution.

[0037] Specifically, the generated feature saliency map is used to guide feature extraction. This feature saliency map, with values ​​between 0 and 1, is element-wise multiplied with the input feature map. This operation is equivalent to assigning a weight to each spatial location in the input feature map. The feature values ​​of highly significant regions remain essentially unchanged or are enhanced, while the feature values ​​of less significant background regions are suppressed and approach zero, resulting in a weighted feature map.

[0038] The weighted feature map is fed into an independent bypass convolution layer, such as a 3×3 convolution layer. The task of the 3×3 convolution layer is not to extract features for classification or positioning, but to learn the offset of the sampling points for the deformable convolution. For a standard 3×3 convolution kernel, it has 9 sampling points, and the bypass convolution layer predicts a two-dimensional offset vector (Δx, Δy) for each sampling point of the center pixel. Since the input is saliency weighted, the learned offset will naturally point to the direction where the feature response is stronger. Finally, in the deformable convolution module of the main path, the position of its sampling points is no longer a fixed grid, but a regular grid point plus the learned offset. For example, a pixel that should be at The point where the position is sampled will now be By sampling at different locations, the receptive field of the convolution operation can dynamically fit the actual shape of the defect, achieving accurate feature extraction of irregular targets.

[0039] In an optional embodiment, the multi-task prediction of the defect candidate area by the anchor-free frame detection head includes: the anchor-free frame detection head includes four parallel prediction branches, each prediction branch consists of two 3×3 convolution layers and one 1×1 convolution layer; the four prediction branches are respectively a first branch, a second branch, a third branch and a fourth branch, the first branch is used to predict the defect center point heat map; the second branch is used to predict the width and height of the defect area; the third branch is used to predict the category probability of the defect; the fourth branch is used to regress a continuous value between 0 and 1 as the defect severity level score.

[0040] Specifically, after the feature map is processed by the feature extraction network and the spatial attention module, it is fed into the anchor-free detection head for final prediction. The anchor-free detection head does not rely on pre-set anchor boxes, but instead makes predictions directly at each position in the feature map. It consists of four parallel prediction branches with different functions. Each prediction branch is a small network structure, for example, deepening features through two 3×3 convolutional layers and then adjusting the number of feature channels to the required output dimension through a 1×1 convolutional layer.

[0041] The first branch outputs a single-channel defect center heat map. The value of each pixel in the map represents the probability that the pixel location is the defect center. For example, a pixel location with a value greater than 0.7 is considered a potential defect center. The second branch outputs two values ​​for each pixel location on the map, predicting the width and height of the defect bounding box centered on this location, respectively, in pixels. The third branch is responsible for classification. If three types of defects need to be identified, three probability values ​​will be output for each pixel location, corresponding to the possibility of being a first-class, second-class, or third-class defect. The fourth branch outputs a single continuous value, such as 0.85. This continuous value directly quantifies the severity of the defect at that location. 0 represents harmless and 1 represents the most serious, thus achieving end-to-end regression of the defect level.

[0042] In an optional embodiment, the calculation of the connection weights between nodes based on the normalized center distance and intersection-over-union ratio between the defect candidate regions includes: for any two nodes i and j, the connection weights Calculated by the following formula: ; in, is the intersection-over-union ratio of the defect candidate regions corresponding to nodes i and j, is the Euclidean distance between the center points of two defect candidate regions, σ is a preset hyperparameter used to control the decay rate of the Gaussian function, and exp() is an exponential function.

[0043] Specifically, the anchor-free detection head generates a large number of dense defect candidate regions, which need to be organized and filtered. Each defect candidate region is regarded as a node in a graph. In order to determine whether any two defect candidate regions describe the same real defect, it is necessary to calculate the connection strength between them, that is, the edge connection weight. . Connection weight It is obtained by multiplying the two parts, taking into account the shape overlap and spatial distance. The first part is the intersection and union ratio , which calculates the ratio of the intersection area to the union area of ​​two defect candidate regions, with a value between 0 and 1. The larger the value, the higher the overlap between the two defect candidate regions. The second part is a Gaussian penalty term related to the center point distance. Calculate the Euclidean distance between the center points of two defect candidate regions, for example, 20 pixels. The square of the Euclidean distance between the center points of the two defect candidate regions is divided by the square of a hyperparameter, sigma, which can be preset to, for example, 50. As the distance increases, the exponential term approaches 0, significantly weakening the connection weight between two distant defect candidate regions, even if they overlap. A high connection weight is achieved only when two defect candidate regions have significant overlap and are close together.

[0044] In an optional embodiment, the graph is pruned using a preset weight threshold, and the defect candidate areas within each connected component are merged into a final defect detection result, including: removing all edges in the topological association graph whose connection weights are less than a preset weight threshold to complete the graph pruning; in the pruned graph, all defect candidate areas belonging to the same connected component are merged into a final defect, the bounding box of the final defect is determined by the minimum circumscribed rectangle of all defect candidate areas in the connected component, and its defect severity level score is the weighted average of the corresponding scores of all defect candidate areas in the connected component, and the weight is the category probability score of each defect candidate area.

[0045] Specifically, after constructing a fully connected graph containing all defect candidate region nodes and the connection weights between them, it needs to be simplified to distinguish different defect instances. Set a weight threshold, such as 0.4. Then traverse all edges in the graph and remove all edges with weights less than 0.4. This process is called graph pruning, which disconnects the connections between defect candidate regions that are not strongly correlated, so that the original complex graph is decomposed into several independent subgraphs, each of which is called a connected component. After pruning, all nodes in the same connected component are considered to describe a single real defect together, so they need to be merged. The merging process first determines the position and size of the final defect, which is achieved by calculating the minimum enclosing rectangle of all defect candidate regions in the connected component, that is, finding a minimum rectangle that can just surround all these defect candidate regions. Determine the severity level of the final defect using weighted average. For example, there are defect candidate regions A and defect candidate region B in a connected component, with defect severity level scores of 0.9 and 0.7 respectively. The model's confidence in their category predictions is 0.95 and 0.80 respectively, so the final defect severity level score is In this way, defect candidate areas with higher confidence have a greater say in determining the final defect severity score.

[0046] The implementation principle of the automobile parts defect detection method of the embodiment of the present invention is: by fusing the color texture information of the two-dimensional image with the geometric depth information of the three-dimensional point cloud, a multi-channel input is constructed, which can overcome the detection limitations of a single data source under complex lighting and diverse defect forms, and significantly improve the comprehensive detection capability and detection stability of various defects. At the feature extraction level, the spatial attention module is used to guide the deformable convolution, so that the feature extraction network can focus on potential defect areas and perform feature sampling based on the actual contours of the defects, thereby achieving accurate capture of complex defect features with irregular shapes and variable sizes. In addition, the present invention not only locates and classifies defects, but also outputs a continuous defect severity grade score, providing a precise quantitative basis for quality assessment. Moreover, the post-processing strategy based on the topological association graph can reasonably merge dense or banded defect clusters to prevent accidental deletion or incomplete merging, thereby ensuring the integrity and accuracy of the final detection results.

[0047] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A method for detecting defects in automobile parts, characterized in that: The steps include: Acquire two-dimensional image data and three-dimensional point cloud data of the automotive part to be inspected, project the three-dimensional point cloud data onto a two-dimensional plane to generate a depth map, and concatenate it with the color channel of the two-dimensional image data to form a multi-channel input tensor; Feeding the multi-channel input tensor into a preset feature extraction network, the feature extraction network generates a feature saliency map through a spatial attention module, and uses the feature saliency map to guide the deformable convolution module to adaptively sample and convolve high-saliency areas in the feature map to extract deep contextual features of the part; Based on the deep context features, a multi-task prediction is performed on the defect candidate area through an anchor-free detection head to obtain the center point coordinates, size, category probability and a corresponding continuous defect severity score of the defect candidate area; A topological association graph of the defect candidate areas is constructed, where each defect candidate area is a graph node. The connection weights between the nodes are calculated based on the normalized center distance and intersection-over-union ratio between the defect candidate areas. The graph is pruned using a preset weight threshold, and the defect candidate areas within each connected component are merged into the final defect detection result.

2. The method for detecting defects in automobile parts according to claim 1, wherein: The three-dimensional point cloud data is projected onto a two-dimensional plane to generate a depth map, and is spliced ​​with the color channels of the two-dimensional image data to form a multi-channel input tensor, including: using a Z-buffer algorithm to orthogonally project the three-dimensional point cloud data along the Z axis onto the XY plane, obtaining the original depth value of each pixel and normalizing it to the range of [0, 255] to generate a depth map; and splicing the depth map with the R, G, and B color channels of the two-dimensional image data to form a four-channel input tensor.

3. The method for detecting defects in automobile parts according to claim 1, wherein: The feature extraction network generates a feature saliency map through a spatial attention module, including: applying maximum pooling and average pooling operations to the input feature map along the channel dimension to generate two two-dimensional feature maps; splicing the two two-dimensional feature maps to form a dual-channel feature description; convolving the dual-channel feature description through a convolutional layer with a convolution kernel size of 7×7, and mapping its output value to a range of 0 to 1 using a sigmoid activation function to generate the feature saliency map.

4. The method for detecting defects in automobile parts according to claim 1, wherein: The method uses the feature saliency map to guide the deformable convolution module to adaptively sample and convolve high-saliency areas in the feature map, including: element-wise multiplication of the feature saliency map and the input feature map to obtain a weighted feature map; feeding the weighted feature map into a 3×3 bypass convolution layer, which specifically learns the two-dimensional offset of the sampling grid; and applying the learned two-dimensional offset to the regular grid sampling points of the deformable convolution module, so that the regular grid sampling points are aggregated towards areas with stronger feature responses, thereby completing adaptive sampling and convolution.

5. The method for detecting defects in automobile parts according to claim 1, wherein: The multi-task prediction of defect candidate areas by the anchor-free frame detection head includes: the anchor-free frame detection head contains four parallel prediction branches, each prediction branch consists of two 3×3 convolutional layers and one 1×1 convolutional layer; the four prediction branches are respectively a first branch, a second branch, a third branch and a fourth branch, the first branch is used to predict the defect center point heat map; the second branch is used to predict the width and height of the defect area; the third branch is used to predict the category probability of the defect; and the fourth branch is used to regress a continuous value between 0 and 1 as the defect severity level score.

6. The automobile parts defect detection method according to claim 1, characterized in that: The calculation of the connection weights between nodes based on the normalized center distance and intersection-over-union ratio between the defect candidate regions includes: For any two nodes i and j, their connection weight Calculated by the following formula: ; in, is the intersection-over-union ratio of the defect candidate regions corresponding to nodes i and j, is the Euclidean distance between the center points of two defect candidate regions, σ is a preset hyperparameter used to control the decay rate of the Gaussian function, and exp() is an exponential function.

7. The method for detecting defects in automobile parts according to claim 1, wherein: The method prunes the graph using a preset weight threshold and merges the defect candidate areas within each connected component into a final defect detection result, including: removing all edges in the topological association graph whose connection weights are less than the preset weight threshold to complete the graph pruning; in the pruned graph, merging all defect candidate areas belonging to the same connected component into a final defect, wherein the bounding box of the final defect is determined by the minimum circumscribed rectangle of all defect candidate areas within the connected component, and the defect severity level score of the final defect is a weighted average of the corresponding scores of all defect candidate areas within the connected component, with the weight being the category probability score of each defect candidate area.

8. The method for detecting defects in automobile parts according to claim 2, wherein: The original depth value is normalized to the range of [0, 255] using the maximum and minimum normalization method.

9. The method for detecting defects in automobile parts according to claim 1, wherein: The method of obtaining the two-dimensional image data and three-dimensional point cloud data of the automobile parts to be inspected includes: obtaining the resolution of The RGB three-channel two-dimensional image of the pixel is used to obtain the three-dimensional point cloud data of the surface of the automobile part using a structured light three-dimensional scanner.

10. The method for detecting defects in automobile parts according to claim 1, wherein: Feeding the multi-channel input tensor into a preset feature extraction network includes: feeding the multi-channel input tensor into a feature extraction network with ResNet-50 as the backbone.

Citation Information

Patent Citations

  • Precise plastic mold detection method, device and equipment and storage medium

    CN119164955A

  • Weld joint positioning and defect detection system and method based on deep learning

    CN119417800A

  • Method and system for analyzing defects in wafer manufacturing based on big data

    CN119580022A

  • Automobile part size correction method based on image processing

    CN120298271A

  • Method for determining whether medication adherence has been fulfilled and server using same

    KR102344101B1

Cited By

  • Road defect self-adaptive detection method and system

    CN121707965A