An image recognition system for detecting defects in industrial parts
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]但是其在实际使用时,仍旧存在一些缺点,如(1)人工目视+手感:依赖经验,检出率受光照角度、人员疲劳影响大,漏检率普遍过高,且无法实现数据闭环
[0060] This invention utilizes a multi-axis controllable light source array module to provide time-division stroboscopic illumination to the surface of parts using different incident angle-intensity combinations and differentiated spectra (Δλ≥30nm). Combined with a high-speed polarization camera module, it synchronously captures multiple frames of polarization image sequences by rotating a polarizer. Edge computing nodes perform photometric stereoscopic computation and differentiable rendering on the images to generate flash-sensitive feature maps. A dual-branch semantic segmentation network fuses geometric and polarization features through a cross-attention mechanism, outputting a pixel-level flash mask. Finally, a physical constraint post-processing module projects the two-dimensional detection results onto the part's CAD coordinate system, performs three-dimensional geometric verification using tolerance bands and height thresholds, and generates a quantitative defect report. This invention possesses the ability to accurately identify minute flashes and suppress interference from complex working conditions, significantly improving the stability and reliability of automated quality inspection, and achieving high-precision, automated detection of flashes on metal/plastic parts.
Smart Images

Figure CN121582224B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of component defect detection technology, and more specifically, to an image recognition system for industrial component defect detection. Background Technology
[0002] In modern manufacturing, the quality and precision of industrial parts directly affect the performance and reliability of the final product. Flash defects, as a common surface flaw in machining, injection molding, and other processes, not only affect the appearance quality of parts but can also adversely impact assembly accuracy, sealing performance, and even the overall structural strength.
[0003] Flash is excess thin, wing-like material that accumulates on the parting surface, cutting edge exit, or ridge line of a part during machining or forming. It typically has a thickness of 10–150 μm and a height of 20–500 μm. Its presence affects subsequent assembly accuracy and may break off in high-speed moving parts, causing equipment failure. In high-end applications such as automotive powertrain systems, hydraulic valve plates, and 3C structural components, "zero flash" has become a mandatory quality standard.
[0004] However, it still has some drawbacks in actual use, such as (1) manual visual inspection + touch: relying on experience, the detection rate is greatly affected by the angle of light and the fatigue of personnel, the missed detection rate is generally too high, and the data closed loop cannot be achieved.
[0005] (2) Contact measuring tools (height gauge, needle gauge): can only be inspected on the accessible edges, which is inefficient and easily crushes thin flash, resulting in "false qualification". Summary of the Invention
[0006] Therefore, this application provides an image recognition system for detecting defects in industrial parts to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, this application provides the following technical solution:
[0008] Firstly, the multi-axis controllable light source array module is configured to provide time-division strobe illumination of the part surface with different incident angle-intensity combinations within the same sampling period.
[0009] High-speed polarization camera module: The polarizer angle θ is synchronously rotated to a preset angle corresponding to the current incident angle-intensity combination at each stroboscopic moment. Capture N≥3 frames of polarization image sequence ;
[0010] Edge computing node module: This module embeds a fly-edge sensitive feature extraction unit. The module obtains the fly-edge sensitive feature map F through the following steps:
[0011] For preset angle Perform photometric stereo calculations to generate an elevation map Z and a surface normal map N;
[0012] Input Z and N into a differentiable rendering layer and reconstruct the ideal contour map Ĉ without fly edges;
[0013] Compare the actual contour map C with By performing a difference operation, an initial defect map is obtained. ;
[0014] right Perform ridge response filtering based on the Hessian matrix to retain pixels with curvature κ>κth and directional continuity ψ>ψth, and obtain the fly-edge sensitive feature map F;
[0015] The dual-branch semantic segmentation network module: the first branch takes the fly-edge sensitive feature map F as input, and the second branch takes... The pseudo-color fusion image is used as input, and the two branches exchange weights bidirectionally in the decoding stage through the attention gating unit, outputting a pixel-level flyedge mask M;
[0016] The physical constraint post-processing module projects pixel-level flash masks onto the coordinate system of the part's CAD model, utilizing the tolerance band T and the flash height threshold. Perform 3D geometric verification while retaining the 2D mask area requirement. And three-dimensional height The example of burr edge is used to obtain the final defect report.
[0017] Optionally, in the multi-axis controllable light source array module, the module is synchronously linked with the industrial camera of the image acquisition unit through a timing control unit. Within a single sampling period, the timing control unit drives each light source unit to perform time-division strobe sequentially with different incident angle-intensity combinations according to a preset illumination sequence. The illumination duration of each combination is set to 5-20ms according to the camera exposure parameters, and the time interval between two adjacent strobe flashes is ≤1ms, completing the acquisition of multiple sets of differentiated illumination images. At the same time, each light source unit is equipped with LEDs of different wavelengths, with the selected wavelengths covering the visible light range of 380nm-780nm; for example, the main wavelength of light source unit 1 is set to 400nm (violet light), light source unit 2 to 450nm (blue light), light source unit 3 to 500nm (green light), light source unit 4 to 560nm (yellow light), light source unit 5 to 620nm (red light), and light source unit 6 to 680nm (near-infrared light). The difference in the main wavelength of any two light source units is ≥60nm, far exceeding Δλ≥30nm. The basic requirements highlight the grayscale differences in different types of flash defects under different spectra.
[0018] Optionally, in the high-speed polarization camera module, the high-speed polarization camera module and the multi-axis controllable light source array module work together via a time synchronization signal; the polarizer at the front end of the camera is driven by a servo motor rotary table; during the time when the incident angle-intensity combined light source is triggered to flicker, the camera control system receives a synchronization pulse signal, and based on the synchronization pulse signal, the polarizer rotates to a preset angle pre-bound to the specific lighting condition with a millisecond-level response speed. For example, when using a low-angle grazing incidence infrared light source, the polarizer will rotate synchronously to... (Relative to a certain reference plane), it captures the polarized light component parallel to the direction of the surface micro-grooves; when switching to a high-angle vertically illuminating blue light source, the polarizer rotates to [a specific position] at the next strobe instant. Subsequently, during the third set of spectral combination illumination, the polarizer will be positioned again. In this way, for the same displacement of the part surface, within the same complete composite illumination sampling period, the camera captures a sequence of polarization images containing N≥3 frames. Each frame of the image contains information about the surface topography and spectral reflectance under illumination from different spectra and angles; this complete image sequence The polarization degree image and polarization angle image of the surface are further calculated using the Stokes vector polarization analysis algorithm.
[0019] Optionally, in the high-speed polarization camera module, the high-speed polarization camera module and the multi-axis controllable light source array module work together via a time synchronization signal; the polarizer at the front end of the camera is driven by a servo motor rotary table; during the time when the incident angle-intensity combined light source is triggered to flicker, the camera control system receives a synchronization pulse signal, and based on the synchronization pulse signal, the polarizer rotates to a preset angle pre-bound to the specific lighting condition with a millisecond-level response speed. For example, when using a low-angle grazing incidence infrared light source, the polarizer will rotate synchronously to... (Relative to a certain reference plane), it captures the polarized light component parallel to the direction of the surface micro-grooves; when switching to a high-angle vertically illuminating blue light source, the polarizer rotates to [a specific position] at the next strobe instant. Subsequently, during the third set of spectral combination illumination, the polarizer will be positioned again. In this way, for the same displacement of the part surface, within the same complete composite illumination sampling period, the camera captures a sequence of polarization images containing N≥3 frames. Each frame of the image contains information about the surface topography and spectral reflectance under illumination from different spectra and angles; this complete image sequence The polarization degree image and polarization angle image of the surface are further calculated using the Stokes vector polarization analysis algorithm.
[0020] Optionally, the dual-branch semantic segmentation network module adopts a symmetrical U-shaped architecture, comprising three parts: a feature encoding layer, a cross-attention gating fusion layer, and a decoding output layer. The inputs of the two branches undergo targeted preprocessing to adapt to feature extraction requirements.
[0021] In the input processing of the first branch (flyedge sensitive feature branch), the flyedge sensitive feature map output by the edge calculation node module is used as input. Normalization processing is performed before input to map pixel values to the [0,1] interval. The specific calculation method is as follows:
[0022]
[0023] in, Represented as coordinates after normalization. Pixel value at; Represented as the original image in coordinates Pixel value at that location, It represents the maximum value among all pixel values in the original image. Represented as the minimum value among all pixel values in the original image;
[0024] Then, a 1×1 convolution kernel is used to increase the channel dimension, providing a multi-channel feature basis for subsequent coding layers, while keeping the feature map resolution unchanged;
[0025] In the second branch (polarization image fusion branch) input processing, the N≥3 frame polarization image sequence acquired by the high-speed polarization camera module is used. The pseudo-color fusion image is used as input; the pseudo-color fusion process uses a multi-channel feature mapping algorithm: first, for each frame... Polarization parameters are calculated, and the polarization degree (DoP) and polarization angle (AoP) features are extracted. Then, DoP, AoP, and the original grayscale values are combined to form a 3-channel feature map. Through HSV color space conversion, DoP is mapped to the Hue channel, AoP to the Saturation channel, and the grayscale values to the Value channel, generating a pseudo-color fusion image. ;
[0026] For each frame Polarization analysis is performed using Stokes vectors to calculate the polarization degree and polarization angle images of the surface; when N=4 ( , , , When ), Stokes vector The specific calculation method is as follows:
[0027]
[0028]
[0029]
[0030] in, The total light intensity component is represented as the Stokes vector. The horizontal-vertical linear polarization components, represented as the Stokes vector, Represented as a Stokes vector - Linear polarization component;
[0031] Note: When N=3, it can be achieved through... , , The light intensity value is obtained by fitting the solution using the least squares method. , , ;
[0032] Based on the Stokes vector, the specific methods for calculating the degree of polarization (DoP) and the polarization angle (AoP) are as follows:
[0033]
[0034]
[0035] The degree of polarization image highlights the birefringence region caused by stress, while the polarization angle image shows the orientation of the surface microstructure; for example, the DoP value of the metal flash region is ≥0.35, while the DoP value of the bulk is ≤0.15; the AoP value of the plastic flash region is offset by ≥5° relative to the bulk. These quantitative characteristics provide a basis for judgment for the flash defect detection module.
[0036] In the feature coding layer, the two branches of the coding layer have the same structure but independent parameters. Each branch contains four downsampling modules, each consisting of two 3×3 convolutional layers, one batch normalization layer, and one 2×2 max pooling layer. Through progressive downsampling, the number of channels is gradually increased from 16 to 256, extracting features from low-level details to high-level semantic features. Specifically, the first branch of the coding layer captures the geometric structural features of the flyedge, such as the length, width, and curvature of the flyedge, and enhances the adaptability to flyedges of different sizes through multi-scale convolutional kernels. The second branch of the coding layer captures polarization characteristics and spectral reflectance features. By adding attention mechanisms to the second and third coding modules, the polarization feature response of the flyedge region is strengthened, and the interference of background noise (such as oil stains and scratches on the surface of the part) is suppressed.
[0037] Optionally, in the physical constraint post-processing module, firstly, the physical constraint post-processing module establishes a mapping relationship between the pixel coordinate system and the coordinate system of the part's CAD model, transforming the two-dimensional flash mask M into a defect candidate region in three-dimensional space; this mapping process is based on system-preset calibration parameters, specifically:
[0038] Camera intrinsic and extrinsic parameter retrieval: Retrieving the intrinsic parameter matrix E (including focal length) of the high-speed polarization camera module. , Principal point coordinates , ) and extrinsic parameter matrix (rotation matrix G, translation vector j), where the extrinsic parameter matrix is obtained in advance through hand-eye calibration, describing the pose relationship between the camera coordinate system and the coordinate system of the part CAD model;
[0039] Pixel coordinates to 3D point cloud conversion: For pixels marked as flash in the flash mask The elevation map generated by the edge computing node module is converted into three-dimensional coordinates in the camera coordinate system using the camera imaging model. The calculation method is as follows:
[0040]
[0041] The coordinates in the camera coordinate system were then converted to three-dimensional coordinates in the part's CAD model coordinate system using an extrinsic parameter matrix. The calculation method is as follows:
[0042]
[0043] The final 3D point cloud of the candidate flash edge region is obtained, and the specific calculation method is as follows:
[0044]
[0045] in, Represented as a 3D point cloud, u represents the total number of fringe pixels in the mask;
[0046] Load the CAD model of the part and extract the surfaces in the model corresponding to the candidate areas for flash. The 3D point cloud is matched with the surface using the ICP algorithm;
[0047] Then, the candidate regions for fly edges are validated through dual physical constraints to eliminate minor interferences:
[0048] In the calculation and selection of 2D mask area, the 2D area V corresponding to the flash pixels in the flash mask is first calculated. The calculation takes into account the image pixel resolution. The specific calculation method is as follows:
[0049]
[0050] Where V represents the area of the two-dimensional mask, Represented as the pixel size in the horizontal direction of the camera. The pixel size is represented in the vertical direction of the camera, and m represents the number of glitter pixels.
[0051] Preset area threshold The tolerances are set according to the part design tolerances and manufacturing process requirements, i.e., precision gear parts. Ordinary plastic casing ;when If it is determined to be a minor interference, the candidate region is directly eliminated; when At that time, proceed to three-dimensional height verification;
[0052] Based on CAD model surface With Fly Edge Candidate Point Cloud Calculate each flash point relative to the surface The normal distance is the height of the flash at that point, and the calculation method is as follows:
[0053]
[0054] in, Represented as in a 3D point cloud On the surface The closest point on, Represented as surface The normal vector at the nearest point;
[0055] Next, calculate the average height of the burr edge area and use the average height as the height of the candidate burr edge area;
[0056] Preset height threshold The threshold is set to ;when When the dimensional deviation is determined to be within the acceptable range, the candidate area is eliminated; when If the candidate region is confirmed to be a real flash defect, it is retained as a flash instance.
[0057] Read the design tolerance data of the currently inspected part from the MES system and automatically match the corresponding... For multi-batch, multi-model parts inspection scenarios, a built-in tolerance zone database is provided, which automatically retrieves the preset T value based on the part model (e.g., T=0.08mm for metal shaft parts, T=0.15mm for plastic connectors), where T represents the tolerance zone.
[0058] After three-dimensional geometric verification, the defect information of the retained flash edge instances is quantified and a report is generated. The final output defect report includes basic defect information, defect quantification parameters, and visualization attachments.
[0059] Compared with the prior art, this application has at least the following beneficial effects:
[0060] This invention utilizes a multi-axis controllable light source array module to provide time-division stroboscopic illumination to the surface of parts using different incident angle-intensity combinations and differentiated spectra (Δλ≥30nm). Combined with a high-speed polarization camera module, it synchronously captures multiple frames of polarization image sequences by rotating a polarizer. Edge computing nodes perform photometric stereoscopic computation and differentiable rendering on the images to generate flash-sensitive feature maps. A dual-branch semantic segmentation network fuses geometric and polarization features through a cross-attention mechanism, outputting a pixel-level flash mask. Finally, a physical constraint post-processing module projects the two-dimensional detection results onto the part's CAD coordinate system, performs three-dimensional geometric verification using tolerance bands and height thresholds, and generates a quantitative defect report. This invention possesses the ability to accurately identify minute flashes and suppress interference from complex working conditions, significantly improving the stability and reliability of automated quality inspection, and achieving high-precision, automated detection of flashes on metal / plastic parts. Attached Figure Description
[0061] To more intuitively illustrate the prior art and this application, exemplary drawings are provided below. It should be understood that the specific shapes and structures shown in the drawings should not generally be regarded as limiting conditions for implementing this application; for example, based on the technical concept disclosed in this application and the exemplary drawings, those skilled in the art are able to easily make conventional adjustments or further optimizations to the addition / reduction / classification, specific shapes, positional relationships, connection methods, size ratios, etc. of certain units (components).
[0062] Figure 1 This is a schematic diagram of the module connection of the present invention.
[0063] Figure 2 This is a schematic diagram of the bi-branch semantic segmentation network process of the present invention.
[0064] Figure 3 This is a schematic diagram of the edge computing node feature extraction process of the present invention; Detailed Implementation
[0065] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0066] Please see Figure 1 As shown, the present invention provides an image recognition system for detecting defects in industrial parts, comprising:
[0067] Multi-axis controllable light source array module: configured to provide time-division strobe illumination of the part surface with different incident angle-intensity combinations within the same sampling period;
[0068] In the multi-axis controllable light source array module, the module is synchronously linked with the industrial camera of the image acquisition unit through a timing control unit. Within a single sampling period, the timing control unit drives each light source unit to perform time-division strobe sequentially with different incident angle-intensity combinations according to a preset illumination sequence. The illumination duration of each combination is set to 5-20ms according to the camera exposure parameters, and the time interval between two adjacent strobe flashes is ≤1ms, completing the acquisition of multiple sets of differentiated illumination images. At the same time, each light source unit is equipped with LEDs of different wavelengths, with selected wavelengths covering the visible light range of 380nm-780nm; for example, the main wavelength of light source unit 1 is set to 400nm (violet light), light source unit 2 to 450nm (blue light), light source unit 3 to 500nm (green light), light source unit 4 to 560nm (yellow light), light source unit 5 to 620nm (red light), and light source unit 6 to 680nm (near-infrared light). The difference in the main wavelength between any two light source units is ≥60nm, far exceeding Δλ≥30nm. The basic requirements highlight the grayscale differences of different types of flash defects under different spectra;
[0069] For burrs at the recessed edges of the part surface, a combination of 30° incident angle and high-power illumination is used to create a clear shadow contrast through strong oblique light, highlighting the height and size of the burrs. For minor burrs in the curved transition area of the part, a combination of 60° incident angle and medium-power illumination is used to reduce interference from surface reflections and clearly present the edge outline of the burrs. Furthermore, by using time-division illumination with different spectra, the material differences between the burrs and the part body are further distinguished. For example, under blue light illumination, the gray value of metal burrs is 20-30 gray levels lower than that of the part body, while under red light illumination, the gray value of plastic burrs is 15-25 gray levels higher than that of the part body. These spectral differences can be captured by the subsequent burr defect detection module and serve as important auxiliary features for defect identification.
[0070] High-speed polarization camera module: The polarizer angle θ is synchronously rotated to a preset angle corresponding to the current incident angle-intensity combination at each stroboscopic moment. Capture N≥3 frames of polarization image sequence ;
[0071] In the high-speed polarization camera module, the high-speed polarization camera module and the multi-axis controllable light source array module work together via a time synchronization signal; the polarizer at the front end of the camera is driven by a servo motor rotary table; during the time when the incident angle-intensity combined light source is triggered to flicker, the camera control system receives a synchronization pulse signal. Based on the synchronization pulse signal, the polarizer rotates to a preset angle pre-bound to the specific lighting condition with a millisecond-level response speed. For example, when using a low-angle grazing incidence infrared light source, the polarizer will rotate synchronously to... (Relative to a certain reference plane), it captures the polarized light component parallel to the direction of the surface micro-grooves; when switching to a high-angle vertically illuminating blue light source, the polarizer rotates to [a specific position] at the next strobe instant. Subsequently, during the third set of spectral combination illumination, the polarizer will be positioned again. In this way, for the same displacement of the part surface, within the same complete composite illumination sampling period, the camera captures a sequence of polarization images containing N≥3 frames. Each frame of the image contains information about the surface topography and spectral reflectance under illumination from different spectra and angles; this complete image sequence The polarization degree image and polarization angle image of the surface are further calculated using the Stokes vector calculation polarization analytical algorithm.
[0072] Each frame of the image The surface morphology and spectral reflectance information contained by illumination from different spectra and angles can be expressed as:
[0073]
[0074] in, Let i be the i-th frame, and let the polarizer rotation angle be . The light intensity value of the image acquired at that time. Represented as a specific spectrum The basic light intensity at the incident angle α.
[0075] Edge computing node module: This module embeds a fly-edge sensitive feature extraction unit. The module obtains the fly-edge sensitive feature map F through the following steps:
[0076] For preset angle Perform photometric stereo calculations to generate an elevation map Z and a surface normal map N;
[0077] Input Z and N into a differentiable rendering layer and reconstruct the ideal contour map without flash by reversing the process. ;
[0078] Compare the actual contour map C with By performing a difference operation, an initial defect map is obtained. ;
[0079] right Perform ridge response filtering based on the Hessian matrix to retain pixels with curvature κ>κth and directional continuity ψ>ψth, and obtain the fly-edge sensitive feature map F;
[0080] The edge computing node module receives a polarization image sequence containing N≥3 frames. Based on photometric stereo vision combined with known illumination parameters (incident angle, light source direction vector) of a multi-axis controllable light source array, photometric stereo calculation is performed on image sequences; this process is based on the Lambertian reflection model, whose reflected light intensity satisfies:
[0081]
[0082] in, Represented as the surface reflectivity of the part. Let be the direction vector of the light source for the i-th illumination. Let G be the surface normal vector, and let G be the ambient light intensity.
[0083] By establishing an overdetermined system of equations for N frames of images, the surface normal vector N(x,y) of each pixel is solved using the least squares method to obtain the surface normal map N. Subsequently, an elevation map Z is generated based on the surface normal map N through integration. The integration process uses the Poisson reconstruction algorithm, and the specific calculation method of the Poisson equation is as follows:
[0084]
[0085] in, Represented as the Laplace operator, Represented as coordinates on a two-dimensional plane The elevation value at that location, Represented as the gradient operator, Represented as coordinates on a two-dimensional plane The normal vector at that point, It is represented as the component of the normal vector N(x,y) in the z direction;
[0086] when At that time, adjacent pixel normal vector interpolation optimization was used, and the final output elevation map Z restored the micro-morphology of the part surface;
[0087] The generated elevation map and surface normal map are input into a differentiable rendering layer, which is based on a physically based rendering (PBR) engine and incorporates the CAD design model of the part. The layer then reconstructs an ideal contour map without flash through reverse rendering. Specifically, the process first extracts the 3D point cloud P(x,y,z) of the actual surface based on the elevation map. Then, the point cloud is registered with the ideal surface of the CAD model. The difference between the actual contour and the ideal contour is minimized using a differentiable loss function, and the geometric parameters of the ideal contour (such as edge curvature and chamfer size) are optimized in reverse. At the same time, the lighting reflection effect of the ideal contour is rendered and corrected by combining the normal vector distribution characteristics of the surface normal map N, ensuring that the grayscale distribution and edge contrast of Ĉ are highly matched with the actual imaging environment. The final output is an ideal contour map without flash. ;
[0088] The specific method for calculating the differentiable loss function is as follows:
[0089]
[0090] Where L represents the loss function, C represents the ideal profile without flash, and C represents the actual profile.
[0091] To obtain the ideal outline without flash. Then, the image sequence acquired by the high-speed polarization camera module is processed first. Edge detection is performed to generate an actual contour map C. The edge detection uses an improved Canny algorithm, which extracts the actual edges of the part through dynamic thresholding, and then eliminates edge breaks through morphological closing operations to obtain a continuous and complete actual contour map C.
[0092] Subsequently, a pixel-by-pixel difference operation is performed between the actual contour image C and the ideal contour image Ĉ. The specific difference calculation method is as follows:
[0093]
[0094] in, Represented as the pixel value of the initial defect map, with a value range of [0,1];
[0095] When pixel value When the deviation from the ideal profile is present, it indicates that there is a burr defect or noise interference at that location; when When this occurs, it indicates that the position completely coincides with the ideal contour; to initially suppress noise, [the following is done]... Perform mean filtering and set a noise threshold. =0.05, will The pixels are set to 0 to obtain the initial defect image after denoising. ;
[0096] Finally, the initial defect image after denoising. Perform ridge response filtering based on the Hessian matrix to filter out pixels with flash defects; calculate... In each pixel The Hessian matrix H at point H is calculated as follows:
[0097]
[0098] in, It is represented as the Hessian matrix at pixel (x,y). It is expressed as the second partial derivative with respect to the x-coordinate. It can be expressed as a mixed second-order partial derivative obtained by first taking the partial derivative with respect to the x-coordinate and then taking the partial derivative with respect to the y-coordinate; It is expressed as the second partial derivative with respect to the y-coordinate;
[0099] By solving the eigenvalues of the Hessian matrix , ( The ridge response value R is calculated using the following method:
[0100]
[0101] Where R represents the ridge response value, , Represented as matrix eigenvalues;
[0102] When the calculated ridge response value When the calculated ridge response value is [value missing], the pixel is located on the center line of the burr edge; when the calculated ridge response value is [value missing], the pixel is located on the center line of the burr edge. When the calculated ridge response value is within a linear structure region, the pixel is likely located within that region. If the pixel does not belong to a linear structure region, then the pixel does not belong to a linear structure region.
[0103] Eigenvalues of the Hessian matrix , The QR algorithm is used to solve the problem by decomposing the matrix into the product of an orthogonal matrix Q and an upper triangular matrix R, and continuously updating the matrix iteratively until the original matrix is transformed into a quasi-upper triangular form, and the elements on the diagonal are the eigenvalues.
[0104] Further introducing curvature k and direction continuity Two key screening criteria:
[0105] Curvature k calculation: pixel-based The coordinates of the graph and its 8 neighboring pixels are used to fit a quadratic curve y = ax² + bx + c using the least squares method. The curvature k is calculated as follows:
[0106]
[0107] Where k represents curvature, a, b, and c represent coefficients obtained by fitting the quadratic curve using the least squares method, a affects the size and direction of the curve's opening, b affects the position of the curve's axis of symmetry, and c represents the curve's intercept on the y-axis.
[0108] Set curvature threshold (Dynamically adjusted according to part type, metal parts) Plastic parts ),reserve Pixels, removing interfering pixels in flat areas;
[0109] Directional continuity The calculation is performed by measuring the edge orientation angle θedge of adjacent pixels (within ≤ 3 pixels), ensuring directional continuity. The specific calculation method is as follows:
[0110]
[0111] in, Represented as directional continuity, Represented as edge direction angle;
[0112] Set directional continuity threshold (i.e., the edge direction deviation between adjacent pixels is ≤36°), retain The pixels are selected to ensure that they have a continuous linear distribution and conform to the geometry of the burr.
[0113] Through the above multi-condition screening, the final selection will retain those that simultaneously meet the criteria. , , The pixels are used to generate a flyedge-sensitive feature map F.
[0114] The dual-branch semantic segmentation network module: the first branch takes the fly-edge sensitive feature map F as input, and the second branch takes... The pseudo-color fusion image is taken as input, and the two branches exchange weights bidirectionally in the decoding stage through cross-attention gating units, outputting a pixel-level flash mask M;
[0115] The dual-branch semantic segmentation network module adopts a symmetrical U-shaped architecture, consisting of three parts: a feature encoding layer, a cross-attention gating fusion layer, and a decoding output layer. The inputs of the two branches undergo targeted preprocessing to adapt to feature extraction requirements.
[0116] In the input processing of the first branch (flyedge sensitive feature branch), the flyedge sensitive feature map output by the edge calculation node module is used as input. Normalization processing is performed before input to map pixel values to the [0,1] interval. The specific calculation method is as follows:
[0117]
[0118] in, Represented as coordinates after normalization. Pixel value at; Represented as the original image in coordinates Pixel value at that location, It represents the maximum value among all pixel values in the original image. Represented as the minimum value among all pixel values in the original image;
[0119] Then, a 1×1 convolution kernel is used to increase the channel dimension, providing a multi-channel feature basis for subsequent coding layers, while keeping the feature map resolution unchanged;
[0120] In the second branch (polarization image fusion branch) input processing, the N≥3 frame polarization image sequence acquired by the high-speed polarization camera module is used. The pseudo-color fusion image is used as input; the pseudo-color fusion process uses a multi-channel feature mapping algorithm: first, for each frame... Polarization parameters are calculated, and the polarization degree (DoP) and polarization angle (AoP) features are extracted. Then, DoP, AoP, and the original grayscale values are combined to form a 3-channel feature map. Through HSV color space conversion, DoP is mapped to the Hue channel, AoP to the Saturation channel, and the grayscale values to the Value channel, generating a pseudo-color fusion image. ;
[0121] For each frame Polarization analysis is performed using Stokes vectors to calculate the polarization degree and polarization angle images of the surface; when N=4 ( , , , When ), Stokes vector The specific calculation method is as follows:
[0122]
[0123]
[0124]
[0125] in, The total light intensity component is represented as the Stokes vector. The horizontal-vertical linear polarization components, represented as the Stokes vector, Represented as a Stokes vector - Linear polarization component;
[0126] Note: When N=3, it can be achieved through... , , The light intensity value is obtained by fitting the solution using the least squares method. , , ;
[0127] Based on the Stokes vector, the specific methods for calculating the degree of polarization (DoP) and the polarization angle (AoP) are as follows:
[0128]
[0129]
[0130] The degree of polarization image highlights the birefringence region caused by stress, while the polarization angle image shows the orientation of the surface microstructure; for example, the DoP value of the metal flash region is ≥0.35, while the DoP value of the bulk is ≤0.15; the AoP value of the plastic flash region is offset by ≥5° relative to the bulk. These quantitative characteristics provide a basis for judgment for the flash defect detection module.
[0131] In the feature coding layer, the two branches of the coding layer have the same structure but independent parameters. Each branch contains four downsampling modules, each consisting of two 3×3 convolutional layers, one batch normalization layer, and one 2×2 max pooling layer. Through progressive downsampling, the number of channels is gradually increased from 16 to 256, extracting features from low-level details to high-level semantic features. Specifically, the first branch of the coding layer captures the geometric structural features of the flyedge, such as its length, width, and curvature, and enhances its adaptability to flyedges of different sizes through multi-scale convolutional kernels. The second branch of the coding layer captures polarization characteristics and spectral reflectance features. By adding attention mechanisms to the second and third coding modules, the polarization feature response of the flyedge region is strengthened, and the interference of background noise (such as oil stains and scratches on the surface of the part) is suppressed.
[0132] In the cross-attention gating unit, a cross-attention gating unit is set between the encoding layer and the decoding layer to perform bidirectional weight exchange and fusion of features from the two branches; this unit adopts a two-stage mechanism of "intra-branch self-attention + inter-branch mutual attention", the specific process of which is as follows:
[0133] Self-attention stage: The two branches perform self-attention calculations on the highest-level feature maps output by the encoding layer, using the scaling dot product attention formula. The specific calculation method is as follows:
[0134]
[0135] Where Q represents the query matrix, K represents the key matrix, and V represents the value matrix. Represented as the key dimension.
[0136] During the mutual attention phase, the Q matrix of the first branch is cross-calculated with the K and V matrices of the second branch to obtain the cross-branch attention weights. Simultaneously, the Q matrix of the second branch is cross-calculated with the K and V matrices of the first branch to obtain... The specific method for calculating weight swapping is as follows:
[0137]
[0138]
[0139] in, This is represented as the feature map after the first branch fusion. This is represented as the feature map after the first branch self-attention. This is represented as the attention weight of the first branch on the second branch; This is represented as the feature map after the second branch fusion. This is represented as the feature map after the second branch self-attention. This is represented as the attention weight of the second branch on the first branch;
[0140] The last layer of the two-branch decoding layer reduces the number of channels to 1 using a 1×1 convolution kernel, and then passes it through a sigmoid activation function to output pixel-level graffiti probability values. (First branch) and (Second branch), with a value range of [0,1]; the closer the probability value is to 1, the higher the probability that the pixel is a glitter. The specific calculation method is as follows:
[0141]
[0142]
[0143] in, Represented as the image coordinates output by the first branch The probability value of a pixel at a given location belonging to a glimpse. Represented as the output of the second branch in image coordinates The probability value of a pixel at a given location belonging to a glimpse. Represented as the Sigmoid function, This is represented as the decoding function of the first branch. This is represented as the decoding function for the second branch;
[0144] Using a weighted fusion strategy and To perform fusion, weighting coefficients , Determined through cross-validation on the training set (dynamically adjusted based on the type of flash, such as metallic flash). , Plastic flash , 5) The specific method of fusion calculation is as follows:
[0145]
[0146] in, This is expressed as the predicted probability after final fusion;
[0147] Then, the segmentation threshold was set. ,when If a pixel is identified as a fringe, its mask value is set to 1; otherwise, it is considered background and its mask value is set to 0. This process generates a pixel-level fringe mask, which marks the pixel-level position of the fringe. The specific calculation method is as follows:
[0148]
[0149] in, Represented as pixel-level glitter mask, This is represented as the segmentation threshold.
[0150] The physical constraint post-processing module projects pixel-level flash masks onto the coordinate system of the part's CAD model, utilizing the tolerance band T and the flash height threshold. Perform 3D geometric verification while retaining the 2D mask area requirement. And three-dimensional height The final defect report is obtained by analyzing examples of burrs.
[0151] In the physical constraint post-processing module, firstly, the module establishes a mapping relationship between the pixel coordinate system and the part CAD model coordinate system, transforming the two-dimensional flash mask M into a defect candidate region in three-dimensional space; this mapping process is based on system-preset calibration parameters, specifically:
[0152] Camera intrinsic and extrinsic parameter retrieval: Retrieving the intrinsic parameter matrix E (including focal length) of the high-speed polarization camera module. , Principal point coordinates , ) and extrinsic parameter matrix (rotation matrix G, translation vector j), where the extrinsic parameter matrix is obtained in advance through hand-eye calibration, describing the pose relationship between the camera coordinate system and the coordinate system of the part CAD model;
[0153] Pixel coordinates to 3D point cloud conversion: For pixels marked as flash in the flash mask The elevation map generated by the edge computing node module is converted into three-dimensional coordinates in the camera coordinate system using the camera imaging model. The calculation method is as follows:
[0154]
[0155] The coordinates in the camera coordinate system were then converted to three-dimensional coordinates in the part's CAD model coordinate system using an extrinsic parameter matrix. The calculation method is as follows:
[0156]
[0157] The final 3D point cloud of the candidate flash edge region is obtained, and the specific calculation method is as follows:
[0158]
[0159] in, Represented as a 3D point cloud, u represents the total number of fringe pixels in the mask;
[0160] Load the CAD model of the part and extract the surfaces in the model corresponding to the candidate areas for flash. The 3D point cloud is matched with the surface using the ICP algorithm;
[0161] Then, the candidate regions for fly edges are validated through dual physical constraints to eliminate minor interferences:
[0162] In the calculation and selection of 2D mask area, the 2D area V corresponding to the flash pixels in the flash mask is first calculated. The calculation takes into account the image pixel resolution. The specific calculation method is as follows:
[0163]
[0164] Where V represents the area of the two-dimensional mask, Represented as the pixel size in the horizontal direction of the camera. The pixel size is represented in the vertical direction of the camera, and m represents the number of glitter pixels.
[0165] Preset area threshold The tolerances are set according to the part design tolerances and manufacturing process requirements, i.e., precision gear parts. Ordinary plastic casing ;when If it is determined to be a minor interference, the candidate region is directly eliminated; when At that time, proceed to three-dimensional height verification;
[0166] Based on CAD model surface With Fly Edge Candidate Point Cloud Calculate each flash point relative to the surface The normal distance is the height of the flash at that point, and the calculation method is as follows:
[0167]
[0168] in, This is expressed as the flash height. Represented as in a 3D point cloud On the surface The closest point on, Represented as surface The normal vector at the nearest point;
[0169] Next, calculate the average height of the burr edge area and use the average height as the height of the candidate burr edge area;
[0170] Preset height threshold The threshold is set to ;when When the dimensional deviation is determined to be within the acceptable range, the candidate area is eliminated; when If the candidate region is confirmed to be a real flash defect, it is retained as a flash instance.
[0171] Read the design tolerance data of the currently inspected part from the MES system and automatically match the corresponding... For multi-batch, multi-model parts inspection scenarios, a built-in tolerance zone database is provided, which automatically retrieves the preset T value based on the part model (e.g., T=0.08mm for metal shaft parts, T=0.15mm for plastic connectors), where T represents the tolerance zone.
[0172] After three-dimensional geometric verification, the defect information of the retained flash edge instances is quantified and a report is generated. The final output defect report includes basic defect information, defect quantification parameters, and visualization attachments.
[0173] Basic defect information includes:
[0174] Part information: Part model, batch number, inspection time, inspection station number;
[0175] Flash number: Multiple flash instances of the same part are numbered according to the inspection sequence;
[0176] Location information: Center coordinates of the flash in the CAD coordinate system , corresponding part surfaces;
[0177] Defect quantification parameters:
[0178] Two-dimensional parameters: flash mask area, maximum two-dimensional length of flash (the longest distance along the surface of the part), and maximum width (the maximum distance perpendicular to the length direction).
[0179] Three-dimensional parameters: average height of flash, maximum height (maximum normal distance within the flash area), and standard deviation of height;
[0180] Defect level: based on Divide by the ratio to T, when At this point, it is considered a minor defect; when At this point, it is considered a moderate defect; when This is a serious defect;
[0181] Visual attachments:
[0182] Two-dimensional image: original image of the part with the flash mask M superimposed, and local magnification of the flash area;
[0183] 3D Model: Mark the location of flash instances in the part's CAD model, and use different colors to distinguish the defect level (green: minor, yellow: moderate, red: severe).
[0184] Data charts: Histogram of flash height distribution, bar chart of defect quantity statistics for parts in the same batch.
[0185] After a defect report is generated, it interacts with the MES system and quality management system in real time via the OPC UA protocol, automatically entering the defect information into the production database and triggering corresponding quality warnings.
[0186] The technical features of the above embodiments can be combined in any way (as long as there is no contradiction in the combination of these technical features). For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described; these embodiments not explicitly written should also be considered to be within the scope of this specification.
Claims
1. An image recognition system for detecting defects in industrial parts, characterized in that, include: Multi-axis controllable light source array module: configured to provide time-division strobe illumination of the part surface with different combinations of incident angle intensities within the same sampling period; High-speed polarization camera module: The polarizer angle θ is synchronously rotated to a preset angle corresponding to the current incident angle intensity combination at each stroboscopic moment. Capture N≥3 frames of polarization image sequence J ; Edge computing node module: This module embeds a fly-edge sensitive feature extraction unit. The module obtains the fly-edge sensitive feature map F through the following steps: For preset angle Perform photometric stereo calculations to generate an elevation map Z and a surface normal map. ; Connect Z with Input a differentiable rendering layer and reconstruct an ideal contour map without flash. ; Compare the actual contour map C with By performing a difference operation, an initial defect map is obtained. ; right Perform ridge response filtering based on the Hessian matrix, preserving curvature κ > κth and directional continuity. > The pixels are used to obtain the flyedge sensitive feature map F; where, Represented as curvature threshold, This is represented as the directional continuity threshold; Dual-branch semantic segmentation network module: The dual-branch semantic segmentation network module adopts a symmetrical U-shaped architecture, comprising three parts: a feature encoding layer, a cross-attention gating fusion layer, and a decoding output layer; the first branch takes the fly-edge sensitive feature map F as input, and the second branch takes... The pseudo-color fusion image is used as input, and the two branches exchange weights bidirectionally in the decoding stage through the attention gating unit, outputting a pixel-level flyedge mask M; The dual-branch semantic segmentation network module adopts a symmetrical U-shaped architecture, which includes three parts: a feature encoding layer, a cross-attention gating fusion layer, and a decoding output layer. In the first branch input processing, the fly-edge sensitive feature map output by the edge computing node module is used as input. Normalization processing is performed before input to map the pixel values to the [0,1] interval. In the second branch input processing, the N≥3 frame polarization image sequence acquired by the high-speed polarization camera module is used. The pseudo-color fusion image is used as input; the pseudo-color fusion process uses a multi-channel feature mapping algorithm: first, for each frame... Polarization parameters are calculated, and the polarization degree (DoP) and polarization angle (AoP) features are extracted. Then, DoP, AoP, and the original grayscale values are combined to form a 3-channel feature map. Through HSV color space conversion, DoP is mapped to the Hue channel, AoP to the Saturation channel, and the grayscale values to the Value channel, generating a pseudo-color fusion image. ; For each frame Polarization analysis is performed using Stokes vectors to calculate the polarization degree and polarization angle images of the surface. In the feature coding layer, the two branches of the coding layer have the same structure, each containing 4 downsampling modules. Each module consists of 2 3×3 convolutional layers, 1 batch normalization layer and 1 2×2 max pooling layer. Through progressive downsampling, the number of channels is gradually increased from 16 channels to 256 channels, from low-level detailed features to high-level semantic features. In the cross-attention gating unit, a cross-attention gating unit is set between the encoding layer and the decoding layer to perform bidirectional weight exchange and fusion of features from the two branches; this unit adopts a two-stage mechanism of intra-branch self-attention and inter-branch mutual attention, the specific process of which is as follows: Self-attention stage: The two branches perform self-attention calculations on the highest-level feature maps output by the coding layer, using the scaling dot product attention formula; During the mutual attention phase, the Q matrix of the first branch is cross-calculated with the K and V matrices of the second branch to obtain the cross-branch attention weights. Simultaneously, the Q matrix of the second branch is cross-calculated with the K and V matrices of the first branch to obtain... ; The last layer of the two-branch decoding layer reduces the number of channels to 1 using a 1×1 convolution kernel, and then passes it through a sigmoid activation function to output pixel-level graffiti probability values. and Then, a weighted fusion strategy is used to... and To merge; Set the segmentation threshold ,when When a pixel is identified as a fringe, its mask value is set to 1; otherwise, it is considered background and its mask value is set to 0. This process generates a pixel-level fringe mask, which marks the pixel-level position of the fringe. This is expressed as the predicted probability after final fusion; The physical constraint post-processing module projects pixel-level flash masks onto the coordinate system of the part's CAD model, utilizing the tolerance band T and the flash height threshold. Perform 3D geometric verification while retaining the 2D mask area requirement. And three-dimensional height The final defect report is obtained by analyzing examples of burrs. This is represented as a preset area threshold.
2. The image recognition system for detecting defects in industrial parts according to claim 1, characterized in that: In the multi-axis controllable light source array module, the industrial camera of the image acquisition unit is synchronously linked with the timing control unit. Within a single sampling period, the timing control unit drives each light source unit to perform time-division strobe sequentially with different incident angle intensities according to a preset lighting sequence, thereby acquiring multiple sets of differentiated lighting images. At the same time, each light source unit is equipped with LEDs of different wavelengths, selecting wavelengths covering the visible light range of 380nm-780nm. Through time-division illumination of different spectra, the material differences between the flash and the part body are distinguished.
3. The image recognition system for detecting defects in industrial parts according to claim 1, characterized in that: In the high-speed polarization camera module, the high-speed polarization camera module and the multi-axis controllable light source array module work together via a time synchronization signal; the polarizer at the front end of the camera is driven by a servo motor rotary table; during the time when the incident angle intensity combined light source is triggered to flicker, the camera control system receives a synchronization pulse signal. Based on the synchronization pulse signal, the polarizer rotates to a preset angle pre-bound to the lighting conditions with a millisecond-level response speed; in this way, for the same displacement of the part surface, within the same complete composite lighting sampling period, the camera captures a set of polarization image sequences J containing N≥3 frames. Each frame of the image contains information about the surface topography and spectral reflectance under illumination from different spectra and angles; this complete sequence of polarized images... The polarization degree image and polarization angle image of the surface are calculated using the Stokes vector polarization analysis algorithm.
4. The image recognition system for detecting defects in industrial parts according to claim 1, characterized in that: The edge computing node module receives a polarization image sequence containing N≥3 frames. Based on photometric stereo vision combined with known illumination parameters of a multi-axis controllable light source array, photometric stereo solution is performed on the image sequence. Using the Lambertian reflection model as a basis, the reflected light intensity satisfies: in, The polarizer rotation angle of the i-th frame is represented as... The light intensity value of the image acquired at that time. Represented as the surface reflectivity of the part. Let be the direction vector of the light source for the i-th illumination. Let G be the surface normal vector, and let G be the ambient light intensity. By establishing an overdetermined system of equations for N frames of images, the surface normal vector of each pixel is solved using the least squares method. (x,y) yields the surface normal map. Subsequently, based on the surface normal map Elevation map Z is generated by performing integral operations, and the Poisson reconstruction algorithm is used in the integration process. The generated elevation map and surface normal map are input into a differentiable rendering layer, which is based on a physical rendering engine and has a built-in CAD design model of the part. The ideal contour map without flash is reconstructed through reverse rendering.
5. The image recognition system for detecting defects in industrial parts according to claim 4, characterized in that: After obtaining the ideal profile map K without flash, the polarization image sequence acquired by the high-speed polarization camera module is first processed. Edge detection is performed to generate an actual contour map C. An improved Canny algorithm is used for edge detection, extracting the actual edges of the part through dynamic thresholding, and then performing morphological closing operations to eliminate edge breaks, resulting in a continuous and complete actual contour map C. Subsequently, the actual contour map C is compared with the ideal contour map. Perform pixel-by-pixel difference operation: Denoising the initial defect image Perform ridge response filtering based on the Hessian matrix to filter out pixels with flash defects; calculate In each pixel The Hessian matrix H at that location, in, It is represented as the Hessian matrix at pixel (x,y). Expressed as the second partial derivative with respect to the x-coordinate, It can be expressed as a mixed second-order partial derivative obtained by first taking the partial derivative with respect to the x-coordinate and then taking the partial derivative with respect to the y-coordinate; It is expressed as the second-order partial derivative with respect to the y-coordinate; By solving for the eigenvalues of the Hessian matrix, the ridge response value is calculated, and further, the curvature k and directional continuity are introduced. Two key screening criteria were used to select the final batch of products that simultaneously met the above criteria. , , The pixels are used to generate a flyedge-sensitive feature map F; R represents the ridge response value. This represents the ridge response threshold; k represents the curvature. Represented as curvature threshold; Represented as directional continuity, This is represented as the directional continuity threshold.
6. The image recognition system for detecting defects in industrial parts according to claim 1, characterized in that: In the physical constraint post-processing module, firstly, the module establishes a mapping relationship between the pixel coordinate system and the coordinate system of the part's CAD model, transforming the pixel-level flash mask M into a defect candidate region in three-dimensional space; the mapping process from the two-dimensional mask to the three-dimensional defect candidate region is based on system-preset calibration parameters, specifically: Camera intrinsic and extrinsic parameter retrieval: Retrieving the intrinsic parameter matrix E and extrinsic parameter matrix of the high-speed polarization camera module. The extrinsic parameter matrix is obtained in advance through hand-eye calibration and describes the pose relationship between the camera coordinate system and the coordinate system of the part's CAD model. Pixel coordinates to 3D point cloud conversion: For pixels marked as flash in the flash mask The elevation map generated by the edge computing node module is converted into 3D coordinates in the camera coordinate system using the camera imaging model. Then, the coordinates in the camera coordinate system are converted into 3D coordinates in the part's CAD model coordinate system using an extrinsic parameter matrix. Finally, the 3D point cloud of the flash edge candidate region is obtained. The specific calculation method is as follows: in, Represented as a 3D point cloud, u represents the total number of fringe pixels in the mask; Load the CAD model of the part and extract the surfaces in the model corresponding to the candidate areas for flash. The 3D point cloud is matched with the surface using the ICP algorithm.
7. The image recognition system for detecting defects in industrial parts according to claim 6, characterized in that: In the calculation and selection of 2D mask area, the 2D area V corresponding to the flash pixels in the flash mask is first calculated, taking into account the image pixel resolution; based on the CAD model surface... With Fly Edge Candidate Point Cloud Calculate each flash point relative to the surface The normal distance is the height of the flash at that point, and the calculation method is as follows: in, This is expressed as the flash height. Represented as in a 3D point cloud On the surface The closest point on, Represented as surface The normal vector at the nearest point; Next, the average height of the burr area is calculated and used as the height of the candidate burr area. After three-dimensional geometric verification, the defect information of the retained burr instances is quantified and a report is generated. The final output defect report includes basic defect information, defect quantification parameters, and visualization attachments. After the defect report is generated, it interacts with the MES system and quality management system in real time through the OPC UA protocol, automatically enters the defect information into the production database, and triggers the corresponding quality warning.
Citation Information
Patent Citations
Surface quality control device for three-dimensional (3D) part formed through metal drop printing and control method of surface quality control device
CN105081325A
Part surface defect detection and process optimization method and system
CN120747065A