Die casting burr detection system and method based on machine vision
By generating a polarization-spectral image set and performing principal component analysis and surface adaptive optical compensation, combined with the U-Net and Mask-R-CNN models, the adaptability problem of complex surface inspection of die-casting parts was solved, high-resolution imaging and stable burr positioning were achieved, and the inspection accuracy and degree of automation were improved.
Patent Information
- Application Number
- CN202510809895.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-10-17
AI Technical Summary
Existing die-casting burr detection technology based on machine vision is not adaptable enough to complex surfaces. The deep learning model does not fully utilize multispectral data and geometric information, making it difficult to achieve stable burr positioning and size measurement under complex lighting conditions.
By collecting multispectral images of die-cast parts, generating polarization-spectral image sets and performing principal component analysis, the optical path is optimized by combining the surface adaptive optics compensation mechanism to generate high-resolution image sets. The U-Net and Mask-R-CNN models are then used for semantic segmentation and detection, shadows are eliminated, and burr detection results are output.
It improves the imaging resolution and clarity of complex surfaces and highly reflective areas, overcomes the deep learning model's dependence on data sets, achieves stable burr positioning and precise size measurement under complex lighting conditions, and provides an efficient and automated detection method.
Smart Images

Figure CN120807406A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine vision detection, and in particular to a system and method for detecting burrs on die castings based on machine vision. Background Art
[0002] Burr detection on the surface of die-cast parts is a key step in quality control in the manufacturing industry. In recent years, automated detection technologies based on machine vision have gradually emerged. High-resolution cameras combined with image processing algorithms can achieve non-contact burr detection. For example, edge detection algorithms and template matching methods based on single-spectrum imaging have been widely used in surface defect identification. In addition, multispectral imaging technology can distinguish the reflective characteristics of metal substrates, oxide layers, and burrs by capturing spectral information in different bands, significantly improving detection accuracy. Deep learning models, such as convolutional neural network models, have further promoted the application of semantic segmentation and target detection, and can achieve burr location and segmentation in complex scenarios by training labeled datasets. Some advanced technologies introduce three-dimensional point cloud modeling and curvature analysis to address the geometric challenges posed by the complex curved surfaces of die-cast parts.
[0003] Although machine vision-based automated inspection technology has made a lot of progress, there is still room for improvement. First, the existing methods are not adaptable enough to the complex surfaces of die-castings and lack adaptive optical path optimization for the surface geometry, resulting in defocus and distortion in the deep cavity area, affecting the clarity of the burr edge. In addition, although the existing deep learning model can achieve high-precision segmentation, it is highly dependent on the training data set and does not fully utilize multispectral data and geometric information, making it difficult to achieve stable burr positioning and size measurement under complex lighting conditions. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a die-casting burr detection method based on machine vision to solve the problems of insufficient adaptability to complex surfaces of die-castings and insufficient utilization of multispectral data and geometric information by deep learning models.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for detecting burrs on die castings based on machine vision, which comprises: Collect multispectral images of die castings to generate a polarization-spectral image set, and perform principal component analysis on the polarization-spectral image set to obtain a spectral feature vector set; Based on the spectral feature vector set, the surface adaptive optics compensation mechanism is used to reacquire the polarization-spectral image set through optical path optimization, obtain the optimized polarization-spectral image set, and perform correction to generate a high-resolution image set. The surface of the die casting is analyzed for light distribution to form a multi-exposure image set, and the high dynamic range image is generated by inputting the U-Net model into the constructed U-Net model for semantic segmentation. The shadow elimination of the high dynamic range image is performed by using the gradient Poisson fusion algorithm, and the pre-trained Mask-R-CNN model is used for detection after the shadow elimination, and the detection result of the burr is output.
[0007] As a preferred scheme of the die casting burr detection method based on machine vision, the spectral feature vector set is obtained, specifically including the following steps, The multi-spectral and polarization information of the surface of the die casting is captured by the hyperspectral camera and the polarization filter array, and the polarization-spectral image set is formed after preprocessing and segmentation; The principal component analysis is applied to each frame of the polarization-spectral image set, the first N principal components are reserved by singular value decomposition, the dimensionality is reduced to an N-dimensional feature matrix, and the spectral feature vector set is generated based on the N-dimensional feature matrix.
[0008] As a preferred scheme of the die casting burr detection method based on machine vision, the spectral feature vector set is obtained, specifically including the following steps, The 3D curvature sensor is used to scan the surface of the die casting by laser triangulation to generate a high-precision three-dimensional point cloud model; The curvature distribution map is generated based on the spectral feature vector set using the Gaussian curvature formula; The geometric characteristics of the surface of the die casting are analyzed by the three-dimensional point cloud model and the curvature distribution map, the high-reflectivity area is identified in combination with the spectral feature vector set, the light path parameters are calculated and the imaging light path is dynamically adjusted, the polarization-spectral image is reacquired using the hyperspectral camera, and the optimized polarization-spectral image set is generated.
[0009] As a preferred scheme of the die casting burr detection method based on machine vision, the high-resolution image set is generated by performing geometric correction and clarity verification on the optimized polarization-spectral image set, and eliminating distortion.
[0010] As a preferred scheme of the die casting burr detection method based on machine vision, the multi-exposure image set is formed, specifically including the following steps, The light distribution of the surface of the die casting is analyzed by using a photosensitive sensor to generate a light distribution map; Based on the region classification of the light distribution map, the image acquisition with different exposure times is set to form a multi-exposure image set.
[0011] As a preferred scheme of the die casting burr detection method based on machine vision provided in the application, the generating a high dynamic range image specifically comprises the following steps, The U-Net model is trained based on the die casting image dataset, and the construction of the U-Net model is completed by optimizing the average intersection-over-union ratio to reach a specified requirement. The multi-exposure image set is input into the constructed U-Net model, and the U-Net model outputs a segmentation mask through the encoder and the decoder. The Debevec algorithm is applied based on the segmentation mask to synthesize the images in the multi-exposure image set, and a high dynamic range image is generated.
[0012] As a preferred scheme of the die casting burr detection method based on machine vision provided in the application, the generating a high dynamic range image specifically comprises the following steps, The high dynamic range image and the segmentation mask are fused by gradient domain processing to generate a high dynamic range image without shadows. The Mask-R-CNN model is trained using the die casting burr image dataset, and the die casting burr image dataset is divided into a training set and a validation set. The accuracy of the classification task is measured by the cross-entropy loss, the accuracy of the bounding box regression and the mask prediction is measured by the L1 loss, and the precision of the Mask-R-CNN model is evaluated using the validation set to reach a specified requirement, and the training of the Mask-R-CNN model is completed. The Mask-R-CNN model comprises a backbone network, a region proposal network, a detection head and a mask branch. The high dynamic range image without shadows is input into the trained Mask-R-CNN model for semantic segmentation and instance detection, and the bounding box, size and binary mask of the burr are output.
[0013] In a second aspect, the application provides a die casting burr detection system based on machine vision, comprising, The acquisition module acquires a multi-spectral image of the die casting, generates a polarimetric-spectral image set, and performs principal component analysis on the polarimetric-spectral image set to obtain a spectral feature vector set. The calibration module uses a curved surface adaptive optical compensation mechanism to reacquire the polarimetric-spectral image set through optical path optimization based on the spectral feature vector set, obtains an optimized polarimetric-spectral image set, and performs correction to generate a high-resolution image set. The segmentation module performs illumination distribution analysis on the surface of the die casting through the high-resolution image set to form a multi-exposure image set, and inputs the multi-exposure image set into the constructed U-Net model for semantic segmentation to generate a high dynamic range image. The detection module applies a gradient Poisson fusion algorithm to the high dynamic range image for shadow elimination, and uses a pre-trained Mask-R-CNN model for detection after shadow elimination, and outputs the detection result of the burr.
[0014] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and wherein the computer program, when executed by the processor, implements any step of the method for detecting burrs of a die casting based on machine vision according to the first aspect of the present application.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any step of the method for detecting burrs of a die casting based on machine vision according to the first aspect of the present application.
[0016] The present application has the following beneficial effects: the polarized-spectral image set is optimized through curved surface adaptive optics, the imaging optimization and geometric correction of complex curved surface regions are realized, a high-resolution image set is generated, the imaging quality of the polarized-spectral image set is improved, and through adaptive light path adjustment and geometric correction, the specular reflection and defocus distortion of high-reflective regions are reduced, the details of burr regions are enhanced, the imaging resolution and clarity of complex curved surfaces and high-reflective regions are significantly improved, the limitations of traditional fixed light path imaging on curved surfaces are overcome, in addition, through full use of the spectral feature vector set and geometric information, combined with high-precision segmentation and detection of the U-Net and Mask-R-CNN models, the dependence of the deep learning model on the data set is overcome, stable burr positioning and accurate size measurement under complex lighting conditions are realized, and an efficient and automated detection means for die casting quality control is provided. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0018] Fig. 1 The flowchart of the method for detecting burrs of a die casting based on machine vision.
[0019] Fig. 2 The schematic diagram of the system for detecting burrs of a die casting based on machine vision.
[0020] Fig. 3 The flowchart of the curved surface adaptive optical compensation mechanism.
[0021] Fig. 4A flowchart of the processing procedure for the U-Net model and the Mask-R-CNN model. DETAILED DESCRIPTION
[0022] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0023] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, so the present application is not limited to the specific embodiments disclosed below.
[0024] Secondly, "one embodiment" or "embodiment" referred to herein means that a specific feature, structure or characteristic can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an embodiment that is independent of or mutually exclusive with other embodiments.
[0025] Reference Figs. 1-4 For one embodiment of the present application, the embodiment provides a die casting burr detection method based on machine vision, comprising the following steps: S1, collecting die casting multispectral images to generate a polarized-spectral image set, and performing principal component analysis on the polarized-spectral image set to obtain a spectral feature vector set; Specifically, the following steps are included, S1.1, the die casting is placed on the conveying belt of the automatic detection platform, the automatic detection platform is equipped with a hyperspectral camera and an electrically controlled rotating polarized filter array, a wide spectrum light source uniformly irradiates the surface of the die casting at an angle of 45°, and the mirror reflection and diffuse reflection characteristics of the high-reflectivity metal are ensured. After the detection starts, the servo motor drives the electrically controlled rotating polarized filter array to rotate from 0° to 180°, with a step unit of 5°, 36 frames of hyperspectral images are captured in turn, each frame of image contains spectral data of 10 wavebands, and each waveband contains pixels with a resolution of 1024x1024. The photosensitive sensor monitors the stability of the light source in real time, and adjusts the light source power through the feedback control circuit to maintain uniform illumination. The 36 frames of captured hyperspectral images form a three-dimensional data cube, which contains the complete reflection information of the die casting under different polarization angles and spectral wavebands, the first dimension of the three-dimensional data cube represents the polarization angle, the second dimension represents the spectral waveband, and the third dimension represents the pixels of each frame of image, thereby forming a three-dimensional data cube.
[0026] S1.2, according to the three-dimensional data cube, 36 frames of hyperspectral images are processed frame by frame for denoising and enhancement to improve image quality, specifically: Gaussian filtering is applied to each frame of hyperspectral image to remove random noise and improve image detail retention rate, then adaptive histogram equalization is used to enhance the image contrast of each frame of hyperspectral image, highlighting the edge and texture features of the burr area, the 36 frames of hyperspectral images after denoising and enhancement retain the original spectral information and improve the visual saliency of the burr area, generating a set of denoised and enhanced hyperspectral images.
[0027] For the set of hyperspectral images after denoising and enhancement, global threshold segmentation is applied to each frame of hyperspectral image to separate the die casting foreground and background, specifically: Otsu algorithm is used to automatically calculate the best threshold value of each frame of image by maximizing the inter-class variance, generating a binary foreground mask, pixel value 0 represents background and pixel value 1 represents foreground, the threshold value is calculated for each waveband in the segmentation process to ensure that the foreground mask accuracy of each waveband reaches the pixel level, by applying the foreground mask to each frame of the 36 frames of hyperspectral images, the pixel value of the die casting foreground area is retained, generating a set of 36 frames of polarization-spectral images, in which the background pixels are set to zero and the foreground area retains the denoised and enhanced spectral information.
[0028] S1.3, based on the set of 36 frames of polarization-spectral images, the pixel values of the 10 wavebands of each frame of image are matrix constructed to prepare for principal component analysis, specifically: for each frame of image, the pixel data of the 10 wavebands is flattened to construct a pixel spectral data matrix. The pixel spectral data matrix is centrally processed to calculate the mean value of each waveband, and each element is subtracted from the mean value of the corresponding waveband to generate a centralized pixel spectral data matrix, each element refers to the spectral intensity value of the pixel point in the corresponding waveband in the pixel spectral data matrix, that is, the gray value, ranging from 0 to 255, the spectral features of each frame of image are retained in the generated centralized pixel spectral data matrix.
[0029] According to the centralized pixel spectral data matrix, principal component analysis (PCA) is applied to the 10 waveband data of each frame of image to extract the main spectral features, and the principal component analysis retains the first five principal components by singular value decomposition of the pixel spectral data matrix. Through principal component projection operation, each waveband data is reduced in dimension, for example, 10 waveband data is reduced to a 5-dimensional feature matrix. The average value of each pixel point of each frame of image is taken to generate a 5-dimensional spectral feature vector, which is extended to about 100 dimensions by linear interpolation, obtaining a spectral feature vector of about 100 dimensions for each frame.
[0030] A pre-constructed spectral feature database is provided, which includes spectral response data of burrs, oxide layers and substrates of aluminum, zinc alloy and other die casting materials, and is generated based on laboratory tests. The 100-dimensional spectral feature vector of each frame is compared with the spectral feature database to obtain a spectral feature vector set. For example, for each 100-dimensional spectral feature vector, the Euclidean distance between each 100-dimensional spectral feature vector and the spectral response data of the burrs, oxide layers and substrates of aluminum, zinc alloy and other materials in the pre-constructed spectral feature database is calculated. If the Euclidean distance is less than the Euclidean distance threshold value 0.05, the corresponding spectral feature vector is preliminarily classified as a burr, an oxide layer or a substrate region. The Euclidean distance threshold value is determined based on the statistical distribution characteristics of the spectral response data of the burrs, oxide layers and substrates of aluminum, zinc alloy and other die casting materials in the pre-constructed spectral feature database.
[0031] Each spectral feature vector is normalized and outliers are removed based on the 3σ criterion. The 3σ criterion refers to vectors that are more than 3 times the standard deviation. The 36 spectral feature vectors after normalization and outlier removal form a spectral feature vector set, which includes the spectral features of burrs, oxide layers and substrates.
[0032] Further, 36 hyperspectral images are captured by the hyperspectral camera and the electrically controlled rotating polarization filter array to form a three-dimensional data cube. This multi-angle and wide-band acquisition method effectively captures the differentiated spectral and polarization features of the burrs, oxide layers and substrates of the die casting, providing high-quality raw data for subsequent processing and significantly improving the comprehensiveness and accuracy of the detection.
[0033] S2: Based on the spectral feature vector set, a curved adaptive optical compensation mechanism is used to reacquire the polarization-spectral image set through optical path optimization, obtain the optimized polarization-spectral image set and perform correction to generate a high-resolution image set; Specifically, the following steps are included, S2.1: Scanning the surface of the die casting by laser triangulation using a 3D curvature sensor, generating a high-precision three-dimensional point cloud model. Specifically, the laser triangulation emits a laser beam, measures the reflected light angle, calculates the three-dimensional coordinates of the die casting surface, forms a three-dimensional point cloud model, and a set of 36 frames of polarization-spectral images provides pixel data of 10 bands per frame, which is used to calibrate the spatial coordinates of the three-dimensional point cloud model to ensure the correspondence between the point cloud and the image pixels. Apply a curvature analysis algorithm to the three-dimensional point cloud model, calculate the local curvature and normal vector of each point cloud region based on the Gaussian curvature formula, and a set of spectral feature vectors are used to identify high-reflectance regions, which are regions where the reflection peak value in the spectral feature vector is greater than the high-reflectance threshold. The high-reflectance threshold is set according to the statistical analysis of the spectral reflection characteristics of the die casting surface and the detection task requirements, for example, 0.8. Prioritize fine curvature calculation for high-reflectance regions. The process of fine curvature calculation is as follows: In the high-reflectance region, based on the Gaussian curvature formula, a high-order polynomial surface fitting is used instead of standard quadratic fitting to calculate the local curvature of each grid point, and the fitting process is optimized by the least squares method to reduce the influence of point cloud noise. Grid processing is performed on the local curvature and normal vector data, mapping the three-dimensional coordinates of the three-dimensional point cloud model to a curvature distribution map consistent with the point cloud resolution. Each grid point records the curvature value and the normal vector calculated by the Gaussian curvature of the corresponding region. The curvature value is obtained by weighted average of the Gaussian curvature of the grid points. The curvature distribution map integrates the fine curvature data of the high-reflectance region, ensuring that the geometric features of the high-reflectance region are accurately described, and retaining the curvature and normal vector information of all other regions. It provides an accurate geometric basis and ensures the targeting of imaging optimization.
[0034] Further explanation, the generated curvature distribution map integrates the fine curvature data of the high-reflectance region, accurately describes the geometric features of all regions, retains the curvature and normal vector information, provides an accurate geometric basis, and ensures the targeting of imaging optimization.
[0035] S2.2: According to the curvature distribution map, the set of 36 frames of polarization-spectral images and the set of spectral feature vectors, calculate the optical path parameters using a geometric optics model. Specifically, the normal vector of each region provided by the curvature distribution map is used to determine the incident direction of the light, combined with the pixel spectral data of the set of 36 frames of polarization-spectral images, the best incident light angle of each region is calculated based on Snell's law formula, which is expressed as: ; wherein, and both represent the refractive index of the medium, represents the incident angle, represents the refraction angle.
[0036] The range of the optimal incident light angle of each region is 0° to 90°, the incident light angle is adjusted preferentially for the high-reflective region to reduce specular reflection and reduce the reflection intensity, and meanwhile, the focal point position of each region is calculated according to the curvature and normal vector of the curvature distribution map, the imaging plane is ensured to be aligned with the curved surface region, the light path parameter set of each frame of image is generated, and the generated light path parameter set of each frame of image includes the optimal incident light angle and the focal point position of each region.
[0037] Through the light path parameter set of each frame of image, the light path is adjusted using the electrically controlled liquid lens and the micro mirror array, and the specific process is as follows: the electrically controlled liquid lens dynamically adjusts the focal length by adjusting the voltage to accurately focus the imaging plane on the curved surface region and eliminate defocus distortion. The micro mirror array adjusts the angle of each micro mirror according to the light path parameter set to optimize the incident light path and reduce the influence of curved surface reflection and distortion. Based on the adjusted light path, 36 frames of polarized-spectral images are reacquired using the hyperspectral camera to ensure that the imaging resolution of the curved surface region is improved, and an optimized polarized-spectral image set is generated.
[0038] It is further illustrated that the generated optimized polarized-spectral image set improves the imaging resolution of the curved surface region, significantly improves the imaging quality of the complex curved surface region, reduces the specular reflection and defocus problem of the high-reflective region, and provides clearer image data for burr detection.
[0039] Based on the optimized polarized-spectral image set, a perspective transformation algorithm is applied to each frame of image for geometric correction, and the specific process is as follows: the perspective transformation algorithm calculates the geometric transformation matrix of each frame of image through the pre-calibrated camera intrinsic parameters and the curvature distribution map, eliminates the residual distortion caused by lens distortion and curved surface, and preferentially performs geometric correction on the high-reflective region in combination with the spectral feature vector set. Specifically, based on the 100-dimensional features of the spectral feature vector set, the high-reflective region is identified, and a binary mask is generated to mark the high-reflective region pixels; for the high-reflective region, high-curvature points (for example, curvature > 0.1 mm⁻ 1 ) are extracted from the curvature distribution map as control points, pixel weights are distributed in combination with the reflection intensity distribution of the spectral feature vector set, for each sub-grid, a local transformation matrix is generated based on the control points and the camera intrinsic parameters by least squares method, and the local transformation matrix is optimized using the random sample consensus algorithm; the high-reflective region is divided into 32x32 pixel sub-grids, the local transformation matrix is calculated for each grid, and the global transformation matrix is formed by splicing; the pixel coordinates are corrected using the geometric transformation matrix and the local transformation matrix, the pixel values are resampled using bilinear interpolation, the spectral intensity of 10 bands is preserved, the pixel alignment of the high-reflective region is optimized, the residual distortion is reduced, and the geometric corrected polarized-spectral image set is obtained. According to the geometric corrected polarized-spectral image set, the sharpness is verified using the Laplacian variance algorithm, and the verification process is as follows: the sharpness of each frame of image is calculated using the Laplacian variance algorithm with a 3x3 convolution kernel, and the formula is as follows: ; wherein, represents the Laplacian variance value, which is a quantitative indicator of image sharpness, the larger the value, the clearer the image edge and detail, represents the rate of change of the image gray value detected by the Laplacian operator, represents the gray value of the image, ranging from [0, 255].
[0040] Based on the sharpness requirement of die casting surface burr detection, the target Laplacian variance threshold is set through image quality evaluation, for example, the target Laplacian variance threshold is 100, if the Laplacian variance value of a certain frame of image is lower than 100, the feedback mechanism is triggered, the electric control liquid lens focal length or the micro mirror array angle is adjusted, and the frame of image is reacquired, if it is higher than 100, the sharpness is satisfied, then the corrected high resolution image set is generated, which has high definition and low distortion characteristics, and retains the spectral characteristics of burrs, oxide layers and substrates.
[0041] Further explanation, the corrected high resolution image set finally generated significantly improves the imaging resolution and sharpness of complex curved surface area through curved surface adaptive optical compensation and geometric correction, provides high quality input for dynamic HDR synthesis and shadow elimination, can effectively support detail extraction in high light area and deep cavity area, and enhances the precision and accuracy of burr detection.
[0042] S3: Through the high resolution image set, light distribution analysis is performed on the die casting surface to form a multi-exposure image set, and is input into the constructed U-Net model for semantic segmentation to generate a high dynamic range image; Specifically, the following steps are included, S3.1: Receive the corrected high resolution image set, use the photosensitive sensor to scan the die casting surface to generate a light distribution map, specifically: use the photosensitive sensor to measure the light intensity of the die casting surface, based on the light intensity, identify the highlight area, shadow area and deep cavity area, generate a light distribution map, the resolution is consistent with the image, the corrected high resolution image set provides 10 waveband pixel spectral data for each frame, which is used to calibrate the spatial coordinates of the light distribution map, to ensure that the light intensity corresponds to the image pixel, the highlight area and the shadow area of the light distribution map are determined based on the light characteristics of the die casting surface and the photosensitive sensor through light distribution statistical analysis, for example, light intensity > 5000 lux is the highlight area, light intensity < 100 lux is the low light area. The light distribution map records the light intensity and classification label of each area, and the classification label includes highlight, shadow and deep cavity.
[0043] Further explanation, based on the statistical analysis of the light distribution, the high light area, the shadow area and the deep cavity area are identified, and the light distribution map records the light intensity and classification label of each area. This accurate light analysis provides regionalized light information for subsequent multi-exposure image acquisition, optimizes the detail capture of high light and deep cavity areas, enhances the adaptability of complex light conditions on the surface of the die casting, and lays the foundation for burr detection S3.2: Based on the light distribution map of the high light area, the shadow area and the deep cavity area, set 3 frames of different exposure time, for example, 0.1ms, 1ms and 10ms; 0.1ms optimizes the detail capture of high light area by short exposure, 1ms optimizes the detail capture of plane area by medium exposure, and 10ms optimizes the detail capture of shadow area and deep cavity area by long exposure. The high dynamic range CMOS sensor captures 3 frames of high resolution images in high speed mode, each frame retains 10 wave bands, and generates a multi-exposure image set, which contains the details of high light, shadow and deep cavity area.
[0044] Further explanation, multi-exposure image set covers the details of high light, shadow and deep cavity area, significantly improves the visibility of deep cavity area, and reduces the problem of overexposure in high light area and underexposure in shadow area.
[0045] S3.3: Based on the 10000 die casting image dataset, the U-Net model is trained to generate the U-Net model for semantic segmentation. The die casting image dataset contains die casting images, and the deep cavity, plane and high light area are labeled, with pixel-level labels of 1 (deep cavity), 2 (plane) and 0 (other), covering 10 spectral wave bands. The die casting image dataset is divided into training set and validation set. The U-Net model adopts the coding-decoding structure, learns the image features through convolution and up-sampling layer, the optimization target is to maximize the average intersection over union, the cross entropy loss function is used, the training is carried out for 100 epochs, the batch size is 16, and the learning rate is 0.001. During the training process, the validation set evaluates the segmentation accuracy, and when the maximum average intersection over union is greater than 0.9, it is ensured that the U-Net model can accurately distinguish the deep cavity, plane and high light area, and the training is completed. The pre-trained U-Net model saves the weight parameters.
[0046] S3.4: Input 3 frames of images in multi-exposure image set into pre-trained U-Net model one by one, the encoder of U-Net model contains multiple convolution blocks, each block uses 3x3 convolution kernel, step 1, combined with ReLU activation function, extracts spatial and spectral features, 10 wave bands of each frame of image as multi-channel input, processed by convolution block in turn, generates high-dimensional feature map. The encoder gradually down-samples through max-pooling operation, reduces the spatial resolution of the feature map, while increases the number of feature channels, captures the deep features of deep cavity, plane and high light area, and generates a feature map containing multi-scale spatial information and spectral features.
[0047] The feature map generated by the encoder is received, and the pre-trained U-Net model is up-sampled by a decoder to restore the spatial resolution, specifically: the decoder includes multiple up-sampling blocks, each of which uses a transpose convolution to gradually enlarge the resolution of the feature map and restore it to the original size, consistent with the resolution of the input image. After each layer of up-sampling, the corresponding feature map of the encoder is combined to fuse the shallow spatial details and deep semantic information through a skip connection to enhance the segmentation accuracy. The fused feature map is further processed by a 3x3 convolution kernel and a ReLU activation function, the processing process being: applying a 3x3 convolution kernel to extract spatial and spectral features through a local receptive field, smoothing the splicing boundary, and outputting a feature map with a spatial resolution; applying a ReLU activation function pixel by pixel to enhance the boundary features of deep cavities and flat areas, suppress noise, and optimize the segmentation boundary clarity; the processed feature map is input into the next up-sampling block or the final 1x1 convolution layer, and the decoder outputs a feature map with the same resolution as the input image, with the number of channels reduced to 3, corresponding to the classification probabilities of deep cavities, flat areas, and other areas.
[0048] Based on the feature map output by the decoder, the U-Net model generates classification results for each pixel through a 1x1 convolution layer and a Softmax activation function, specifically: for each pixel, the U-Net model outputs probability values for 3 categories, corresponding to deep cavities, flat areas, and others, respectively, and takes the maximum probability value as the classification result, with pixel values of 1 indicating deep cavity regions, 2 indicating flat regions, and 0 indicating other regions. The classification result is used to determine the category of each pixel. The classification result forms a segmentation mask, with one mask generated for each frame of image, a total of 3 masks generated for 3 frames. The segmentation mask is used to mark deep cavity and flat area regions, and the segmentation mask is an image representation of the classification result. The light distribution map provides a reference for highlight, shadow, and deep cavity regions, and assists in verifying the accuracy of the segmentation mask to ensure that the deep cavity and flat area segmentation accuracy is maximized, with a greater-than-0.9 intersection over union ratio.
[0049] Further explanation, by using the U-Net model for accurate labeling of deep cavity and flat area regions through high-precision semantic segmentation, accurate region guidance is provided for HDR synthesis, and the segmentation reliability of the burr region is enhanced.
[0050] S3.5: Apply Debevec algorithm for HDR synthesis by segmentation mask and multi-exposure image set, generate high dynamic range image, the specific process is: based on the light distribution map, set the weight of Debevec algorithm, the weight of highlight area is low, the weight of shadow / deep cavity area is high, in order to enhance the visibility of deep cavity area. Debevec algorithm fuses 3 images of multi-exposure image set, calculates weighted average for 10 bands one by one, and normalizes pixel value to [0, 255] to generate high dynamic range image. In the fusion process, the segmentation mask provides guidance, preferentially enhances the spectral data of deep cavity area, retains the spectral characteristics of burr, oxidation layer and substrate, and generates high dynamic range image with high definition and enhanced deep cavity details.
[0051] Further explanation, the generated high dynamic range image integrates the spectral information and semantic segmentation results of multi-exposure images, retains the spectral characteristics of burr, oxidation layer and substrate, and enhances the visibility of deep cavity area.
[0052] S4: Apply gradient-based Poisson fusion algorithm to high dynamic range image for shadow elimination, and use pre-trained Mask-R-CNN model for detection after shadow elimination to output burr detection results.
[0053] Specifically includes the following steps, S4.1: For high dynamic range image and light distribution map, use SIFT feature point matching algorithm to extract feature points of shadow area, generate feature point pair set, specifically: SIFT feature point matching algorithm applies scale invariant feature transform to each band of high dynamic range image, detects feature points of shadow area and adjacent non-shadow area, calculates 128-dimensional feature descriptor, matches feature descriptors of shadow area and non-shadow area, and generates feature point pair set, records texture and edge information.
[0054] For each band of high dynamic range image, use the light distribution map to mark the shadow area pixels, initialize the shadow area pixel value as the average pixel value of the adjacent non-shadow area, and the initialized shadow area pixel value retains the spectral characteristics of 10 bands. Next, according to the initialized shadow area pixel value and high dynamic range image, combined with the feature point pair set, apply gradient-based Poisson fusion algorithm to calculate gradient field and fuse pixels, the specific process is: process 10 bands of high dynamic range image one by one, based on the texture and edge information in the feature point pair set, calculate the gradient field of non-shadow area, and extract the horizontal and vertical pixel changes. The gradient-based Poisson fusion algorithm solves the Poisson equation to generate a preliminary fusion image, and its expression is: ; Wherein, represents the Laplacian operator, represents the initial pixel value of the shadow area provided by the preliminary fusion image, represents the divergence operator, represents the gradient field.
[0055] After obtaining the preliminary fusion image, an iterative optimization is applied, the iterative process prioritizes the deep cavity area, controls the fusion error by comparing the gradient difference between the shadow area and the non-shadow area, and optimizes the pixel value of each waveband. After normalization, the pixel value is adjusted to 0-255, and the generated fusion image retains the spectral characteristics of burrs, oxidation layers and substrates. The preliminary fusion image is reconstructed by the gradient field The pixel value of the shadow area is reconstructed, and the pixel value of the shadow area is seamlessly connected with the texture and edge of the non-shadow area, eliminating the influence of the shadow area.
[0056] Based on the generated fusion image, the shadow elimination effect is verified, and a high-quality shadow-free HDR image is generated. The specific process is as follows: for each waveband of the fusion image, calculate the pixel value difference between the shadow area and the adjacent non-shadow area, ensure that the fusion error is less than the specified requirement, for example, less than 0.5%. Combined with the light distribution map, verify whether the shadow area is effectively filled, check the consistency of the details of the deep cavity area, confirm that the spectral characteristics of the burrs, oxidation layers and substrates are not lost, and verify that the fusion image forms a high-quality shadow-free HDR image, retaining high definition and spectral characteristics.
[0057] S4.2: Train the Mask-R-CNN model through the die casting image dataset labeled with burrs. The die casting image dataset labeled with burrs includes die casting images, labeled bounding boxes including burrs, and pixel-level masks, and is divided into a training set and a validation set, wherein each image contains 10 spectral wavebands. Each waveband is a 1024x1024 single-channel grayscale image, and the 10 wavebands together form a multi-channel input tensor. The Mask-R-CNN model includes a backbone network, a region proposal network, a detection head, and a mask branch. The backbone network uses ResNet-50-FPN to extract multi-scale features. ResNet-50 contains 50 layers of convolution, processes 10 waveband inputs, performs convolution operations on the 10 waveband inputs, extracts spatial and spectral features, generates feature maps, and FPN fuses multi-scale features through upsampling and horizontal connection to enhance the detection capability of the burr area. The region proposal network generates candidate regions on the feature map, outputs bounding boxes and object scores, and uses a 3x3 convolution kernel to predict candidate box coordinates and classification probabilities. Anchor box scales are 8, 16, 32, 64, and 128 pixels, and aspect ratios are 1:1, 1:2, and 2:1. The detection head performs bounding box regression and classification on the candidate regions, adjusts the bounding box coordinates using a fully connected layer, and predicts the burr class and confidence (range 0-1). The mask branch generates a pixel-level mask for each candidate region, and uses a 28x28 resolution convolution layer to predict a binary mask to mark the burr area.
[0058] The training process of the Mask-R-CNN model is as follows: cross-entropy loss is used to quantify the accuracy of the classification task, L1 loss is used to quantify the accuracy of the boundary box regression and mask prediction, the optimization goal is to maximize the average precision, 100 epochs are trained, the batch size is 8, the learning rate is 0.001, and the validation set is used to evaluate the Mask-R-CNN model. When the average precision reaches the specified requirement, for example, 0.9, it means that the Mask-R-CNN model can ensure accurate detection of burr regions, complete the training of the Mask-R-CNN model, and save the parameters of the trained Mask-R-CNN model.
[0059] The shadow-free high-quality HDR image is input into the trained Mask-R-CNN model. The shadow-free high-quality HDR image contains 10 bands, forming a 10-channel input tensor. The backbone network ResNet-50-FPN processes the 10-channel input. ResNet-50 contains 50 layers of convolution, extracts deep features through residual connection, and generates initial feature maps. FPN fuses the features of the initial feature maps at different convolution layers (different scales) through upsampling and horizontal connection to generate multi-scale feature maps. The multi-scale feature maps capture the texture, edge and spectral characteristics of the burr region, and are suitable for the detection of burrs of different sizes.
[0060] The region proposal network applies a 3x3 convolution kernel to each scale feature map to generate the boundary box and object score of the candidate region. The anchor box scale is 8, 16, 32, 64, 128 pixels, and the width-height ratio is 1:1, 1:2, 2:1, covering small to large size burr features. The region proposal network predicts whether each anchor box contains a burr (foreground or background) and outputs a candidate region. Each candidate region contains preliminary boundary box coordinates and object score (range 0-1). The candidate region is generated based on multi-scale feature maps to ensure the capture of burr features of different bands and sizes.
[0061] The detection head performs feature pooling on each candidate region to extract fixed-size features from the multi-scale feature maps. The boundary box regression is performed through the fully connected layer to adjust the candidate region coordinates to the accurate boundary box and predict the confidence, which is in the range of 0-1, representing the probability of the burr region. The mask branch generates a 28x28 resolution binary mask for each candidate region. The convolution layer is used to predict the pixel-level segmentation result. The pixel value of 1 represents the burr region, and 0 represents the background. Each band is processed one by one to generate the preliminary detection result of each band, which includes the boundary box, confidence and binary mask of the candidate burr region.
[0062] According to the preliminary detection result, qualified detection is screened and the size is converted, and the detection result of burrs is generated. Specifically, only the candidate region with a confidence greater than 0.95 is reserved as qualified detection, low-confidence regions are removed, non-maximum suppression is applied, an intersection over union threshold of 0.5 is set, the bounding boxes of the candidate regions are compared, the bounding box with the highest confidence is reserved, and the overlapping bounding boxes with an IoU greater than 0.5 are removed to ensure that each burr region corresponds to a unique detection result. The generated qualified detection result includes the bounding box, confidence (candidate burr region bounding box > 0.95), and binary mask of each burr region. For the bounding box of each qualified burr region, the pixel coordinates are converted to millimeter units using the camera intrinsic parameters to obtain the detection result of the burr. The detection result of the burr includes the bounding box, size, and binary mask of the burr. The binary mask value is a pixel-level contour that marks the burr region (pixel value 1 for burr and 0 for background). The confidence greater than 0.95 is set based on the high precision requirement of burr detection. The intersection over union threshold of 0.5 is set according to the detection task requirements, which means that when the overlapping area of two bounding boxes exceeds 50%, the bounding box with higher confidence is retained, which is suitable for the dense distribution of burr regions.
[0063] The embodiment also provides a die casting burr detection system based on machine vision, comprising: A collection module collects a die casting multispectral image, generates a polarized-spectral image set, and performs principal component analysis on the polarized-spectral image set to obtain a spectral feature vector set. A calibration module uses a curved surface adaptive optical compensation mechanism to reacquire the polarized-spectral image set through light path optimization based on the spectral feature vector set, obtains an optimized polarized-spectral image set, and performs correction to generate a high-resolution image set. A segmentation module performs illumination distribution analysis on the die casting surface to form a multi-exposure image set, and inputs the completed U-Net model for semantic segmentation to generate a high dynamic range image. A detection module applies a gradient Poisson fusion algorithm to the high dynamic range image to eliminate shadows, and uses a pre-trained Mask-R-CNN model for detection after shadow elimination to output the detection result of the burr.
[0064] The embodiment also provides a computer device suitable for the die casting burr detection method based on machine vision, comprising a memory and a processor. The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to implement the die casting burr detection method based on machine vision as described in the above embodiment.
[0065] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved by WIFI, an operator network, NFC (Near Field Communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.
[0066] The embodiment also provides a storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the method for detecting a die burr based on machine vision. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk.
[0067] In summary, the present application: through curved surface adaptive optical optimization of polarization-spectral image set, realizes the imaging optimization and geometric correction of complex curved surface area, generates high-resolution image set, improves the imaging quality of polarization-spectral image set, and through adaptive optical path adjustment and geometric correction, reduces the specular reflection and defocus distortion of high-reflective area, enhances the details of burr area, significantly improves the imaging resolution and clarity of complex curved surface and high-reflective area, overcomes the limitations of traditional fixed light path imaging on curved surface, in addition, by fully utilizing the spectral feature vector set and geometric information, combined with the high-precision segmentation and detection of U-Net and Mask-R-CNN model, the dependence of deep learning model on data set is overcome, stable burr positioning and accurate size measurement under complex lighting conditions are realized, and an efficient and automatic detection means for die casting quality control is provided.
[0068] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, which should be covered in the scope of claims of the present application.
Claims
1. A method for detecting burrs on die castings based on machine vision, characterized by: include, Collect multispectral images of die castings to generate a polarization-spectral image set, and perform principal component analysis on the polarization-spectral image set to obtain a spectral feature vector set; Based on the spectral feature vector set, the surface adaptive optics compensation mechanism is used to reacquire the polarization-spectral image set through optical path optimization, obtain the optimized polarization-spectral image set, and perform correction to generate a high-resolution image set. Using high-resolution image sets, we analyze the illumination distribution of the die-casting surface to form a multi-exposure image set. This image set is then fed into the constructed U-Net model for semantic segmentation to generate high dynamic range images. The gradient Poisson fusion algorithm is applied to remove shadows from high dynamic range images. After the shadows are removed, the pre-trained Mask-R-CNN model is used for detection to output the burr detection results.
2. The method for detecting burrs on die castings based on machine vision according to claim 1, wherein: The method of obtaining the spectral feature vector set specifically includes the following steps: The multispectral and polarization information of the die casting surface is captured by a hyperspectral camera and a polarization filter array, and a polarization-spectral image set is formed after preprocessing and segmentation; Principal component analysis is applied to each frame of the polarization-spectral image set. The first N principal components are retained through singular value decomposition, and the dimension is reduced to an N-dimensional feature matrix. A spectral feature vector set is generated based on the N-dimensional feature matrix.
3. The method for detecting burrs on die castings based on machine vision according to claim 2, wherein: The step of obtaining the optimized polarization-spectral image set specifically includes the following steps: Use a 3D curvature sensor to scan the die casting surface through laser triangulation to generate a high-precision 3D point cloud model; Generate a curvature distribution map using the Gaussian curvature formula based on the spectral feature vector set; The geometric characteristics of the die-casting surface are analyzed through the three-dimensional point cloud model and curvature distribution map. The high-reflective areas are identified by combining the spectral feature vector set. The optical path parameters are calculated and the imaging optical path is dynamically adjusted. The polarization-spectral image is re-collected using a hyperspectral camera to generate an optimized polarization-spectral image set.
4. The method for detecting burrs on die castings based on machine vision according to claim 3, wherein: Generating a high-resolution image set refers to performing geometric correction and clarity verification on the optimized polarization-spectral image set, and eliminating distortion.
5. The method for detecting burrs on die castings based on machine vision according to claim 4, wherein: The forming of the multi-exposure image set specifically includes the following steps: Use photosensors to analyze the light distribution on the surface of the die casting and generate a light distribution map; Based on the regional classification of the illumination distribution map, image acquisition with different exposure times is set to form a multi-exposure image set.
6. The method for detecting burrs on die castings based on machine vision according to claim 5, wherein: The generating of high dynamic range image specifically comprises the following steps: A U-Net model was trained based on a die-casting image dataset and optimized to meet specified requirements, completing the construction of the U-Net model. The multi-exposure image set is input into the constructed U-Net model, which outputs the segmentation mask through the encoder and decoder; The Debevec algorithm is applied based on the segmentation mask to synthesize images from a multi-exposure image set to generate a high dynamic range image.
7. The method for detecting burrs on die castings based on machine vision according to claim 6, wherein: Apply the gradient Poisson fusion algorithm to remove shadows from high dynamic range images, and use the pre-trained Mask-R-CNN model to detect burrs after shadow removal, and output the burr detection results. The specific steps include the following: The high dynamic range image is fused with the segmentation mask through gradient domain processing to generate a shadow-free high dynamic range image; The Mask-R-CNN model was trained using a dataset of burr die-casting images. The dataset was divided into a training set and a validation set. The accuracy of the classification task was measured using cross-entropy loss, and the accuracy of bounding box regression and mask prediction was measured using L1 loss. The Mask-R-CNN model training was completed after the validation set was used to evaluate the accuracy of the model and found that it met the specified requirements. The Mask-R-CNN model includes a backbone network, a region proposal network, a detection head, and a mask branch; The shadow-free high dynamic range image is input into the trained Mask-R-CNN model for semantic segmentation and instance detection, and the bounding box, size and binary mask of the burr are output.
8. A die-casting burr detection system based on machine vision, based on the die-casting burr detection method based on machine vision according to any one of claims 1 to 7, characterized in that: include, An acquisition module collects multispectral images of die castings, generates a polarization-spectral image set, and performs principal component analysis on the polarization-spectral image set to obtain a spectral feature vector set; The calibration module uses a surface adaptive optics compensation mechanism based on the spectral feature vector set to reacquire the polarization-spectral image set through optical path optimization, obtain the optimized polarization-spectral image set, and perform correction to generate a high-resolution image set; The segmentation module uses a high-resolution image set to analyze the illumination distribution of the die-casting surface to form a multi-exposure image set. This image set is then input into the constructed U-Net model for semantic segmentation to generate a high dynamic range image. The detection module applies the gradient Poisson fusion algorithm to remove shadows from high dynamic range images, and uses the pre-trained Mask-R-CNN model for detection after shadow removal, outputting the burr detection results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the die-casting burr detection method based on machine vision are implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting burrs on die castings based on machine vision are implemented.
Citation Information
Patent Citations
Casting burr detection method and system based on 3D vision
CN117237360A
Die-casting alloy workpiece defect detection method and system based on visual identification
CN118130477A
Flexible circuit board defect detection method and system
CN118644483A
Machine vision dynamic defect detection method and device for precise structural part
CN119887745A
Cited By
Thermal imaging defect quantitative detection method and system based on continuous laser line scanning
CN121164368A
Tomato sugar degree and hardness online detection method based on hyperspectral depth characteristics
CN121431404A
Shoe body quality detection method based on image visual analysis
CN121459160A