Visual inspection system and application method thereof
By designing a multi-modular vision detection system based on deep learning, the existing system's high error rate and insufficient adaptability in complex environments is solved, and high-precision, automation and stable defect detection are achieved, meeting the high efficiency and high stability requirements of industrial production.
Patent Information
- Application Number
- CN202510638248.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing vision detection system based on deep learning lacks high-precision detection capabilities in complex environments, especially in the fusion of multi-angle and multi-modal data, and at the same time, it lacks an effective adaptive optimization mechanism and feedback calibration function, making it difficult to meet the requirements of high precision, high efficiency and high stability in industrial production.
A visual detection system based on deep learning is designed, including image acquisition module, data preprocessing module, feature extraction module, defect detection module, adaptive optimization module, result output module, feedback calibration module, equipment control module and performance monitoring module. The system synchronously acquires multi-angle optical images and three-dimensional point cloud data through a multi-spectral camera and laser scanning unit, combines a deep learning algorithm to extract image features, and uses multi-scale feature fusion and attention mechanism for defect detection. At the same time, the system has the ability to dynamically adjust the detection threshold and classifier parameters, and perform model weight correction through manual review results to realize automated sorting and performance monitoring.
It realizes high-precision defect detection, can accurately identify cracks, scratches and foreign objects on the surface of the object, and improves the accuracy and robustness of the detection. Through adaptive optimization and feedback calibration mechanisms, the system can maintain stable detection accuracy under different operating conditions, and improve production efficiency and long-term operating stability of the system.
Smart Images

Figure CN120177494A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial automation detection and intelligent manufacturing. More specifically, the present invention relates to a vision detection system and its application method. Background Art
[0002] In modern industrial production, vision detection technology is widely used in product quality control and defect detection. Traditional vision detection methods mainly rely on manual visual inspection or rule-based image processing algorithms, which have many limitations in terms of detection accuracy, efficiency, and stability. With the development of industrial automation, the requirements for detection systems are getting higher and higher, especially for high-precision detection and fast response capabilities in complex environments. In recent years, deep learning technology has made remarkable progress in the field of image recognition and processing, bringing new development opportunities to vision detection technology. However, most existing deep learning-based vision detection systems only focus on two-dimensional image information, ignoring the three-dimensional morphological features of objects, resulting in a high misjudgment rate when detecting complex surface defects. In addition, the adaptability of existing systems in dynamic environments is poor, and it is difficult to perform adaptive optimization and model calibration based on actual detection results.
[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the prior art: on the one hand, the traditional vision detection system has insufficient recognition ability for complex defects, especially with obvious shortcomings in multi-angle and multi-modal data fusion; on the other hand, the existing system lacks an effective adaptive optimization mechanism and feedback calibration function, and it is difficult to meet the requirements of high precision, high efficiency, and high stability in industrial production. Summary of the Invention
[0004] The present invention provides a vision detection system and its application method.
[0005] In the first aspect of the present invention, a deep learning-based vision detection system is provided, including: An image acquisition module for real-time acquisition of multi-angle optical images and laser three-dimensional point cloud data of an object to be measured; A data preprocessing module for denoising, geometric correction, and multi-modal data alignment of the original image; A feature extraction module for extracting texture, edge, and defect features in the image through a convolutional neural network; A defect detection module for identifying cracks, scratches, and foreign objects on the surface of the object based on the feature fusion result; An adaptive optimization module for dynamically adjusting the detection threshold and classifier parameters according to the detection result; A result output module for generating a detection report and marking the defect location; A feedback calibration module for correcting the weights of the detection model according to the results of manual review; An equipment control module for triggering the sorting device to remove defective products; A performance monitoring module for statistically analyzing the detection accuracy and system response latency.
[0006] Further, the image acquisition module includes: A multispectral camera unit that synchronously acquires visible light and near-infrared band images at a rate of 200 ms / frame, with a resolution of not less than 4096×2160; A laser scanning unit that uses a line array laser to generate a three-dimensional point cloud with a density of 0.1 mm / point; A trigger control unit that starts the global shutter mode when the conveyor belt speed
[0007] Further, the data preprocessing module includes: A noise suppression unit that performs adaptive median filtering using the following formula: ; where is a 5×5 filtering window, is the average value of pixels within the window, is the standard deviation; A distortion correction unit that eliminates lens distortion through a perspective transformation matrix based on camera calibration parameters; A point cloud registration unit that aligns multi-viewpoint clouds to a unified coordinate system using the ICP algorithm.
[0008] Further, the feature extraction module uses an improved ResNet-50 network, specifically including: A multi-scale feature fusion layer that fuses the feature maps output by the 3rd, 4th, and 5th residual blocks using the following formula: ; where is the weight coefficient of the i-th layer ( ), is bilinear interpolation upsampling, is the channel attention weight, represents element-wise multiplication; An attention mechanism unit that dynamically adjusts the feature channel weights through the SE module.
[0009] Further, the objective function of the defect detection module is: ; where is the defect category cross-entropy loss, is the IoU loss for the defect area, is the weight decay term, is the loss weight coefficient, satisfying and .
[0010] Furthermore, the threshold adjustment rule of the adaptive optimization module is: ; wherein, is the updated detection threshold, is the original threshold, is the learning rate, is the current recall rate, is the false positive rate, is the cumulative number of samples.
[0011] Furthermore, the weight correction formula of the feedback calibration module is: ; wherein, is the updated weight of the k-th layer, is the manually reviewed label, is the model prediction value, M is the number of calibration samples, is the sparsification coefficient.
[0012] Furthermore, the performance monitoring module performs the following calculations: Accuracy metric: ; wherein, is the number of correctly detected defects, is the total number of actual defects; Response delay: ; wherein, is the triggering time of the k-th detection, is the result output time, K is the number of detections.
[0013] Furthermore, the device control module includes: A priority queue unit that sorts the sorting priorities according to the defect levels; A pulse control unit that generates a PWM signal to drive the solenoid valve, and the pulse width ; wherein, v is the conveyor belt speed, is the length of the defective product, is the safety margin.
[0014] In the second aspect of the present invention, an application method of a vision detection system based on deep learning is provided, including: Step 1: Synchronously collect the surface image and 3D topography data of the object through a multispectral camera and a laser scanning unit; Step 2: Perform median filtering, distortion correction, and point cloud registration preprocessing on the original data; Step 3: Use a multi-scale feature fusion network to extract texture and geometric features in the image; Step 4: Adopt a cascade classifier to identify the defect type, and determine the defect area through non-maximum suppression; Step 5: Dynamically adjust the classification confidence threshold according to the real-time detection accuracy; Step 6: Output a detection report with defect marks and store it in the database; Step 7: Backpropagate and correct the model weights based on the manual review results; Step 8: Generate a sorting control signal according to the defect position and grade; Step 9: Statistically analyze the system operation indicators and generate performance optimization suggestions.
[0015] According to the above embodiments of the present invention, it has at least the following beneficial effects: The vision detection system of the present invention can achieve high-precision defect detection. By synchronously collecting multi-angle optical images and 3D point cloud data through a multispectral camera and a laser scanning unit, and using deep learning algorithms to extract texture, edge, and defect features in the image, the system can accurately identify defects such as cracks, scratches, and foreign objects on the object surface. At the same time, the data preprocessing module can effectively remove noise, correct distortion, and align multi-modal data to ensure the reliability of the detection results. In addition, the feature extraction module adopts an improved ResNet-50 network and multi-scale feature fusion technology, and combines an attention mechanism to dynamically adjust the feature weights, further improving the accuracy of defect detection.
[0016] The system can also dynamically adjust the detection threshold and classifier parameters according to the real-time detection results to optimize the detection performance. The adaptive optimization module can adjust the detection threshold according to the current recall rate and false alarm rate to ensure stable detection accuracy under different working conditions. The feedback calibration module corrects the weights of the detection model according to the manual review results, further improving the adaptability and accuracy of the model. The device control module generates a sorting control signal according to the defect position and grade to realize automatic removal of defective products and improve production efficiency. The performance monitoring module statistically analyzes the detection accuracy and system response delay to provide data support for system optimization and ensure the high-efficiency and stability of the system during long-term operation. Brief Description of the Drawings
[0017] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become easily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, wherein: Figure 1 Schematic diagram of the structure of a vision detection system based on deep learning provided by an embodiment of the present invention; Figure 2 Flow schematic diagram of the application method of a vision detection system based on deep learning provided by an embodiment of the present invention; Figure 3 Schematically shows the structure diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0018] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present invention, rather than limiting the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to be able to fully convey the scope of the present invention to those skilled in the art.
[0019] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method, or computer program product. Therefore, the present invention can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0020] It should be noted that the quantity of any element in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0021] Below with reference to Figure 1 , Figure 1 Schematic diagram of the structure of a vision detection system based on deep learning provided by an embodiment of the present invention. As Figure 1 shown, a vision detection system 100 based on deep learning includes: An image acquisition module 101, configured to acquire multi-angle optical images and laser three-dimensional point cloud data of an object to be measured in real time; A data preprocessing module 102, configured to denoise, geometrically correct, and align multi-modal data for the original image; A feature extraction module 103, configured to extract texture, edge, and defect features in the image through a convolutional neural network; A defect detection module 104, configured to identify cracks, scratches, and foreign objects on the surface of the object based on the feature fusion result; An adaptive optimization module 105, configured to dynamically adjust the detection threshold and classifier parameters according to the detection result; A result output module 106, configured to generate a detection report and mark the defect position; A feedback calibration module 107, configured to correct the weights of the detection model according to the manual review result; The device control module 108 is used to trigger the sorting device to reject defective products; The performance monitoring module 109 is used to count the detection accuracy rate and the system response delay.
[0022] It should be noted that the vision detection system of the present invention includes multiple modules for realizing high-precision detection of surface defects of objects. The image acquisition module is used to obtain multi-angle optical images and laser three-dimensional point cloud data of the object to be detected in real time. The optical images can provide texture and color information of the object surface, while the laser three-dimensional point cloud data is used to obtain geometric shape and depth information of the object surface. The data preprocessing module is used to denoise, geometrically correct and align multi-modal data of the collected original images to ensure the data quality and consistency for subsequent processing. The feature extraction module extracts texture, edge and defect features in the image through a convolutional neural network. These features are key information for defect recognition. The defect detection module identifies cracks, scratches and foreign objects on the object surface based on the feature fusion result, and improves the detection accuracy by fusing different features. The adaptive optimization module can dynamically adjust the detection threshold and classifier parameters according to the detection results to adapt to different detection environments and requirements. The result output module is used to generate a detection report and mark the defect location for subsequent processing and recording. The feedback calibration module can correct the weights of the detection model according to the results of manual review to further improve the accuracy and adaptability of the model. The device control module is used to trigger the sorting device to reject defective products to achieve automated quality control. The performance monitoring module is used to count the detection accuracy rate and the system response delay to provide data support for system optimization.
[0023] Specifically, the multispectral camera unit in the image acquisition module can synchronously acquire visible light and near-infrared band images at a rate of 200 ms / frame, with a resolution of 4096×2160. This high-resolution and multi-band image acquisition method can provide rich detail information, which helps to more accurately identify defects. The laser scanning unit uses a line array laser to generate a three-dimensional point cloud with a density of 0.1 mm / point. This high-density point cloud data can accurately reflect the geometric shape of the object surface and provide support for the three-dimensional positioning of defects. The trigger control unit dynamically adjusts the shutter mode according to the conveyor belt speed. When the conveyor belt speed is greater than 2 m / s, the global shutter mode is activated to avoid motion blur and ensure that the acquired images are clear and usable. The noise suppression unit in the data preprocessing module uses an adaptive median filtering algorithm to dynamically adjust the filtering intensity by calculating the pixel mean and standard deviation within the filtering window, thereby effectively removing noise without losing image details. The distortion correction unit uses camera calibration parameters and a perspective transformation matrix to eliminate lens distortion and ensure the geometric accuracy of the images. The point cloud registration unit aligns multi-viewpoint clouds to a unified coordinate system through the ICP (Iterative Closest Point) algorithm, providing a basis for subsequent three-dimensional analysis. The feature extraction module uses an improved ResNet-50 network. The multi-scale feature fusion layer in it fuses feature maps at different levels through specific weight coefficients and bilinear interpolation upsampling techniques, and combines a channel attention mechanism to dynamically adjust feature weights, thereby more effectively extracting feature information related to defects. The objective function of the defect detection module combines classification, localization, and regularization losses. By optimizing the weight coefficients of these loss functions, the accuracy and robustness of the detection can be balanced. The adaptive optimization module dynamically adjusts the detection threshold according to the current recall rate and false alarm rate, and the learning rate setting can be adjusted according to actual needs to achieve better performance optimization. The weight correction formula of the feedback calibration module is based on the results of manual review, and corrects the model weights by calculating the difference between the predicted value and the true label. The introduction of the sparsification coefficient can prevent the model from overfitting and improve the generalization ability of the model. The performance monitoring module calculates indicators such as accuracy and response delay to monitor and evaluate the operation of the system in real time, providing data support for subsequent optimization.
[0024] Preferably, the multispectral camera unit in the image acquisition module can select a camera with higher resolution or faster speed according to the actual application scenario to meet different detection requirements. For example, in high-precision detection scenarios, the resolution can be increased to 8K or higher to obtain more delicate image details. For the laser scanning unit, a higher-density point cloud generation technology can be adopted, such as increasing the point cloud density to 0.05 mm / point, to further improve the accuracy of three-dimensional topography. In the data preprocessing module, the filter window size of the noise suppression unit can be adjusted according to the noise level of the image. For example, a 7×7 or larger filter window is used in high-noise environments. The camera calibration parameters of the distortion correction unit can be obtained through more accurate calibration methods to improve the accuracy of correction. The ICP algorithm of the point cloud registration unit can be combined with other optimization algorithms, such as genetic algorithms or particle swarm optimization algorithms, to improve the efficiency and accuracy of registration. In the feature extraction module, the improved ResNet-50 network can be further optimized. For example, the feature extraction ability can be enhanced by adjusting the number of residual blocks or introducing a deeper network structure. In the objective function of the defect detection module, the loss weight coefficient can be dynamically adjusted according to the defect type and the importance of the detection task to achieve a more flexible detection strategy. The threshold adjustment rule of the adaptive optimization module can introduce more performance indicators, such as the weighted average of precision and recall, to more comprehensively evaluate the detection performance and optimize the threshold. The weight correction formula of the feedback calibration module can combine more calibration samples and use optimization methods such as batch gradient descent to improve the efficiency and stability of weight correction. The performance monitoring module can introduce more performance indicators, such as detection speed and resource occupancy rate, to more comprehensively evaluate the operating status of the system and generate more detailed performance optimization suggestions based on these indicators.
[0025] In some embodiments, the image acquisition module includes: A multispectral camera unit that synchronously acquires visible light and near-infrared band images at a rate of 200 ms / frame, with a resolution of not less than 4096×2160; A laser scanning unit that uses a line array laser to generate a three-dimensional point cloud with a density of 0.1 mm / point; A trigger control unit that activates the global shutter mode when the conveyor belt speed
[0026] It should be noted that the image acquisition module is one of the core components of the vision detection system, and its function is to obtain multi-angle optical images and laser three-dimensional point cloud data of the object to be measured in real time. The multi-spectral camera unit can synchronously acquire visible light and near-infrared band images. Among them, the visible light image is mainly used to capture the color and texture information of the object surface, while the near-infrared band image can provide additional information about the object material and internal structure. The laser scanning unit generates a three-dimensional point cloud through a line array laser, which is used to accurately measure the geometric shape and depth information of the object surface. The trigger control unit dynamically adjusts the shutter mode according to the conveyor belt speed to ensure that the captured images are clear and not affected by motion blur, thereby providing a high-quality data basis for subsequent defect detection.
[0027] Specifically, the acquisition rate of the multi-spectral camera unit is 200 ms / frame, which can obtain high-resolution images in a short time, and its resolution is not less than 4096×2160 pixels. This means that the images can provide extremely high detail information, facilitating the system to detect tiny defects. The laser scanning unit uses a line array laser, and the density of the generated three-dimensional point cloud is 0.1 mm / point. This high-density point cloud can accurately reflect the microscopic geometric features of the object surface, such as the depth and shape of cracks and scratches. The trigger control unit will dynamically select the shutter mode according to the conveyor belt speed. When the conveyor belt speed is greater than 2 m / s, the system will switch to the global shutter mode to avoid image blur caused by the rapid movement of the object. The main difference between the global shutter mode and the rolling shutter mode is that the global shutter can expose the entire image sensor simultaneously, while the rolling shutter exposes line by line. Therefore, in high-speed motion scenarios, the global shutter can provide clearer images. The setting of these parameters and the application of these concepts enable the image acquisition module to stably and efficiently obtain high-quality image and point cloud data in complex industrial environments.
[0028] Preferably, the acquisition rate of the multi-spectral camera unit can be further optimized according to the actual application scenario. For example, when detecting high-speed moving objects, the acquisition rate can be increased to 100 ms / frame or higher to reduce the impact of motion blur on image quality. At the same time, the resolution can also be further improved according to the requirements of detection accuracy. For example, a camera with 8K or higher resolution can be used to capture more subtle defect features. For the laser scanning unit, the point cloud density can be further increased to 0.05 mm / point or higher to improve the detection ability for complex surface defects. In addition, the trigger control unit can set multiple thresholds according to the change range of the conveyor belt speed to implement a more refined shutter mode switching strategy. For example, when the conveyor belt speed is between 1 - 2 m / s, the rolling shutter mode is adopted to increase the image brightness; when the speed exceeds 2 m / s, it switches to the global shutter mode to ensure image clarity. This flexible control strategy can better adapt to the detection requirements at different speeds and improve the versatility and adaptability of the system.
[0029] In some embodiments, the data preprocessing module includes: A noise suppression unit that performs adaptive median filtering using the following formula: ; where is a 5×5 filtering window, is the mean value of pixels within the window, is the standard deviation; A distortion correction unit that eliminates lens distortion through a perspective transformation matrix based on camera calibration parameters; A point cloud registration unit that aligns multi-view point clouds to a unified coordinate system using the ICP algorithm.
[0030] It should be noted that the data preprocessing module is a key part of the vision detection system for improving the quality of image and point cloud data. This module processes the acquired raw data through three units: noise suppression, distortion correction, and point cloud registration, to ensure the accuracy and reliability of subsequent feature extraction and defect detection. The noise suppression unit uses an adaptive median filtering algorithm, which can dynamically adjust the filtering intensity according to the local features of the image, effectively removing noise while retaining image details. The distortion correction unit uses camera calibration parameters and a perspective transformation matrix to eliminate lens distortion and ensure the geometric accuracy of the image. The point cloud registration unit aligns multi-view point clouds to a unified coordinate system through the ICP (Iterative Closest Point) algorithm, providing a basis for 3D analysis.
[0031] Specifically, the adaptive median filtering algorithm of the noise suppression unit uses a 5×5 filtering window, and determines whether a pixel is a noise point by calculating the mean value and standard deviation of pixels within the window. If the difference between the pixel value and the window mean exceeds 3 times the standard deviation, the pixel is considered a noise point and filtered. This algorithm can dynamically adjust the filtering intensity according to the local statistical characteristics of the image, avoiding loss of image details caused by over-smoothing. The distortion correction unit uses camera calibration parameters to perform geometric correction on the image through a perspective transformation matrix, eliminating the influence of lens distortion on the shape and size of the image. Camera calibration parameters include focal length, principal point coordinates, and distortion coefficients, etc., which can be obtained through the calibration process to ensure that the corrected image conforms to the geometric characteristics of the actual object. The point cloud registration unit uses the ICP algorithm, iteratively calculates the closest point pairs between multi-view point clouds, and adjusts the pose of the point cloud, finally aligning all point clouds to a unified coordinate system to provide accurate geometric information for subsequent 3D defect detection.
[0032] Preferably, the filter window size of the noise suppression unit can be adjusted according to the resolution and noise level of the image. For example, in the case of high-resolution images or low noise levels, a 3×3 filter window can be used to reduce the computational load; while in the case of low-resolution images or high noise levels, a 7×7 filter window can be used to enhance the filtering effect. In the distortion correction unit, the camera calibration parameters can be obtained by a more precise calibration method, such as using the checkerboard calibration method or the circular calibration method, to improve the calibration accuracy. In addition, for the point cloud registration unit, the initial pose of the ICP algorithm can be estimated by a rough registration method (such as feature point-based registration) to accelerate the convergence speed of the algorithm and improve the registration accuracy. In some complex scenarios, other point cloud registration algorithms (such as deep learning-based registration methods) can also be combined as an alternative to further improve the efficiency and accuracy of point cloud alignment.
[0033] In some embodiments, the feature extraction module employs an improved ResNet-50 network, specifically including: A multi-scale feature fusion layer that fuses the feature maps output by the 3rd, 4th, and 5th residual blocks through the following formula: ; where is the weight coefficient of the i-th layer ( ), is bilinear interpolation upsampling, is the channel attention weight, represents element-wise multiplication; An attention mechanism unit that dynamically adjusts the feature channel weights through the SE module.
[0034] It should be noted that the feature extraction module is the core part of the visual detection system, and its role is to extract texture, edge, and defect features related to defect detection from the image through a convolutional neural network. This system adopts an improved ResNet-50 network structure, which introduces a multi-scale feature fusion layer and an attention mechanism unit to enhance the feature expression ability and discrimination. The multi-scale feature fusion layer fuses feature maps at different levels through specific weight coefficients and upsampling techniques to capture more detailed information. The attention mechanism unit dynamically adjusts the weights of the feature channels through the SE module, enabling the network to focus on features more relevant to defect detection, thereby improving the detection accuracy and robustness.
[0035] Specifically, the improved ResNet-50 network is optimized based on the traditional residual network. The multi-scale feature fusion layer fuses the feature maps output by the 3rd, 4th, and 5th residual blocks, aligns the feature maps of different resolutions through bilinear interpolation upsampling technology, and performs weighted fusion on the features using channel attention weights. Among them, the weight coefficients are 0.3, 0.5, and 0.2 respectively, and these weight coefficients can be adjusted according to the importance of different feature layers to better balance the feature contributions. The attention mechanism unit adopts the SE module (Squeeze-and-Excitation module), evaluates the importance of feature channels through global average pooling and fully connected layers, and dynamically adjusts the channel weights, enabling the network to adaptively focus on the feature channels more relevant to defect detection. This mechanism can effectively improve the discriminability of features, especially in the case of complex backgrounds or noise interference, further improving the detection performance.
[0036] Preferably, the multi-scale feature fusion layer in the feature extraction module can be further optimized according to the actual application scenario. For example, the weight coefficients can be adjusted according to the defect type and the distribution of image features to better adapt to different detection tasks. In addition, other methods can also be selected for the upsampling technology, such as nearest neighbor interpolation or deconvolution operation, to improve the quality of the fused features. For the attention mechanism unit, the structure of the SE module can be further expanded, for example, by adding more hidden layers or introducing more complex activation functions, to enhance the flexibility and accuracy of feature weight adjustment. In addition, other attention mechanisms (such as spatial attention mechanism) can be combined with the channel attention mechanism to further enhance the feature expression ability. In some specific application scenarios, it is also possible to consider introducing lightweight convolutional neural network structures (such as MobileNet or ShuffleNet) as alternatives to reduce the computational complexity and improve the real-time performance of the system.
[0037] In some embodiments, the objective function of the defect detection module is: ; where is the defect category cross-entropy loss, is the defect region IoU loss, is the weight decay term, is the loss weight coefficient, satisfying and .
[0038] It should be noted that the objective function of the defect detection module is a key part for optimizing the defect detection model, which improves the detection performance by comprehensively considering the losses in three aspects: classification, localization, and regularization. Among them, the classification loss is used to measure the model's ability to identify defect types, the localization loss is used to evaluate the accuracy of the model in defect regions, and the regularization loss is used to prevent the model from overfitting and ensure the generalization ability of the model. By reasonably setting the weight coefficients of these three losses, the accuracy, robustness, and complexity of the detection can be balanced, so as to achieve high-precision detection of defects such as cracks, scratches, and foreign objects on the object surface.
[0039] Specifically, the objective function consists of three parts: classification loss , localization loss and regularization loss . The classification loss usually adopts the cross-entropy loss function to measure the difference between the predicted defect category of the model and the true category; the localization loss adopts the IoU (Intersection over Union) loss to evaluate the coincidence degree between the predicted defect region of the model and the true region; the regularization loss adopts a weight decay term to prevent overfitting by restricting the magnitude of the model weights. In the objective function, the weight coefficients , and are respectively used to balance the contributions of these three losses, where the sum of and is 1, and the value of is 0.001. These parameters can be adjusted according to the requirements of the actual detection task to optimize the performance of the model.
[0040] Preferably, the weight coefficients in the objective function can be further adjusted according to the specific requirements of defect detection. For example, in some application scenarios, if higher precision requirements for defect localization are required, the value of can be increased, and the value of can be correspondingly reduced to enhance the sensitivity of the model to the localization loss. For the weight of the regularization loss, it can be dynamically adjusted according to the complexity of the model and the scale of the training data. When the amount of data is small, the value of can be appropriately increased to enhance the regularization effect and prevent overfitting; while when the amount of data is sufficient, the value of can be appropriately reduced to improve the fitting ability of the model. In addition, in addition to the cross-entropy loss and the IoU loss, other types of loss functions, such as the Dice loss or the Focal loss, can also be considered to further improve the detection performance of the model for unbalanced data or small target defects.
[0041] In some embodiments, the threshold adjustment rule of the adaptive optimization module is: ; Among them, is the updated detection threshold, is the original threshold, is the learning rate, is the current recall rate, is the false positive rate, is the cumulative number of samples.
[0042] It should be noted that the threshold adjustment rule of the adaptive optimization module is a key mechanism for dynamically optimizing the detection performance in the visual detection system. This mechanism dynamically adjusts the detection threshold according to the cumulative number of samples by combining the current recall rate and the false positive rate, so as to effectively reduce the false positive rate while ensuring the detection recall rate. This adaptive adjustment method enables the system to automatically optimize the detection parameters under different working conditions to adapt to the complex and changeable detection environment and ensure the accuracy and stability of the detection results.
[0043] Specifically, the threshold adjustment formula of the adaptive optimization module is: ; Among them, the threshold is a key parameter for judging whether the detection result is a defect; the recall rate refers to the proportion of actual defects that are correctly detected; the false positive rate refers to the proportion of non-defect samples that are misjudged as defects; the number of samples is the total number of samples accumulated in the detection; is the learning rate, which is used to control the amplitude of threshold adjustment. When the recall rate is higher than the false positive rate, the threshold will be appropriately reduced to improve the detection sensitivity; otherwise, the threshold will be increased to reduce the false positive. Through this dynamic adjustment mechanism, the system can continuously optimize its own detection performance during the detection process.
[0044] Preferably, the learning rate in the adaptive optimization module can be adjusted according to the complexity of the actual detection task and the sample distribution. For example, when the detection task is relatively complex and the sample categories are unbalanced, the learning rate can be appropriately reduced to avoid unstable system performance caused by too fast threshold adjustment. At the same time, the number of samples in the threshold adjustment formula can be calculated by means of a sliding window, only considering the number of samples in the recent period of time, rather than all the cumulative number of samples, so as to improve the real-time performance and adaptability of the system. In addition, in addition to the linear adjustment formula, a non-linear adjustment strategy can also be adopted, such as introducing a sigmoid function to smoothly adjust the threshold to better cope with the threshold changes in extreme cases.
[0045] In some embodiments, the weight correction formula of the feedback calibration module is: ; Among them, is the updated weight of the k-th layer, is the manually reviewed label, is the model prediction value, M is the number of calibration samples, is the sparsification coefficient.
[0046] It should be noted that the feedback calibration module is a key part of the visual detection system for correcting the weights of the detection model according to the results of manual review. This module calculates the difference between the manually reviewed label and the model prediction value, and adjusts the model weights in combination with the sparsification coefficient, thereby optimizing the detection performance of the model. This feedback mechanism can effectively make up for the deficiencies of the model in practical applications, improve the accuracy and adaptability of detection, and ensure that the system maintains high performance during long-term operation.
[0047] Specifically, the weight correction formula of the feedback calibration module is: ; where, and represent the updated weight and the original weight respectively; is the label of manual review, representing the actual defect situation; is the prediction value of the model; M is the number of calibration samples; is the learning rate, used to control the amplitude of weight update; is the sparsification coefficient, used to prevent the model from overfitting. By minimizing the mean square error between the prediction value and the true label, and combining the L1 regularization of the weights, the model can maintain simplicity while correcting the weights, and improve the generalization ability.
[0048] Preferably, the learning rate and the sparsification coefficient in the feedback calibration module can be adjusted according to the actual application scenario. For example, when the model deviation is large, the learning rate can be appropriately increased to speed up the weight correction speed; while when the model is close to convergence, the learning rate is reduced to avoid overcorrection. The sparsification coefficient can be adjusted according to the complexity of the model. When the model has more parameters, appropriately increase to enhance the regularization effect. In addition, in addition to L1 regularization, other regularization methods such as L2 regularization or Dropout can also be considered to further optimize the generalization ability of the model. In practical applications, the batch normalization technique can also be combined to standardize the input data, thereby improving the efficiency and stability of weight correction.
[0049] In some embodiments, the performance monitoring module performs the following calculations: Accuracy metric: ; where, The number of defects correctly detected is the total number of actual defects; Response delay: ; Among them, is the triggering time of the k-th detection, is the result output time, and K is the number of detections.
[0050] It should be noted that the performance monitoring module is a key part of the vision detection system for evaluating and optimizing the system operation status. This module provides a quantitative basis for the system performance evaluation by calculating two core indicators: detection accuracy and system response delay. The detection accuracy reflects the correctness of the system in defect detection, while the response delay measures the time efficiency of the system from detection triggering to result output. By monitoring these indicators in real time, the system can timely discover potential problems and optimize them, thus ensuring the efficiency and reliability of the detection process.
[0051] Specifically, the accuracy calculation formula of the performance monitoring module is: ; Among them, the number of defects correctly detected refers to the number of samples that the system successfully identifies and marks as defects, and the total number of actual defects refers to the total number of defect samples confirmed by manual review or other reliable means. The calculation formula of the response delay is: ; Among them, K represents the number of detections, is the starting time point of the system detection, is the time point when the system completes the detection and outputs the result. Through these formulas, the performance monitoring module can accurately quantify the detection performance and time efficiency of the system, providing data support for subsequent optimization.
[0052] Preferably, the performance monitoring module can further refine the monitoring indicators and optimization strategies. For example, when calculating the accuracy, the weights of different types of defects can be introduced to more accurately reflect the performance of the system in actual applications. For the response delay, different priorities can be set, and stricter delay monitoring can be carried out for high-priority defect detection tasks. In addition, in addition to accuracy and response delay, other performance indicators such as recall rate, false alarm rate, and system resource occupancy rate can be introduced to comprehensively evaluate the system operation status. In actual applications, the performance monitoring module can combine machine learning algorithms to automatically adjust the monitoring threshold according to historical data to achieve intelligent performance optimization.
[0053] In some embodiments, the device control module includes: A priority queue unit that sorts the sorting priorities according to the defect levels; The pulse control unit generates a PWM signal to drive the solenoid valve, and the pulse width ; where v is the conveyor belt speed, is the length of the defective product, is the safety margin.
[0054] It should be noted that the equipment control module is a key part of the vision inspection system for realizing automatic sorting. This module triggers the sorting device according to the defect detection result, removes the detected defective products, so as to realize the automatic quality control of the production process. The core functions of the equipment control module include priority queue management and pulse signal control. The priority queue is used to divide the sorting order according to the severity of the defect, and the pulse signal is used to drive the execution action of the sorting device. Through this automatic control mechanism, the system can efficiently process the detected defective products, improve production efficiency and reduce manual intervention.
[0055] Specifically, the priority queue unit in the equipment control module divides the sorting priority according to the defect level to ensure that products with a high defect level can be processed first. The defect level is usually comprehensively evaluated according to factors such as the type, size and position of the defect. The pulse control unit generates a PWM signal according to the conveyor belt speed, the length of the defective product and the safety margin, and drives the solenoid valve or other sorting devices. The calculation formula for the pulse width is: ; where, is the conveyor belt speed, the length of the defective product is the length of the detected defective area, and the safety margin is an additional time or space buffer set to ensure the reliability of the sorting action. In this way, the equipment control module can accurately control the action of the sorting device to ensure that the defective products are accurately removed.
[0056] Preferably, the priority queue in the equipment control module can be set more flexibly according to the actual production requirements. For example, in addition to dividing the priority based on the defect level, other factors in the production process, such as product batches and production order, can be combined to further optimize the sorting strategy. For the generation of the pulse signal, more complex control algorithms can be introduced, such as pulse width adjustment based on PID control, to adapt to the sorting requirements under different speeds and load conditions. In addition, redundant design can be considered, such as setting multiple sorting channels or standby sorting devices, to improve the reliability and fault tolerance of the system. In practical applications, the equipment control module can also be integrated with the production management system (MES) to realize real-time feedback of production data and optimized scheduling.
[0057] The above embodiments of the present invention have the following beneficial effects: The image acquisition module of the present invention can obtain multi-angle optical images and laser three-dimensional point cloud data of the object to be measured in real time, providing a rich information basis for subsequent defect detection. The data preprocessing module can denoise, geometrically correct, and align multi-modal data for the original images, ensuring the quality and consistency of the input data. The feature extraction module extracts texture, edge, and defect features in the images through a convolutional neural network. Combining multi-scale feature fusion and attention mechanism, it can capture defect features more accurately. The defect detection module identifies cracks, scratches, and foreign objects on the object surface based on the feature fusion results. Its objective function fuses classification, localization, and regularization losses, effectively improving the accuracy and robustness of detection. The adaptive optimization module can dynamically adjust the detection threshold and classifier parameters according to the detection results, enabling the system to automatically optimize its performance according to the actual working conditions. The feedback calibration module can correct the weights of the detection model according to the results of manual review, further improving the adaptability and accuracy of the model. The performance monitoring module can count the detection accuracy and system response delay, providing data support for system optimization and ensuring the system remains efficient and stable during long-term operation.
[0058] The multi-spectral camera unit in the image acquisition module synchronously acquires visible light and near-infrared band images at a high frame rate. The laser scanning unit generates high-density three-dimensional point clouds. The trigger control unit flexibly adjusts the shutter mode according to the conveyor belt speed. These designs can ensure the system stably acquires high-quality data under different working conditions. The noise suppression unit in the data preprocessing module uses an adaptive median filtering algorithm. The distortion correction unit eliminates lens distortion through a perspective transformation matrix. The point cloud registration unit aligns multi-viewpoint clouds using the ICP algorithm. These technical means can effectively improve the effect and efficiency of data preprocessing. The feature extraction module uses an improved ResNet-50 network, combined with a multi-scale feature fusion layer and channel attention mechanism, to more comprehensively extract and utilize image features. The defect detection module can improve the comprehensive performance of detection by optimizing the objective function and balancing classification, localization, and regularization losses. The threshold adjustment rule of the adaptive optimization module can dynamically adjust the detection threshold according to the current recall rate and false alarm rate. The weight correction formula of the feedback calibration module can correct the model weights according to the results of manual review. These mechanisms can further improve the adaptive ability and detection accuracy of the system.
[0059] As Figure 2 shown, an application method 200 of a vision detection system based on deep learning in some embodiments, the method 200 includes: Step 1: Synchronously acquire the surface image and three-dimensional topography data of the object through a multi-spectral camera and a laser scanning unit; Step 2: Perform median filtering, distortion correction, and point cloud registration preprocessing on the original data; Step 3: Use a multi-scale feature fusion network to extract texture and geometric features in the image; Step 4: Employ a cascade classifier to identify the defect type and determine the defect area through non-maximum suppression; Step 5: Dynamically adjust the classification confidence threshold according to the real-time detection accuracy; Step 6: Output a detection report with defect markings and store it in the database; Step 7: Backpropagate and correct the model weights based on the results of manual review; Step 8: Generate sorting control signals according to the defect location and grade; Step 9: Statistically analyze the system operation metrics and generate performance optimization suggestions.
[0060] It can be understood that the steps described in the application method 200 of the deep learning-based visual detection system correspond to the respective modules in the deep learning-based visual detection system described in the reference Figure 1 Therefore, the modules, features, and beneficial effects described above for the deep learning-based visual detection system also apply to the application method 200 of the deep learning-based visual detection system and the operations included therein, and will not be elaborated herein.
[0061] Next, refer to Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present invention. The electronic device in some embodiments of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The terminal device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0062] As Figure 3 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage device 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0063] Typically, the following devices can be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had. Figure 3 Each block shown in can represent one device or multiple devices as needed.
[0064] Furthermore, the storage medium of the embodiment of the present application stores program instructions capable of implementing all the above methods. Among them, the program instructions can be stored in the above storage medium in the form of a software product, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. And the foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.
[0065] The above description is only some preferred embodiments of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present invention.
Claims
1. A visual inspection system, characterized in that: Includes the following modules: Image acquisition module, used to obtain multi-angle optical images and laser three-dimensional point cloud data of the object to be measured in real time; Data preprocessing module, used for denoising, geometric correction and multimodal data alignment of original images; Feature extraction module, used to extract texture, edge and defect features in images through convolutional neural network; Defect detection module, used to identify cracks, scratches and foreign objects on the surface of objects based on feature fusion results; Adaptive optimization module, used to dynamically adjust detection thresholds and classifier parameters according to detection results; Result output module, used to generate inspection reports and mark defect locations; Feedback calibration module, used to correct the detection model weights according to manual review results; Equipment control module, used to trigger the sorting device to remove defective products; The performance monitoring module is used to count the detection accuracy and system response delay.
2. The system according to claim 1, characterized in that The image acquisition module comprises: A multispectral camera unit that simultaneously collects visible and near-infrared band images at a preset rate; Laser scanning unit, which uses linear array laser to generate 3D point cloud; The control unit is triggered to start the global shutter mode when the conveyor belt speed is greater than a preset speed threshold, otherwise the rolling shutter mode is adopted.
3. The system according to claim 1, characterized in that The data preprocessing module comprises: The noise suppression unit uses the following formula for adaptive median filtering: ; in, is a 5×5 filter window, is the mean value of pixels in the window, is the standard deviation; is the pixel intensity value of the original image at the coordinate (x, y), is the median filter operator; A distortion correction unit that eliminates lens distortion through a perspective transformation matrix based on camera calibration parameters; The point cloud registration unit uses the ICP algorithm to align multi-view point clouds to a unified coordinate system.
4. The system according to claim 1, characterized in that The feature extraction module adopts an improved ResNet-50 network, which specifically includes: The multi-scale feature fusion layer fuses the feature maps output by the 3rd, 4th, and 5th residual blocks through the following formula: ; in, is the weight coefficient of the i-th layer, is bilinear interpolation upsampling, is the channel attention weight, Represents element-wise multiplication; It is the feature map output by the i-th residual block; The attention mechanism unit dynamically adjusts the feature channel weights through the SE module.
5. The system according to claim 1, characterized in that The objective function of the defect detection module is: ; in, is the defect category cross entropy loss, is the IoU loss of the defect area, is the weight decay term, is the loss weight coefficient.
6. The system according to claim 5, characterized in that The threshold adjustment rule of the adaptive optimization module is: ; in, is the updated detection threshold, is the original threshold, is the learning rate, is the current recall rate, is the false alarm rate, is the cumulative sample size.
7. The system according to claim 1, characterized in that The weight correction formula of the feedback calibration module is: ; in, is the updated weight of the kth layer, is the original weight of the kth layer, To manually review the labels, is the model prediction value, M is the number of calibration samples, is the sparsification coefficient, is the set of all weight parameters of the neural network.
8. The system according to claim 1, characterized in that The performance monitoring module is used to calculate the accuracy index and response delay, including: The calculation accuracy index is shown in the following formula: ; in, is the number of defects detected correctly, is the actual total number of defects; The response delay is calculated as follows: ; in, is the kth detection trigger time, is the result output time, and K is the number of detections.
9. The system according to claim 1, characterized in that The device control module comprises: Priority queue unit, which divides sorting priorities according to defect levels; The pulse control unit generates a PWM signal to drive the solenoid valve. The pulse width is shown in the following formula: ; Where v is the conveyor belt speed, is the length of defective product, For safety margin.
10. An application method of a visual inspection system, applied to a visual inspection system according to any one of claims 1 to 9, characterized in that: The following steps are involved: Step 1: synchronously collect the surface image and three-dimensional shape data of the object through the multispectral camera and the laser scanning unit; Step 2: Perform median filtering, distortion correction and point cloud registration preprocessing on the original data; Step 3: Use a multi-scale feature fusion network to extract texture and geometric features from the image; Step 4: Use cascade classifier to identify defect types and determine defect areas through non-maximum suppression; Step 5: Dynamically adjust the classification confidence threshold according to the real-time detection accuracy; Step 6: Output the inspection report with defect marks and store it in the database; Step 7: Back-propagate and modify the model weights based on the manual review results; Step 8: Generate sorting control signals according to defect locations and levels; Step 9: Count system operation indicators and generate performance optimization suggestions.
Citation Information
Patent Citations
Text quality inspection automatic training method, electronic device and computer equipment
CN109740760A
Textile fabric product defect detection method and system
CN119290896A
Industrial product quality detection method and system based on machine vision
CN119600032A
Intelligent detection and evaluation method and system for underground pipeline defects
CN119848675A
Ice landslide monitoring and early warning method and system based on AI image recognition
CN120014378A
Cited By
Traditional Chinese medicine quality detection method and system
CN120404616A
Method and system for detecting secondary machining defect of internal thread of copper pipe
CN120427646A
Equipment calibration method and system for photoelectric sorting machine
CN120733998A
A calibration method and system for photoelectric sorting machines
CN120733998B
Visual detection method and device for shield tunnel segment assembly
CN120852305A