Product defect detection method, system and device based on multi-camera and deep learning fusion, and storage medium
Through the fusion method of multi-camera and deep learning, the limitations of multi-view video fusion and detection accuracy in the prior art are solved, and more efficient, accurate and economical product defect detection is achieved.
Patent Information
- Application Number
- CN202510348033.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art has limitations in multi-view video fusion, effective video extraction, detection accuracy, detection inference performance and cost, and it is difficult to meet the industrial site's demand for real-time and limited resources.
The panoramic video and images of the product are obtained through multiple cameras, and the fusion algorithm of SIFT and ORB operators is used to form a synthetic video of multi-view information. The keyframes are screened in combination with the dual attention mechanism of A2-Nets, and defect detection and fine inspection are performed based on Yolo and Swin Transformer operators.
It realizes a more comprehensive capture of the three-dimensional characteristics of the product, improves the accuracy and coverage of detection, reduces the computing resource requirements, optimizes the detection and reasoning performance, and reduces hardware and labor costs.
Smart Images

Figure CN120182243A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of industrial automation and intelligent manufacturing, and particularly relates to a product defect detection method, system, device and storage medium based on the fusion of multiple cameras and deep learning. Background Art
[0002] In the fields of industrial automation and intelligent manufacturing, product defect detection is a key link to ensure product quality and production efficiency. With the development of technology, defect detection methods have evolved from manual inspection to rule-based methods and then to automated detection based on deep learning. However, there are still some limitations in the existing technologies:
[0003] Manual Inspection:
[0004] Traditional manual inspection methods are inefficient and are easily affected by the experience and fatigue of inspectors, making it difficult to ensure the stability and consistency of inspection results.
[0005] Rule-based Methods:
[0006] Early automated defect detection systems relied on preset rules and thresholds. Although they had certain effects on known types of defects, they showed great limitations and lacked flexibility and adaptability when faced with complex and variable defect types and environmental changes.
[0007] Deep Learning-based Methods:
[0008] In recent years, defect detection methods based on deep learning have gradually emerged. By training models such as convolutional neural networks (CNNs), they have shown superior performance. However, these models usually require a large amount of computing resources and data for training and inference, resulting in an increase in system complexity and making it difficult to meet the requirements of real-time performance and limited resources in industrial sites.
[0009] In terms of the conditions of data acquisition and sources, currently, methods based on deep learning mainly adopt single-view cameras, binocular cameras, and multi-camera solutions. In the related technologies of single-view camera detection, the biggest problem is the inability to comprehensively capture the three-dimensional features of products, resulting in limited accuracy of defect detection; in the aspect of binocular camera collaborative detection, there are problems such as difficult stereo matching, high calibration complexity, large cost and maintenance issues, and poor environmental adaptability; in the aspect of multi-camera collaborative detection, some patented technologies have also been explored. For example, a multi-camera collaborative object tracking method based on deep learning uses Faster R-CNN for object detection, obtains the real position information of the object through the mapping relationship between image coordinates and planar geospatial coordinates, and realizes object detection and tracking in the multi-camera video surveillance scenario. However, this method mainly focuses on object tracking rather than defect detection, and there are certain limitations in processing multi-view image fusion and feature extraction.
[0010] In addition, in video-based streaming computing, there is currently a serious problem: the video stream contains a large number of irrelevant frames, increasing the computational burden, especially a serious challenge to large deep learning models; existing detection models are difficult to balance between accuracy and speed; at the same time, the high demand for computing power resources also limits its wide application in actual production.
[0011] Regarding the deep learning model for defect detection, the YOLO series of models have been widely used in industrial defect detection due to their high detection speed and good detection accuracy. For example, a deep learning chip packaging crack defect detection method based on YOLO realizes fast and accurate detection of chip packaging crack defects by optimizing the YOLO model structure and training strategy. However, the YOLO model still has room for improvement in its recognition ability when dealing with complex scenarios and small target defects.
[0012] In addition, as an emerging deep learning model, Swin Transformer demonstrates powerful capabilities in image feature extraction. A surface defect detection method based on Swin Transformer can accurately detect and locate the defect positions through a U-shaped symmetric encoder-decoder structure network with skip connections and attention mechanisms added. However, this method mainly focuses on defect detection for single-view images and has deficiencies in multi-view image fusion and comprehensive analysis. Summary of the Invention
[0013] The purpose of this application is to provide a product defect detection method, system, device, and storage medium based on the fusion of multi-cameras and deep learning to solve the problems of multi-view video fusion, effective video extraction, detection accuracy, detection inference performance, and cost limitations in existing defect detection technologies.
[0014] To achieve the above object, an embodiment of the present application provides a product defect detection method based on the fusion of multiple cameras and deep learning, including the following steps:
[0015] Obtain the panoramic video and images of the product through multiple cameras, and use a fusion algorithm of the Scale-Invariant Feature Transform (SIFT) operator and the Oriented FAST and Rotated BRIEF (ORB) operator for multi-view vision matching to synthesize the video streams captured by the cameras in the same frame to form a synthesized video containing multi-view information of the product;
[0016] Based on the dual attention mechanism of the A2-Nets operator, screen out the key frame images related to product defect detection from the video streams of the synthesized video, and reduce the number of frames of the same product to a calibrated minimum value;
[0017] Based on the Yolo operator, comprehensively analyze the size and surface features of the product in the key frame images. If any frame in the key frame images of the same product is recognized as defective, then initially determine that the product is a defective product;
[0018] Based on the Swin Transformer operator, perform integrated weighted voting on all the key frame images of the product initially determined to be defective in two dimensions: external dimension detection and appearance surface detection. If the voting result in any dimension of the external dimension detection and appearance surface detection is defective, then finally determine that the corresponding product is defective; otherwise, determine that it is not defective.
[0019] Optionally, before obtaining the panoramic video and images of the product through multiple cameras, it further includes:
[0020] Calibrate the camera parameters, correct the image distortion of the images captured by the cameras, and perform data fusion calibration, video frame extraction calibration, labeled data generation, and data consistency verification on the images from multiple cameras.
[0021] Optionally, the camera parameters include internal parameters and external parameters. The internal parameters include focal length and principal point coordinates, and the external parameters include position and attitude.
[0022] Optionally, the video frame extraction calibration specifically includes: calibrating the key area of the recognized product in the video according to the external shape characteristics of the product, the physical parameters of the camera, and the installation environment factors, and at the same time, calibrating the number of video frames to be extracted.
[0023] Optionally, the labeled data generation specifically includes: adding labeled information to the captured image data during the calibration process to generate a labeled data set, and the labeled information includes defect type, defect location, and defect area size.
[0024] Optionally, the data consistency check specifically includes: performing a consistency check on the same product data collected by different cameras.
[0025] Optionally, after the final determination that the corresponding product is defective, it further includes:
[0026] Transmitting the defect data detected in the shape and size detection and the appearance surface detection to other systems or devices through a network or a signal. The defect data includes the defect type, defect location, defect area size, and defect qualification rate.
[0027] To achieve the above object, the present application further provides a product defect detection system based on the fusion of multiple cameras and deep learning, including:
[0028] A video acquisition module, configured to obtain panoramic videos and images of a product through multiple cameras, and use a fusion algorithm of the Scale-Invariant Feature Transform (SIFT) operator and the Oriented FAST and Rotated BRIEF (ORB) operator for multi-view vision matching to synthesize the video streams collected by the cameras in the same frame to form a synthesized video containing multi-view information of the product;
[0029] A video frame extraction module, configured to screen out key frame images related to product defect detection from the video stream of the synthesized video based on the dual attention mechanism of the A2-Nets operator, and reduce the number of frames of the same product to a calibrated minimum value;
[0030] A defect preliminary inspection module, configured to comprehensively analyze the size and surface features of the product in the key frame images based on the Yolo operator. If any one of the key frame images of the same product is identified as defective, the product is preliminarily determined to be a defective product;
[0031] A defect refined inspection module, configured to perform integrated weighted voting on all key frame images of the product preliminarily determined to be defective based on the Swin Transformer operator in two dimensions: shape and size detection and appearance surface detection. If the voting result of any dimension of the shape and size detection and the appearance surface detection is defective, the corresponding product is finally determined to be defective; otherwise, it is determined to be non-defective.
[0032] To achieve the above object, the present application further provides a product defect detection device based on the fusion of multiple cameras and deep learning, including: a memory; and
[0033] A processor connected to the memory, and the processor is configured to execute the steps of the method as described above.
[0034] To achieve the above object, the present application further provides a computer storage medium, on which a computer program is stored, and when the computer program is executed by a machine, it implements the steps of the method as described above.
[0035] The embodiments of this application have the following advantages:
[0036] Advantages of multi-view video fusion:
[0037] Comprehensively capture three-dimensional features: By placing 3 or 4 industrial cameras at different angles, it is possible to simultaneously obtain the front and side videos of the product, forming a synthetic video containing multi-view information. This multi-view fusion method captures the three-dimensional features of the product more comprehensively than single-view video acquisition, providing a richer data basis for defect detection and helping to improve the accuracy of detection.
[0038] Improve the detection coverage: In traditional single-view detection, some defects may be blocked or difficult to observe due to view limitations. However, the dual-camera system of this application can observe the product from different angles, effectively expanding the detection coverage and reducing the possibility of missed detections.
[0039] Innovation in video frame extraction:
[0040] Intelligently screen key frames: Introduce the dual attention mechanism of the A2-Nets (Double Attention Networks) operator, which can intelligently screen out the key frames related to product defect detection from a large number of video frames. This mechanism can accurately focus on the feature changes and abnormal areas in the frames, reducing the number of frames for the same product to an extremely small value, such as about 5 frames. Significantly reduce the computational amount: By greatly reducing the number of frames to be processed, this application can geometrically reduce the computational amount. This not only improves the processing efficiency of the system but also reduces the demand for computing resources, enabling the defect detection system to complete the detection task in a shorter time.
[0041] Improve the accuracy of defect detection:
[0042] Combine initial inspection and refined inspection: The defect initial inspection module is built based on the Yolo operator and can quickly screen out potential defective products and initially detect defective products. The defect refined inspection module is based on the Swin Transformer operator. Although the detection speed is relatively slow, the detection effect is more accurate. This combination of initial inspection and refined inspection not only ensures the detection efficiency but also improves the detection accuracy.
[0043] Comprehensively analyze dimensions and surface features: In the defect initial inspection module, the dimension detection module and the surface detection module work together to comprehensively analyze the shape, size, and surface features of the product. Especially in the determination based on the shape and size, when the distance between the camera and the product is within 5 meters and the speed of the product assembly line or conveyor belt is within 5 m / s, it is possible to achieve sub-millimeter-level detection and determination of the shape and size.
[0044] Optimize the detection inference performance:
[0045] Efficient Frame Extraction Improves Inference Speed: The efficient frame selection mechanism of the video frame extraction module reduces the number of frames that need to be inferred, thus accelerating the inference speed of the entire defect detection process. This is particularly important for real-time online detection systems and can meet the requirements for detection speed on fast production lines.
[0046] Balancing Speed and Precision: Although the detection speed of the defect refined inspection module is relatively slow, its precision is higher, and it can further verify and supplement the detection results based on the initial inspection. Overall, this application optimizes the detection inference performance by optimizing the collaborative work of each module while ensuring the detection accuracy.
[0047] Cost Control and Economic Benefits:
[0048] Reducing Hardware Costs: Compared with multi-view detection systems that require multiple cameras or complex sensors, this application can be achieved by only using three industrial cameras in the hundreds of yuan range + a computer in the thousands of yuan range, reducing the input cost of hardware devices.
[0049] Reducing Labor Costs: The highly automated defect detection system can replace a large amount of manual inspection work, reduce the dependence on manual inspectors, and lower labor costs. At the same time, the system can provide real-time feedback on detection results, facilitating the timely handling of defective products and avoiding greater losses caused by defective products flowing into subsequent production processes.
[0050] Real-time Feedback and Data Consistency:
[0051] Real-time Data Output: The data output module can transmit the detected defect data, such as defect type, location, size, etc., to other systems or devices in real time through the network or signals. This real-time feedback mechanism helps to adjust production management and quality control in a timely manner, improving the flexibility and response speed of production.
[0052] Data Consistency Verification: In the data calibration module, consistency verification is performed on the data of the same product collected by different cameras to ensure the accuracy and consistency of the data from different perspectives. This provides a reliable data basis for subsequent defect detection and avoids detection errors caused by inconsistent data.
[0053] In summary, this application has significant beneficial effects in aspects such as multi-view video fusion, video frame extraction, defect detection accuracy, detection inference performance, cost control, and real-time feedback and data consistency. It can effectively solve the problems existing in the existing defect detection technology and provide a more efficient, accurate, and economical solution for product defect detection in industrial production. Brief Description of the Drawings
[0054] To more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.
[0055] Figure 1 Flowchart of a product defect detection method based on the fusion of multiple cameras and deep learning provided by at least one embodiment of the present application;
[0056] Figure 2 Logic block diagram of a product defect detection method based on the fusion of multiple cameras and deep learning provided by at least one embodiment of the present application;
[0057] Figure 3 Connection block diagram of a product defect detection system based on the fusion of multiple cameras and deep learning provided by at least one embodiment of the present application;
[0058] Figure 4 Module block diagram of a product defect detection device based on the fusion of multiple cameras and deep learning provided by at least one embodiment of the present application. Specific embodiments
[0059] The following specific embodiments illustrate the embodiments of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of them. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0060] It should be noted that in the claims and the specification of the present application, the steps can be executed basically in parallel or in the reverse order under appropriate circumstances, depending on the functions involved.
[0061] In addition, the technical features involved in different embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0062] An embodiment of the present application provides a product defect detection method based on the fusion of multiple cameras and deep learning. Refer to Figure 1 and Figure 2 , Figure 1 and Figure 2The flowchart and logic block diagram of a product defect detection method based on the fusion of multiple cameras and deep learning provided in at least one embodiment of the present application should be understood that the method may further include additional blocks not shown and / or may omit the shown blocks, and the scope of the present application is not limited in this regard.
[0063] At step 101, panoramic videos and images of the product are obtained through multiple cameras, and a fusion algorithm of the Scale-Invariant Feature Transform (SIFT) operator and the Oriented FAST and Rotated BRIEF (ORB) operator for multi-view vision matching is used to synthesize the video streams collected by the cameras in the same frame to form a synthesized video containing multi-view information of the product.
[0064] Specifically, multiple industrial cameras are started, and the system timestamp is synchronized to each industrial camera. The video streams of the product are synchronously collected by multiple cameras. By selecting the SIFT operator of the Scale-Invariant Feature Transform and the ORB operator for vision matching, the collected video streams are synthesized in the same frame to form a synthesized video containing multi-view information of the product. During the synthesis process, it is ensured that the video frames of multiple views strictly correspond in time so that subsequent processing can accurately associate the product features under different views.
[0065] In some embodiments, before obtaining the panoramic videos and images of the product through multiple cameras, it further includes:
[0066] Calibrating the parameters of multiple cameras, correcting the image distortion of the images collected by the cameras, performing data fusion calibration, video frame extraction calibration, labeled data generation, and data consistency verification on the images from multiple cameras.
[0067] Specifically, before the system is put into use, precise calibration of the camera parameters is carried out. In some embodiments, the camera parameters include internal parameters (such as focal length, principal point coordinates, etc.) and external parameters (such as position, attitude, etc.). In some embodiments, the number of cameras supports 3 or 4: when the number of cameras is 3, they are placed at an angle of 120 degrees to each other in the imaging direction; when the number of cameras is 4, they are placed at an angle of 90 degrees to each other in the imaging direction. To ensure that the images collected from different perspectives can be accurately fused. Image distortion correction is performed on the images collected by the cameras to eliminate image distortion caused by lens distortion and improve image quality. Data fusion calibration is performed on the video data from multiple cameras to establish the correspondence between images from different perspectives and achieve the accurate synthesis of the front view and side view of the product. The video frame extraction calibration specifically includes: calibrating the key area of the product to be recognized in the video according to factors such as the product shape characteristics, camera physical parameters, and installation environment, and at the same time, calibrating the number of video frames to be extracted. Let the number of video frames be n. Generally, the number of video frames extracted is n ∈ {3, 5, 7, 9, 11}. The labeled data generation specifically includes: during the calibration process, adding labeling information to the collected image data to generate a high-quality labeled data set for training and optimizing the defect detection model. The labeling information includes defect type, defect location, defect area size, etc. The data consistency verification specifically includes: performing consistency verification on the data of the same product collected by different cameras to ensure the accuracy and consistency of the data from different perspectives and provide a reliable data basis for subsequent defect detection.
[0068] In some embodiments, a checkerboard plate is used for internal parameter calibration to calculate parameters such as focal length and principal point coordinates; the external parameters are recorded through a calibration tool. The imaging directions of the 3 cameras are perpendicular at 120 degrees, and the linear distance between any two of them is 1500 millimeters; the cv2.calibrateCamera() method of OpenCV is used for internal parameter calculation, and distortion correction is achieved through cv2.undistort(); the geometric relationship calibration of the dual-camera view is achieved through SIFT feature point matching.
[0069] At step 102, based on the dual attention mechanism of the A2-Nets operator, key frame images related to product defect detection are screened out from the video stream of the synthesized video, and the number of frames of the same product is reduced to the calibrated minimum value.
[0070] Specifically, the A2-Nets (Double Attention Networks) operator is activated to analyze the video stream using its dual attention mechanism. The operator intelligently filters out the key frames related to product defect detection from the video stream, reducing the number of frames for the same product to a calibrated minimum value (usually around 5 frames). Specifically, the operator focuses on key information such as the feature changes of the product in the frame and the appearance of abnormal areas, thereby achieving precise frame selection. In this way, the computational amount can be geometrically reduced, the processing efficiency of the system can be improved, and at the same time, it is ensured that the extracted frames can effectively represent the key features of the product. The key frames obtained through screening and extraction will be transmitted to the subsequent defect detection module.
[0071] At step 103, based on the Yolo operator, comprehensively analyze the size and surface features of the product in the key frame images. If any one of the key frame images of the same product is identified as defective, then the product is initially inspected and determined to be a defective product.
[0072] Specifically, receive all the key frame images and input these key frame images into the size detection module and the surface detection module respectively. The size detection module, based on the Yolo operator, detects the external shape and size of the product to determine whether it meets the preset standard size range. This module will identify the main contour of the product and compare it with the standard size data to calculate the size deviation. The surface detection module, also based on the Yolo operator, is used to detect the features of the product surface, such as defects like scratches, dents, and foreign objects. This module will analyze information such as the texture, color, and shape in the frame image to identify the surface abnormal areas.
[0073] Integrate the detection results of the size detection module and the surface detection module to make a defect determination for each key frame image. If any one of all the key frame images is determined to be defective, then the initial inspection result is defective; otherwise, it is determined to be non-defective. Record the initial inspection result and transmit the information of the product determined to be defective to the defect refined inspection module.
[0074] At step 104, based on the Swin Transformer operator, perform integrated weighted voting on all the key frame images of the product initially inspected and determined to be defective in two dimensions: external shape size detection and appearance surface detection. If the voting result in any one of the dimensions of external shape size detection and appearance surface detection is defective, then the corresponding product is finally determined to be defective; otherwise, it is determined to be non-defective.
[0075] Specifically, for products initially inspected and determined to be defective, the defect refined inspection module is activated. The defective product information passed from the initial inspection module, including all key-frame images, is input into the defect refined inspection module. The defect refined inspection module further precisely detects these key-frame images based on the Swin Transformer operator. The Swin Transformer operator can capture more subtle defect features, such as tiny cracks, complex surface texture changes, etc.
[0076] Different from the initial inspection module, the refined inspection determination integrates and votes separately for the results of all key-frame images extracted from the same product in two dimensions: shape and size inspection and appearance surface inspection. If any result of the shape and size inspection or the appearance surface inspection is defective, it is determined that the final product is defective; otherwise, it is determined that there is no defect.
[0077] Record the refined inspection results, and verify and supplement the initial inspection results to improve the accuracy of defect detection.
[0078] In some embodiments, after the final determination that the corresponding product is defective, it further includes:
[0079] Transmit the defect data detected in the shape and size inspection and the appearance surface inspection to other systems or devices through the network or signals. The defect data includes information such as defect type (scratches, dents, dimensional deviations, etc.), defect location (specific coordinates on the product, which can be represented by pixel coordinates), defect area size (the size or area of the defect area, which can be represented by pixel area or size), defect qualification rate (statistical by day, week, month), etc.
[0080] Specifically, these output data can be used in the production reporting system to generate quality reports; trigger an alarm to remind relevant personnel to handle defective products in a timely manner; or transmit to the quality control center for real-time monitoring of the production process and quality improvement.
[0081] In some embodiments, the detection results are transmitted to the manufacturing execution system MES through the Modbus TCP protocol, and the robotic arm is triggered to remove and offline the defective products. At the same time, the detection process images and results can also be synchronously viewed on the quality inspection large screen.
[0082] Compared with the prior art, the present application has the following beneficial effects:
[0083] Advantages of multi-view video fusion:
[0084] Comprehensive capture of three-dimensional features: By placing three or four industrial cameras at different angles, it is possible to simultaneously obtain the front and side videos of the product, forming a synthetic video containing multi-perspective information. This multi-perspective fusion method captures the three-dimensional features of the product more comprehensively than single-perspective video acquisition, providing a richer data basis for defect detection and helping to improve the accuracy of detection.
[0085] Improve the detection coverage: In traditional single-perspective detection, some defects may be blocked or difficult to observe due to perspective limitations. However, the dual-camera system of this application can observe the product from different angles, effectively expanding the detection coverage and reducing the possibility of missed detections.
[0086] Innovation in video frame extraction:
[0087] Intelligent screening of key frames: Introduce the dual attention mechanism of the A2-Nets (Double Attention Networks) operator, which can intelligently screen out the key frames related to product defect detection from a large number of video frames. This mechanism can accurately focus on the feature changes and abnormal areas in the frames, reducing the number of frames for the same product to an extremely small value, such as about 5 frames. Significantly reduce the computational volume: By greatly reducing the number of frames to be processed, this application can geometrically reduce the computational volume. This not only improves the processing efficiency of the system but also reduces the demand for computing resources, enabling the defect detection system to complete the detection task in a shorter time.
[0088] Improve the accuracy of defect detection:
[0089] Combination of initial inspection and precise inspection: The defect initial inspection module is built based on the Yolo operator and can quickly screen out potential defective products and initially detect defective products. The defect precise inspection module is based on the Swin Transformer operator. Although the detection speed is relatively slow, the detection effect is more accurate. This combination of initial inspection and precise inspection not only ensures the detection efficiency but also improves the detection accuracy.
[0090] Comprehensively analyze size and surface features: In the defect initial inspection module, the size detection module and the surface detection module work together to comprehensively analyze the shape, size, and surface features of the product. Especially in the determination of the shape and size, when the distance between the camera and the product is within 5 meters and the speed of the product assembly line or conveyor belt is within 5 m / s, it is possible to achieve sub-millimeter-level detection and determination of the shape and size.
[0091] Optimize the detection and inference performance:
[0092] Efficient Frame Extraction Improves Inference Speed: The efficient frame selection mechanism of the video frame extraction module reduces the number of frames that need to be inferred, thus accelerating the inference speed of the entire defect detection process. This is particularly important for real-time online detection systems and can meet the requirements for detection speed on fast production lines.
[0093] Balancing Speed and Precision: Although the detection speed of the defect fine inspection module is relatively slow, its precision is higher, and it can further verify and supplement the detection results based on the initial inspection. Overall, while ensuring the detection accuracy, this application optimizes the detection inference performance by optimizing the collaborative work of each module.
[0094] Cost Control and Economic Benefits:
[0095] Reducing Hardware Costs: Compared with multi-view detection systems that require multiple cameras or complex sensors, this application can be achieved by only using three industrial cameras in the hundreds of yuan range + a computer in the thousands of yuan range, reducing the investment cost of hardware equipment.
[0096] Reducing Labor Costs: A highly automated defect detection system can replace a large amount of manual inspection work, reducing the dependence on manual inspection personnel and lowering labor costs. At the same time, the system can provide real-time feedback on detection results, facilitating the timely handling of defective products and avoiding greater losses caused by defective products flowing into subsequent production links.
[0097] Real-time Feedback and Data Consistency:
[0098] Real-time Data Output: The data output module can transmit the detected defect data, such as defect type, location, size, etc., to other systems or devices in real time through the network or signal. This real-time feedback mechanism helps to adjust production management and quality control in a timely manner, improving the flexibility and response speed of production.
[0099] Data Consistency Verification: In the data calibration module, consistency verification is performed on the data of the same product collected by different cameras to ensure the accuracy and consistency of the data from different perspectives. This provides a reliable data basis for subsequent defect detection and avoids detection errors caused by inconsistent data.
[0100] In summary, this application has significant beneficial effects in aspects such as multi-view video fusion, video frame extraction, defect detection accuracy, detection inference performance, cost control, and real-time feedback and data consistency. It can effectively solve the problems existing in the existing defect detection technology and provide a more efficient, accurate, and economical solution for product defect detection in industrial production.
[0101] Reference Figure 3, an embodiment of the present application also provides a product defect detection system based on the fusion of multi-cameras and deep learning, including:
[0102] A data calibration module, which is used to calibrate camera parameters, correct image distortion of the images collected by the cameras, perform data fusion calibration, video frame extraction calibration, labeled data generation, and data consistency verification on the images from multiple said cameras;
[0103] A video acquisition module, which is used to obtain panoramic videos and images of the product through multiple cameras, and use a fusion algorithm of the Scale-Invariant Feature Transform (SIFT) operator and the Oriented FAST and Rotated BRIEF (ORB) operator for multi-view vision matching to synthesize the video streams collected by the cameras in the same frame to form a synthesized video containing multi-view information of the product;
[0104] A video frame extraction module, which is used to screen out key frame images related to product defect detection from the video stream of the synthesized video based on the dual attention mechanism of the A2-Nets operator, and reduce the number of frames of the same product to a calibrated minimum value;
[0105] A defect preliminary inspection module, which is used to comprehensively analyze the size and surface features of the product in the key frame images based on the Yolo operator. If any frame in the key frame images of the same product is identified as defective, the product is preliminarily inspected and determined to be a defective product;
[0106] A defect refined inspection module, which is used to perform integrated weighted voting on all key frame images of the product preliminarily inspected and determined to be defective based on the Swin Transformer operator in two dimensions of external dimension detection and appearance surface detection. If the voting result of any dimension of the external dimension detection and the appearance surface detection is defective, the corresponding product is finally determined to be defective; otherwise, it is determined to be non-defective;
[0107] A data output module, which is used to transmit the defect data detected in the external dimension detection and the appearance surface detection to other systems or devices through a network or a signal. The defect data includes defect type, defect location, defect area size, and defect qualification rate.
[0108] The following embodiments take the defect detection of non-ferrous metal handicrafts as an example to explain the system provided by the present application:
[0109] Data calibration module: The video frame extraction calibration is 5 frames per piece (5 frames for each handicraft); calibrate the length, width, height dimensions and allowable errors of the handicrafts; calibrate the surface shape and structure of the handicrafts.
[0110] Video acquisition module: Synchronize the system time to two cameras every hour. Meanwhile, adopt a multi-threaded architecture to achieve synchronous video stream acquisition by the cameras, and synchronize video frames with the help of system timestamps; Based on the fusion algorithm of the Scale-Invariant Feature Transform (SIFT) operator and the multi-view vision matching ORB operator, synthesize the video streams captured by the cameras frame by frame to form a synthesized video stream containing multi-view information of the product.
[0111] Video frame extraction module: Based on the A2-Nets operator, use a dual attention mechanism to screen key frames, ensuring that the most effective key frames containing defect areas are extracted for each handicraft according to the calibrated parameters, and discarding a large number of redundant frames.
[0112] Initial defect detection module: Input all 5 video key frames of each handicraft extracted into the size detection and surface detection module based on the YOLO v10 operator to preliminarily screen defective products. If any frame is determined to be defective in size detection or surface detection, the handicraft is initially determined to be a defective handicraft.
[0113] Defect refinement detection module: Adopt the Swin Transformer v2.0 operator to conduct a detailed analysis of the defective frames screened by the initial detection to capture tiny or complex defect features. Input all 5 video key frames of each handicraft extracted into the size detection module and the surface detection module respectively, and conduct integrated voting on the recognition results of the size detection module and the surface detection module. If the result of the size detection module or the surface detection module is defective, the handicraft is initially determined to be a defective handicraft.
[0114] Data output module: Transmit the detection results to the production management system (MES) and the audible and visual alarm through the TCP / IP protocol or the industrial fieldbus (such as Modbus), and take the defective handicrafts offline in real time.
[0115] Figure 4 This is a module block diagram of a product defect detection device based on the fusion of multi-cameras and deep learning provided by at least one embodiment of the present application. The device includes:
[0116] A memory 201; and a processor 202 connected to the memory 201, the processor 202 being configured to: obtain panoramic videos and images of the product through multiple cameras, and use the fusion algorithm of the Scale-Invariant Feature Transform (SIFT) operator and the multi-view vision matching ORB operator to synthesize the video streams captured by the cameras frame by frame to form a synthesized video containing multi-view information of the product;
[0117] Based on the dual attention mechanism of the A2-Nets operator, screen out key frame images related to product defect detection from the video streams of the synthesized video, and reduce the number of frames under the same product to a calibrated minimum value;
[0118] Based on the Yolo operator, comprehensively analyze the size and surface features of the products in the key frame images. If any frame in the key frame images of the same product is identified as defective, then initially determine that the product is a defective product.
[0119] Based on the Swin Transformer operator, for all the key frame images of the products initially determined to be defective, perform integrated weighted voting separately in two dimensions: external dimension detection and appearance surface detection. If the voting result in any dimension of the external dimension detection and appearance surface detection is defective, then finally determine that the corresponding product is defective; otherwise, determine that it is not defective.
[0120] In some embodiments, the processor 202 is further configured to: before obtaining the panoramic video and images of the products through multiple cameras, further include:
[0121] Calibrate the camera parameters, correct the image distortion of the images collected by the cameras, and perform data fusion calibration, video frame extraction calibration, labeled data generation, and data consistency verification on the images from multiple cameras.
[0122] In some embodiments, the processor 202 is further configured to: the camera parameters include internal parameters and external parameters. The internal parameters include focal length and principal point coordinates, and the external parameters include position and attitude.
[0123] In some embodiments, the processor 202 is further configured to: the video frame extraction calibration specifically includes: calibrate the key area of the identified product in the video according to the external shape characteristics of the product, the physical parameters of the camera, and the installation environment factors, and at the same time, calibrate the number of video frames to be extracted.
[0124] In some embodiments, the processor 202 is further configured to: the labeled data generation specifically includes: during the calibration process, add labeled information to the collected image data to generate a labeled data set. The labeled information includes defect type, defect location, and defect area size.
[0125] In some embodiments, the processor 202 is further configured to: the data consistency verification specifically includes: perform consistency verification on the data of the same product collected by different cameras.
[0126] In some embodiments, the processor 202 is further configured to: after finally determining that the corresponding product is defective, further include:
[0127] Transmit the defect data detected in the above-mentioned external dimension detection and appearance surface detection to other systems or devices through a network or signals. The defect data includes defect type, defect location, defect area size, and defect pass rate.
[0128] For the specific implementation methods of the system and device, refer to the foregoing method embodiments, which will not be elaborated herein.
[0129] This application can be a method, a device, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of this application.
[0130] A computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0131] The computer-readable program instructions described herein can be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0132] The computer program instructions for performing the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of this application.
[0133] Aspects of the present application are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0134] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that when these instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, and these instructions cause a computer, a programmable data - processing apparatus, and / or other devices to work in a specific manner. Thus, the computer - readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0135] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0136] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0137] Note that unless otherwise directly stated, all features disclosed in this specification (including any appended claims, abstract, and drawings) may be replaced by alternative features serving the same, equivalent, or similar purpose. Therefore, unless otherwise explicitly stated, each feature disclosed is only an example of a set of equivalent or similar features. In cases where it is used, further, preferably, furthermore, and more preferably are simply the starting points for elaborating another embodiment based on the foregoing embodiments. The content following the further, preferably, furthermore, or more preferably in combination with the foregoing embodiments constitutes the complete composition of another embodiment. Combinations can be made arbitrarily among several further, preferably, furthermore, or more preferably settings following the same embodiment to form yet another embodiment.
[0138] Although the present application has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it on the basis of the present application, which will be obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present application fall within the scope of protection required by the present application.
Claims
1. A product defect detection method based on multi-camera and deep learning fusion, characterized in that: The following steps are involved: The panoramic video and image of the product are obtained through multiple cameras. The fusion algorithm of SIFT operator based on scale-invariant feature transformation and ORB operator based on multi-view matching is used to synthesize the video streams collected by the cameras in the same frame to form a synthetic video containing multi-view information of the product. Based on the dual attention mechanism of the A2-Nets operator, key frame images related to product defect detection are screened out from the video stream of the synthetic video, and the number of frames under the same product is reduced to a calibrated minimum value; Based on the Yolo operator, the size and surface features of the product in the key frame image are comprehensively analyzed. If any frame in the key frame image of the same product is identified as defective, the product is initially inspected and determined to be a defective product; Based on the Swin Transformer operator, all key frame images of the products judged as defective in the initial inspection are subjected to integrated weighted voting according to the two dimensions of shape dimension detection and appearance surface detection. If the voting result of any dimension of shape dimension detection and appearance surface detection is defective, the corresponding product is finally judged to be defective. Otherwise, it is judged to be not defective.
2. The product defect detection method based on multi-camera and deep learning fusion according to claim 1 is characterized in that: Before obtaining the panoramic video and image of the product through multiple cameras, it also includes: The camera parameters are calibrated, image distortion correction is performed on the images captured by the camera, and data fusion calibration, video frame extraction calibration, annotation data generation, and data consistency verification are performed on the images from multiple cameras.
3. The product defect detection method based on multi-camera and deep learning fusion according to claim 2 is characterized in that: The camera parameters include internal parameters and external parameters, the internal parameters include focal length and principal point coordinates, and the external parameters include position and posture.
4. The product defect detection method based on multi-camera and deep learning fusion according to claim 2 is characterized in that: The video frame extraction calibration specifically includes: calibrating the key areas of the identified product in the video according to the product appearance characteristics, camera physical parameters and installation environment factors, and at the same time, calibrating the number of extracted video frames.
5. The product defect detection method based on multi-camera and deep learning fusion according to claim 2 is characterized in that: The generation of the annotation data specifically includes: in the calibration process, adding annotation information to the collected image data to generate an annotation data set, wherein the annotation information includes defect type, defect location, and defect area size.
6. The product defect detection method based on multi-camera and deep learning fusion according to claim 2 is characterized in that: Data consistency verification specifically includes: consistency verification of the same product data collected by different cameras.
7. The product defect detection method based on multi-camera and deep learning fusion according to claim 1 is characterized in that: After the final determination that the corresponding product is defective, the method further includes: The defect data detected in the shape dimension inspection and appearance surface inspection are transmitted to other systems or devices through a network or signal, and the defect data includes defect type, defect location, defect area size and defect qualification rate.
8. A product defect detection system based on multi-camera and deep learning fusion, characterized in that: include: The video acquisition module is used to obtain panoramic videos and images of products through multiple cameras, and synthesize the video streams collected by the cameras into the same frame by using the fusion algorithm of SIFT operator based on scale-invariant feature transformation and ORB operator based on multi-view vision matching to form a synthetic video containing multi-view information of the product; A video frame extraction module, which is used to filter out key frame images related to product defect detection from the video stream of the synthetic video based on the dual attention mechanism of the A2-Nets operator, and reduce the number of frames under the same product to a calibrated minimum value; A defect initial inspection module is used to comprehensively analyze the size and surface features of the product in the key frame image based on the Yolo operator. If any frame in the key frame image of the same product is identified as defective, the product is initially inspected and determined to be a defective product; The defect inspection module is used to perform integrated weighted voting on all key frame images of the product judged as defective in the initial inspection according to the two dimensions of shape size detection and appearance surface detection based on the Swin Transformer operator. If the voting result of any dimension of shape size detection and appearance surface detection is defective, the corresponding product is finally judged to be defective, otherwise, it is judged to be not defective.
9. A product defect detection device based on multi-camera and deep learning fusion, characterized in that: include: Memory; as well as A processor connected to the memory, the processor being configured to perform the steps of the method according to any one of claims 1 to 7.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a machine, the steps of the method according to any one of claims 1 to 7 are implemented.