Multi-modal fusion automobile door panel visual defect detection system and detection method

The multimodal fusion automotive door panel visual defect detection system utilizes a multi-camera array, adaptive light source, LiDAR, and ultrasonic sensors combined with a deep learning model to achieve high-speed, accurate, and full-area detection of automotive door panels. This solves the problems of low efficiency, poor accuracy, and weak adaptability in existing technologies, thereby improving detection efficiency and accuracy.

CN121324352APending Publication Date: 2026-01-13JIANGSU LINQUAN AUTO DECORATIONS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511697587.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing automotive door panel vision inspection systems are inefficient, inaccurate, unresponsive, and limited in their inspection dimensions. They cannot perform rapid and accurate inspection of multiple doors and require repeated disassembly for front and back inspection.

Method used

The automotive door panel visual defect detection system employing multimodal fusion includes a data acquisition module, a data processing module, a decision output module, and a feedback optimization module. It acquires multi-source data through a multi-camera array, an adaptive multispectral light source, LiDAR, and an ultrasonic sensing unit, and combines deep learning models and multimodal data fusion technology to achieve full-domain detection.

Benefits of technology

It achieves high-speed, accurate, and full-area inspection of automotive door panels, shortening the inspection cycle to within 20 seconds, achieving a defect identification accuracy rate of 99.2%, and reducing the material replacement adaptation time to 3 seconds. It requires no manual intervention, reducing inspection costs and operational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121324352A_ABST
    Figure CN121324352A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal fusion automobile door panel visual defect detection system and detection method. The system comprises a data acquisition module, a data processing module, a decision output module and a feedback optimization module which are in communication connection in sequence to form a closed-loop detection system. The data acquisition module adopts a multi-camera array, a self-adaptive multispectral light source, a laser radar and a multi-source acquisition architecture of ultrasonic sensing, and is matched with the high-speed transmission unit to realize high-fidelity data acquisition; an improved DoorNet deep learning model, a multi-modal data fusion unit and a parameter self-optimization unit are arranged in the data processing module, and accurate defect recognition and dynamic scene adaptation are achieved; the decision output module is linked with the production line PLC system to complete defect interception, and the feedback optimization module realizes continuous improvement of system performance through incremental training. The method solves the problems of low efficiency, poor precision and weak adaptability in the prior art, and is suitable for global defect detection of the high-speed automobile production line.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of automotive door panel inspection technology, specifically, it relates to a multimodal fusion automotive door panel visual defect detection system and method. Background Technology

[0002] As living standards improve and cars become more widespread, people have a more detailed understanding of automobiles. As a representative of modern assembly line production, car door production is characterized by highly automated operations. After the car door is produced, its front and back surfaces need to be inspected. Traditional manual assembly and inspection can lead to many loopholes, resulting in misjudgments and missed judgments. This causes defective products to enter the market, which has a significant impact on brand reputation. In order to improve production efficiency, reduce the defect rate, and lower labor costs, some companies have started to use visual inspection devices to complete the defect detection of car doors.

[0003] Referring to the visual inspection system for defects in car door interiors disclosed in patent application CN109187561A, this visual inspection system uses machine vision methods to inspect car door interiors. It is fast, has a low error rate, and can eliminate the large economic losses caused by various misjudgments and omissions in manual inspection during long-term operation. It is not only applicable to the inspection of interiors of a certain type of car door, but can be reused on the same platform to save costs. It adopts fixed-position visual inspection, which can overcome the unsatisfactory inspection results caused by movement on the conveyor belt.

[0004] The aforementioned visual inspection system can effectively detect defects in car doors through its image acquisition device. However, this system can only fix a single car door in a single inspection and cannot quickly fix multiple car doors before inspection. Therefore, after a single inspection, the system needs to be stopped to load and unload materials before inspection can continue, resulting in low inspection efficiency and time-consuming and labor-intensive operation, which is not conducive to the continuous and stable operation of inspection work. Meanwhile, the aforementioned visual inspection system cannot switch between the front and back of the car door during inspection, which greatly limits the inspection work. Therefore, when it is necessary to inspect the front and back of the car door, it is necessary to repeatedly disassemble the car door to complete the switching inspection, which further increases the workload of inspection. Summary of the Invention

[0005] In view of this, the technical problem to be solved by the present invention is to provide a multimodal fusion automotive door panel visual defect detection system and detection method. Through the collaborative design of customized hardware upgrades and deep algorithm innovation, the system achieves high-speed, accurate, and full-domain detection of automotive door panel defects, meets the intelligent detection needs of modern automotive production lines, and overcomes the shortcomings of existing automotive door panel defect detection technologies, such as low efficiency, poor accuracy, weak adaptability, and single detection dimension.

[0006] To address the aforementioned technical issues, this invention discloses a multimodal fusion-based visual defect detection system for automotive door panels, comprising a data acquisition module, a data processing module, a decision output module, and a feedback optimization module. Each module is sequentially connected via industrial Ethernet and fiber optic transmission links to form a closed-loop detection system for acquisition, processing, decision-making, and optimization.

[0007] The data acquisition module, serving as the system's "perception layer," is responsible for acquiring surface images, 3D contours, and internal structural data of the automotive door panel. Its core comprises a multi-camera array unit, an adaptive multispectral light source unit, a lidar unit, an ultrasonic sensing unit, and a high-speed transmission unit. The multi-camera array unit uses three 80-megapixel CMOS industrial cameras arranged linearly along the length of the automotive door panel. Adjacent cameras have a 10-15% overlap in their detection areas to ensure no blind spots. Each camera is equipped with a 24mm fixed-focus telecentric lens and a global shutter. The telecentric lens effectively eliminates perspective distortion, while the global shutter adaptively adjusts within a shutter speed range of 1 / 500s to 1 / 2000s, avoiding image blurring caused by door panel movement on high-speed production lines. Ultimately, this achieves an image resolution of 0.01mm / pixel and a frame rate output of 60fps, providing high-fidelity image data for capturing minute defects.

[0008] The adaptive multispectral light source unit is the core component for resolving imaging interference from various door panels. It includes a ring-shaped LED main light source, an adjustable (0-45°) strip auxiliary light source, and an infrared supplementary lighting module, while also incorporating a material identification sensor based on spectral analysis. The unit's operating logic is as follows: the material identification sensor first acquires the spectral reflectance curve of the door panel surface. The acquired spectral data is then matched with light source parameter templates for eight common door panel materials (metal, ABS, slush molding, fabric covering, etc.) pre-set by the system to automatically determine the optimal light source combination. For example, when inspecting a matte black slush molding door panel, the system automatically adjusts the ring light source brightness to 500 lux and the strip light source angle to 30° to enhance the shadow characteristics of dent defects. When inspecting a chrome-plated metal door panel, the system turns off the strip direct light source and activates the polarizer filter of the ring light source to filter reflections, retaining only diffused light for imaging, ensuring that high-contrast defect images are obtained for door panels of different materials.

[0009] The lidar unit employs line laser scanning technology, achieving a depth measurement accuracy of 0.02mm. It acquires three-dimensional contour data of the door panel surface, accurately quantifying the depth and height parameters of defects such as dents and protrusions. The ultrasonic sensing unit emits 200kHz ultrasonic signals into the door panel and detects structural defects such as weld quality and reinforcing rib integrity based on the difference in echo signal intensity and propagation time. The high-speed transmission unit serves as the "highway" for data flow, employing a three-tier transmission architecture consisting of a high-speed image acquisition card with a Camera Link HS interface, a 128GB DDR4 high-speed cache, and EtherCAT industrial Ethernet. The Camera Link HS interface boasts a transmission rate of 10Gbps, controlling the image transmission delay of a single camera to within 15ms. The high-speed cache enables parallel storage and preprocessing of multi-source data, while the EtherCAT bus ensures real-time data transmission to the back-end processing module. The end-to-end data transmission latency is ≤50ms, meeting the real-time requirements of high-speed production lines.

[0010] The data processing module, as the "core computing layer" of the system, is responsible for defect feature extraction, multi-source data fusion, and self-optimization of detection parameters. It has a built-in improved deep learning model DoorNet, a multimodal data fusion unit, and a parameter self-optimization unit. The DoorNet model is a deep learning architecture custom-developed for door panel defect recognition. It innovatively integrates a spatial attention module (SAM), a channel attention module (CAM), and a multi-scale feature extraction network. The spatial attention module and the channel attention module learn weights during training, enabling the model to automatically focus on potential defect areas on the door panel surface and weaken the influence of interfering textures such as seams and decorative strips. The multi-scale feature extraction network consists of 4 convolutional pooling layers and 2 deconvolutional layers, extracting image features at three scales: 16×16, 32×32, and 64×64, respectively, achieving full coverage recognition of different types of defects such as tiny scratches (small scale) with a width of 0.05mm, large-area dents (large scale), and thin cracks (medium scale). To solve the problem of imbalanced defect samples (many samples of slight scratches and few samples of deep dents), the model adopts a combined loss function of "cross-entropy loss + Dice loss". Cross-entropy loss ensures the model's classification accuracy for common defects, while Dice loss improves the model's sensitivity to rare defects. The DoorNet model is trained using an efficient strategy of "transfer learning + incremental training". First, it is pre-trained based on the publicly available NEU-DET steel defect dataset to obtain basic defect feature extraction capabilities. Then, it is fine-tuned using 50,000 labeled samples covering 12 types of door panel defects and 8 materials to adapt the model to the car door panel detection scenario, ultimately achieving an average defect recognition accuracy of 99.2%.

[0011] The multimodal data fusion unit is key to achieving full-domain detection of "surface-internal-3D". Its fusion process consists of three core steps: The first step is data synchronization, which uses high-precision timestamp alignment technology to achieve millisecond-level synchronization of visual images (2D surface information), LiDAR point clouds (3D contour information), and ultrasonic echo signals (1D internal structure information), ensuring accurate matching of multi-source data at the same detection location. The second step is feature-level fusion, which extracts the 3D contour features of the LiDAR (such as the depth value and slope change of the defect area), the internal structure features of the ultrasonic waves (such as the echo intensity of the weld point and the density change of the substrate), and the surface texture features of the visual images (such as the grayscale gradient of scratches and the boundary features of bubbles), and then integrates these three types of features according to their dimensions. The first step involves stitching together multi-dimensional feature vectors to achieve complementary defect information. The second step is decision-level fusion, which uses a weighted voting method to fuse the defect identification results of each modality. Weights are assigned based on the advantages of different modalities in various defect detections (visual modality has a weight of 0.4 for surface scratches, LiDAR has a weight of 0.4 for indentation depth, and ultrasound has a weight of 0.5 for internal defects). For example, when the visual modality determines "surface anomaly exists", LiDAR detects "abnormal area depth of 0.3mm", and ultrasound determines "normal internal structure", the system outputs a comprehensive and accurate conclusion of "surface indentation (depth of 0.3mm)". If the ultrasound modality detects "abnormal attenuation of echo signal", it is determined to be a composite defect type of "internal weld defect + surface indentation".

[0012] The parameter self-optimization unit enables the system to adapt to dynamic scenes. Based on reinforcement learning algorithms, it constructs a parameter adjustment model with the optimization goal of "false detection rate <1% and false negative rate <0.5%", automatically adjusting 12 key detection parameters such as camera exposure time, light source brightness, and model inference threshold. The system has a built-in scene perception mechanism. By analyzing indicators such as the average brightness, texture complexity, and defect distribution characteristics of the detected image in real time, it determines whether the current detection scene has changed. When the brightness fluctuation of the detected image exceeds ±20%, or the matching degree between the texture features and the preset template is less than 85%, the adaptive adjustment process is immediately triggered. The parameter adjustment response time is ≤3 seconds. It can adapt to scenarios such as door panel material changes, process adjustments, or changes in ambient lighting without manual intervention, solving the problem of "false detection when the scene changes" in traditional systems.

[0013] The decision output module is responsible for converting the inspection results into production-usable information. Its core functions include: First, quantitative output of defect information, accurately labeling the defect type (such as scratches, bubbles, and poor weld joints), two-dimensional position (X / Y coordinates based on the door panel coordinate system), three-dimensional parameters (depth, height, accuracy ±0.02mm), and defect level (divided into four levels: A / B / C / D according to GB / T 23978-2009 standard); Second, production line linkage control, feeding back the defect judgment results to the production line PLC system in real time. If a C-level or higher serious defect is detected, a door panel interception signal is immediately triggered, controlling the conveyor line to transfer the defective door panel to the rework area, realizing an automated closed loop of "inspection-interception-rework"; Third, inspection data storage, classifying and storing the inspection report, original images, and multi-source data of each door panel according to the rule of "vehicle model-batch-serial number", providing data support for subsequent quality traceability and process optimization.

[0014] The feedback optimization module is the core of the system's "continuous evolution," comprising an incremental data pool and a lightweight training unit. The incremental data pool, linked with the manual review workstation, automatically collects and stores manually confirmed "misjudged samples" (such as misclassifying textures as defects) and "missed samples" (such as undetected micro-pinholes) during the detection process, and automatically labels the samples (labeling defect type, location, and manual judgment results). The lightweight training unit sets trigger conditions to automatically initiate the model optimization process every 500 incremental samples. To avoid excessively long full-scale training times, this unit only updates the parameters of the DoorNet model's fully connected layers and attention modules, keeping training time within 5 minutes, enabling real-time online model updates and continuously improving the system's detection accuracy over time.

[0015] Based on the above system, this invention also discloses a multimodal fusion method for detecting visual defects in automotive door panels, comprising the following steps: Step S10: Data Acquisition Preparation. After the car door panel enters the inspection area via the conveyor line, the production line's photoelectric sensor sends a position signal. The multi-camera array unit, lidar unit, and ultrasonic sensing unit complete collaborative positioning based on the three-dimensional dimensions of the door panel (preset with a parameter library for different car models), ensuring that the inspection range completely covers the entire door panel area. At the same time, the material recognition sensor of the adaptive multispectral light source unit collects the spectral data of the door panel surface, and after matching it with the preset template, completes the parameter initialization of light source type, brightness, and angle, preparing for high-quality imaging.

[0016] Step S20: Dynamic Image Acquisition. The conveyor line drives the door panel through the detection area at a constant speed of 0.5m / s. The multi-camera array unit starts acquiring data synchronously, with each camera responsible for a detection area of ​​40-60cm. The multi-view images are fused into a complete full-area image of the door panel using an image stitching algorithm based on SIFT feature point matching, with a stitching error ≤1 pixel. During this process, the lidar unit simultaneously completes the 3D contour scanning of the door panel, and the ultrasonic sensing unit acquires internal structural data at a density of 500 points / second. The high-speed transmission unit transmits the acquired images, point clouds, and ultrasonic data to the data processing module in real time, with an image transmission delay of ≤15ms for a single camera, ensuring data real-time performance.

[0017] Step S30: Defect Feature Processing. After receiving the data, the data processing module first performs preprocessing on the image, such as noise reduction and ROI (Region of Interest) extraction, to remove background interference. Then, the DoorNet model performs feature extraction and defect identification on the preprocessed image, outputting preliminary judgment results of surface defects. The multimodal data fusion unit performs feature fusion and decision fusion with the DoorNet recognition results, the 3D data of the LiDAR, and the internal data of the ultrasound, to complete the accurate determination of defect type and parameters. The DoorNet model has a recognition rate of ≥95% for 0.05mm wide scratches and a recognition rate of ≥93% for bubbles of the same color as the door panel. The LiDAR's measurement error for defect depth is controlled within ±0.02mm.

[0018] Step S40: Output of Inspection Results. The decision output module generates a standardized inspection report from the fused defect information and transmits it to the production line PLC system and the workshop MES system via industrial Ethernet. If a qualified door panel is detected, the conveyor line continues to the next process. If a defect of grade C or above is detected, the PLC system controls the pneumatic baffle to rise, intercepting the defective door panel and transferring it to the rework area. At the same time, the MES system records the defect information for quality traceability.

[0019] Step S50: System self-optimization. The manual review workstation manually confirms the intercepted defective door panels and the "suspected defective" door panels marked by the system. The confirmed misjudged and missed samples are automatically stored in the incremental data pool of the feedback optimization module. When the number of samples in the incremental data pool reaches 500, the lightweight training unit starts updating the model parameters, completing one system self-optimization cycle, so that the model detection performance is continuously improved.

[0020] Through the above steps, the system can complete the inspection cycle of a single door panel in ≤20 seconds, which can meet the needs of a high-speed production line with a speed of 3-4 pieces per minute.

[0021] Compared with the prior art, the present invention can achieve the following technical effects: With the help of multi-camera arrays and high-speed transmission architecture, the inspection cycle of a single door panel is reduced to less than 20 seconds, which meets the needs of high-speed production lines. The application of DoorNet model and multimodal fusion technology enables the average defect recognition accuracy to reach 99.2%, and the recognition rate of small scratches and bubbles of the same color is more than 93%. The false negative rate and false positive rate are controlled below 0.5% and 1% respectively, which far exceeds the traditional inspection technology.

[0022] The adaptive multispectral light source unit enables automatic adaptation to 8 common door panel materials, and the parameter self-optimization unit reduces the system's adaptation time to material changes and environmental changes from 2 hours to 3 seconds. It can maintain stable detection performance without manual intervention, reducing reliance on the professional skills of operators.

[0023] By integrating visual, lidar, and ultrasonic multimodal data, it achieves full-domain coverage from surface defect identification to internal structure detection and three-dimensional parameter quantification, providing complete data support for automotive door panel quality assessment and overcoming the limitations of traditional technologies that can only examine the surface.

[0024] The feedback optimization module constructs a closed loop of detection, feedback, and training, which continuously improves the system's detection accuracy over time, extends the equipment's life cycle value, and reduces later maintenance costs.

[0025] Of course, any product implementing this invention does not necessarily need to achieve all of the technical effects described above at the same time. Attached Figure Description

[0026] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a module architecture diagram of the improved automotive door panel visual defect detection system of the present invention; Figure 2 This is a schematic diagram of the network structure of the DoorNet model of the present invention; Figure 3 This is a schematic diagram of the multimodal data fusion process of the present invention; Figure 4 This is a flowchart illustrating the steps of the improved visual defect detection method for automotive door panels according to the present invention.

[0027] Figure label: 1-Data acquisition module, 11-Multi-camera array unit, 12-Adaptive multispectral light source unit, 13-LiDAR unit, 14-Ultrasonic sensing unit, 15-High-speed transmission unit, 2-Data processing module, 21-DoorNet model, 22-Multimodal data fusion unit, 23-Parameter self-optimization unit, 3-Decision output module, 4-Feedback optimization module, 41-Incremental data pool, 42-Lightweight training unit. Detailed Implementation

[0028] The following will describe in detail the implementation of the present invention with reference to the accompanying drawings and embodiments, so that the process of how the present invention uses technical means to solve technical problems and achieve technical effects can be fully understood and implemented accordingly.

[0029] like Figure 1 As shown, this embodiment discloses an improved visual defect detection system for automotive door panels, including a data acquisition module 1, a data processing module 2, a decision output module 3, and a feedback optimization module 4. Each module is connected to the system via EtherCAT industrial Ethernet and fiber optic transmission links to achieve gigabit-level communication, forming a closed-loop detection system.

[0030] The specific configuration of data acquisition module 1 is as follows: The multi-camera array unit 11 uses three Basler acA2500-14gc 80-megapixel CMOS industrial cameras, deployed in an array along the length of the car door panel 5 (1.8m door panel), with an adjacent camera spacing of 50cm and a 12% overlap in the detection area. Each camera is equipped with a Kowa LM24HC fixed-focus telecentric lens and a global shutter. The shutter speed is set to 1 / 1000s according to the door panel's movement speed, achieving a resolution of 0.01mm / pixel and a frame rate of 60fps. The adaptive multispectral light source unit 12 uses a CCS VLG series ring LED main light source, a strip auxiliary light source, and an infrared fill light module, with a built-in Ocean Optics USB4000 spectral sensor and preset light source parameter templates for eight types of materials. The lidar unit 13 uses a SICK LMS511-20100 model line lidar with a measurement range of 0.1-20m and a depth accuracy of ±0.02mm. The ultrasonic sensing unit 14 uses a Panasonic... The EVM30A sensor operates at a frequency of 200kHz. The high-speed transmission unit 15 is equipped with an NI PCIe-1435 Camera Link HS acquisition card and a 128GB DDR4 3200MHz high-speed cache. It communicates with the data processing module 2 via the EtherCAT bus, and the measured end-to-end latency is 42ms.

[0031] refer to Figure 2 and Figure 3Data processing module 2 is built based on Intel Xeon Gold 6330 processors and NVIDIA A100 GPUs to construct computing nodes. The training process of DoorNet model 21 is as follows: First, it is pre-trained on the NEU-DET dataset for 100 epochs with a learning rate of 0.001 and a batch size of... The size was set to 32. Subsequently, 50,000 door panel labeled samples (covering 12 types of defects such as scratches, bubbles, and dents, and including 8 materials such as metal and molten plastic) were used to fine-tune the model for 80 epochs, with the learning rate reduced to 0.0001. The "cross-entropy loss + Dice loss" combined function was adopted, and the final model achieved an average recognition accuracy of 99.3% on the test set. The weights of the multimodal data fusion unit 22 were set as follows: visual modality 0.4, LiDAR 0.4, and ultrasound 0.5. Feature stitching and weighted voting were implemented using Python's NumPy library. The parameter self-optimization unit 23 was implemented based on the reinforcement learning framework PPO, with "false positive rate <1% and false negative rate <0.5%" as the reward function. The parameters adjusted included 12 items such as camera exposure time (5-20ms) and light source brightness (300-1000 lux). The actual response time was 2.8 seconds.

[0032] The decision output module 3 uses LabVIEW to develop a human-computer interaction interface, which displays the defect type, location (accuracy ±1mm), depth parameters and level in real time. It communicates with Siemens S7-1200 PLC through the Profinet protocol to realize defect door panel interception control. The incremental data pool 41 of the feedback optimization module 4 uses a MySQL database to store samples. The lightweight training unit 42 is implemented based on the transfer learning interface of PyTorch, which only updates the fully connected layer and attention module of the model. The actual training time for 500 samples is 4.5 minutes.

[0033] Based on the above system, the detection method in this embodiment is as follows: Figure 4 As shown, the specific steps are as follows: Step S10: Data Acquisition Preparation. The car door panel is conveyed to the detection area by the conveyor line. The photoelectric sensor triggers the positioning signal. The multi-camera array unit 11, the lidar unit 13, and the ultrasonic sensing unit 14 complete the positioning based on the preset parameters of the B-class sedan door panel (length 1.6m, width 0.8m). The spectral sensor collects the spectral data of the door panel surface, which is matched with the black matte slush plastic material. The adaptive light source unit 12 adjusts the brightness of the ring light source to 500 lux and the angle of the strip light source to 30°.

[0034] Step S20: Dynamic image acquisition. The conveyor line runs at a speed of 0.5 m / s, and multiple camera arrays synchronously acquire images. The images are stitched together using the SIFT algorithm to form a 1.6 × 0.8 m full-area image with a stitching error of 0.8 pixels. The lidar scan acquires three-dimensional contour data, and the ultrasonic sensor acquires internal data at 500 points / second. The high-speed transmission unit 15 transmits the data to the processing module 2 in real time.

[0035] Step S30: Defect feature processing. The preprocessed image is input into the DoorNet model 21 to identify two suspected defects on the door panel surface; the multimodal fusion unit 22 fuses the three-dimensional data (depth 0.25mm, 0.08mm) of the suspected defect area with the ultrasonic data (normal echo) and determines them as "surface depression (0.25mm)" and "surface scratch (0.08mm)".

[0036] Step S40: Output of detection results. The decision module 3 outputs a detection report. Since the defect level is D (minor defect), the conveyor line continues to flow. If a dent with a depth of 1.2mm (B-level defect) is detected, the PLC interception signal is triggered.

[0037] Step S50: System self-optimization. A missed 0.05mm pinhole defect is manually verified and the sample is automatically stored in the incremental data pool 41; when the number of samples reaches 500, the lightweight training unit 42 is activated, and the model's recognition rate of pinhole defects increases from 92% to 95%.

[0038] The actual test data of this embodiment shows that the detection cycle of a single door panel is 18 seconds, the recognition rate of 0.05mm scratches is 95.2%, the recognition rate of bubbles of the same color is 93.7%, the measurement error of defect depth is ±0.02mm, and the system adaptation time when changing materials is 2.8 seconds, which fully meets the detection requirements of high-speed production lines.

[0039] The foregoing description illustrates and describes several preferred embodiments of the present invention. However, as previously stated, it should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the inventive concept described herein through the foregoing teachings or techniques or knowledge in related fields. Any modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A multimodal fusion-based visual defect detection system for automotive door panels, characterized in that, It includes a data acquisition module, a data processing module, a decision output module, and a feedback optimization module that are connected in sequence via communication. The data acquisition module includes: The multi-camera array unit consists of three 80-megapixel CMOS industrial cameras arranged linearly along the length of the car door panel. Each camera is equipped with a 24mm fixed-focus telecentric lens and a global shutter. The system is configured to achieve an image resolution of 0.01mm / pixel and a frame rate of 60fps through this unit. The adaptive multispectral light source unit includes a ring-shaped LED main light source, an angle-adjustable strip auxiliary light source, and an infrared supplementary light module. It also integrates a material identification sensor to collect spectral reflectance data of the door panel surface. The lidar unit is used to acquire the three-dimensional contour point cloud data of the door panel; An ultrasonic sensing unit is used to acquire ultrasonic echo signals inside the door panel. The high-speed transmission unit uses a Camera Link HS interface acquisition card, 128GB DDR4 memory as a high-speed cache, and communicates via EtherCAT industrial Ethernet. The end-to-end data transmission delay of the data acquisition module is configured to be no more than 50ms. The data processing module includes: The improved deep learning model DoorNet integrates a spatial attention module, a channel attention module, and a multi-scale feature extraction network in its network structure, and uses a linear combination of the cross-entropy loss function and the Dice loss function as the loss function. The multimodal data fusion unit is configured to perform feature-level and decision-level fusion of visual images, lidar point clouds, and ultrasonic echo signals from the data acquisition module. The parameter self-optimization unit is used to dynamically adjust the system operating parameters; The decision output module is configured to output the type, location coordinates, and quantification parameters of the defect, and communicates and links with the PLC system of the production line. The feedback optimization module includes an incremental data pool and a lightweight training unit, which are used to implement online updates of the DoorNet model.

2. The multimodal fusion-based visual defect detection system for automotive door panels according to claim 1, characterized in that, In the multi-camera array unit, the field of view detection areas of two adjacent cameras have an overlap of 10% to 15%, and are fused into a complete full-area image of the door panel by an image stitching algorithm based on SIFT feature point matching. The pixel error of the image stitching is no more than 1 pixel. The shutter speed of the global shutter can be adaptively adjusted within the range of 1 / 500 second to 1 / 2000 second.

3. The multimodal fusion-based visual defect detection system for automotive door panels according to claim 1, characterized in that, The adaptive multispectral light source unit is configured to perform the following working logic: acquire the spectral data of the door panel through the material identification sensor, match the data with the light source parameter templates pre-stored in the system for different door panel materials, and then automatically adjust the type, brightness and illumination angle of the light source; specifically, for black matte slush-molded door panels, adjust the brightness of the ring main light source to 500 lux and adjust the illumination angle of the strip auxiliary light source to 30 degrees; for chrome-plated metal door panels, turn off the direct light path of the ring LED main light source and the strip auxiliary light source, and enable the polarizer module, retaining only the diffuse light source for illumination.

4. The multimodal fusion-based visual defect detection system for automotive door panels according to claim 1, characterized in that, The multi-scale feature extraction network in the DoorNet model contains four convolutional pooling layers and two deconvolutional layers, used to extract image features at three scales: 16×16 pixels, 32×32 pixels, and 64×64 pixels, respectively. The training of the model adopts a strategy combining transfer learning and incremental training. First, it is pre-trained on the publicly available NEU-DET metal surface defect dataset, and then fine-tuned using 50,000 sample images of car door panels containing 12 types of defects and 8 materials.

5. The multimodal fusion-based visual defect detection system for automotive door panels according to claim 1, characterized in that, The fusion process of the multimodal data fusion unit includes: Step S1: Using hardware synchronization signal and software timestamp alignment technology, time synchronization is achieved among visual images, LiDAR point clouds and ultrasonic echo signals, with synchronization accuracy within milliseconds. Step S2: Extract 3D contour features from the lidar point cloud, extract internal structure features from the ultrasonic echo signal, and extract surface texture features from the visual image, and then stitch these features together to form a unified multi-dimensional feature vector. Step S3: The multi-dimensional feature vector is fused at the decision level using a confidence-based weighted voting method; when a detection instance shows that the visual modality determines that the surface is abnormal, the lidar modality detects a depth change of 0.3 mm, and the ultrasonic modality determines that the internal structure is normal, the unit outputs the defect conclusion "surface depression, depth 0.3 mm".

6. The multimodal fusion-based visual defect detection system for automotive door panels according to claim 1, characterized in that, The parameter self-optimization unit uses a reinforcement learning algorithm, with the system false detection rate of less than 1% and false negative rate of less than 0.5% as the optimization target. It automatically adjusts 12 system parameters, including camera exposure time, light source brightness, and model inference threshold. When the system detects that the image brightness fluctuation exceeds the baseline value ±20% or the image texture matching degree is less than 85%, the adaptive adjustment process is triggered. The response time from triggering to completing the parameter adjustment does not exceed 3 seconds.

7. The multimodal fusion-based visual defect detection system for automotive door panels according to claim 1, characterized in that, The incremental data pool in the feedback optimization module is configured to automatically store misjudged and missed samples that have been manually verified. When the number of samples accumulated in the incremental data pool reaches 500, the lightweight training unit is automatically triggered. This training unit only updates the fully connected layer parameters and attention module parameters of the DoorNet model, and the training time for a single session does not exceed 5 minutes.

8. A detection method for an automotive door panel visual defect detection system based on the multimodal fusion of any one of claims 1 to 7, characterized in that, Includes the following steps: step S10: Data acquisition preparation, the multi-camera array unit, lidar unit and ultrasonic sensing unit perform coordinated positioning, and the adaptive multispectral light source unit completes the initial configuration of light source parameters based on the data obtained by the material identification sensor; Step S20: Dynamic data acquisition. The multi-camera array unit synchronously acquires multi-view images of the door panel and synthesizes a full-area image of the door panel through an image stitching algorithm. At the same time, the lidar unit acquires the three-dimensional contour data of the door panel, the ultrasonic sensing unit acquires the internal structure information of the door panel, and the high-speed transmission unit transmits all the acquired data to the data processing module in real time. Step S30: Defect feature processing and analysis. The DoorNet model extracts defect features from the global image, and the multimodal data fusion unit fuses multi-source feature vectors from vision, lidar, and ultrasound to complete defect identification and classification. Step S40: Output and execution of detection results. The decision output module outputs the type, location and quantification parameters of the defect, and feeds the result back to the production line PLC system to trigger the interception mechanism for the defective door panel. Step S50: System self-optimization. The feedback optimization module updates the incremental data pool based on the manual review results and starts the lightweight training unit when the conditions are met to optimize the parameters of the DoorNet model.

9. The detection method of the multimodal fusion automotive door panel visual defect detection system according to claim 8, characterized in that, In step S30, the DoorNet model has a recognition rate of no less than 95% for fine scratches with a width of 0.05mm and a recognition rate of no less than 93% for bubble defects with the same color as the door panel body; the LiDAR unit has a measurement error of ±0.02mm for defect depth.

10. The detection method of the multimodal fusion automotive door panel visual defect detection system according to claim 8, characterized in that, In step S20, the delay in transmitting a single frame image captured by a single camera to the data processing module does not exceed 15ms, and the total cycle of the system performing a complete inspection task on a single car door panel does not exceed 20 seconds.

Citation Information

Patent Citations

  • Visual detection system for defects of vehicle door interior trim

    CN109187561A