Intelligent visual inspection system and method for automotive parts assembly line

By employing an intelligent visual inspection method that combines multimodal synchronous perception and dynamic compensation, the challenges of inspection in automotive parts assembly lines under unstable lighting and dynamic environments have been solved. This method enables high-precision and real-time inspection of parts status, thereby improving the stability and accuracy of the production line.

CN122109114APending Publication Date: 2026-05-29WUXI JIN CHENGLI PRECISION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUXI JIN CHENGLI PRECISION CO LTD
Filing Date
2025-12-31
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Visual inspection systems on automotive parts assembly lines struggle to achieve high precision and real-time inspection in environments with unstable lighting, reflection interference, and dynamic conditions. This results in image distortion and reduced recognition accuracy, impacting production efficiency and product quality.

Method used

An intelligent visual inspection method employing multimodal synchronous perception, dynamic compensation, and adaptive correction is adopted. Data is collected synchronously through a high-precision camera, laser sensor, and illumination sensor. Image preprocessing and motion compensation are performed by combining polarization filtering, multi-source synthesis, and inertial measurement unit. Deep learning detection algorithms are applied to extract component features and assess the risk of accuracy instability to generate correction instructions.

Benefits of technology

It achieves high-precision inspection under complex lighting and high-speed motion environments, significantly improving the visibility of surface details and imaging stability of parts, enhancing inspection accuracy and production line stability, and ensuring product quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122109114A_ABST
    Figure CN122109114A_ABST
Patent Text Reader

Abstract

The application discloses an intelligent visual detection system and method for automobile part assembly lines, and particularly relates to the technical field of intelligent visual detection, through the cooperative collection of high-precision cameras, lasers and light sensors, the consistency of multi-source data in time and space is maintained, through the light optimization processing of polarization filtering and multi-light source synthesis, the reflection spot and highlight interference are effectively inhibited, through the coupling compensation of inertial measurement units and conveyor belt motion data, the image space alignment and attitude consistency are maintained, the deep learning detection algorithm further realizes the high-precision identification of the space position, attitude and contour of the parts, and the defect type, size and space coordinate information with confidence are output in combination with the defect detection process, through the comprehensive evaluation mechanism of light change parameters, reflection interference degree and position deviation, the correction instruction is generated in real time and the visual system parameters are automatically adjusted, so that the self-correction of detection precision and long-term stable operation are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent vision inspection technology, and more specifically, to an intelligent vision inspection system and method for automotive parts assembly lines. Background Technology

[0002] In intelligent vision inspection systems for automotive parts assembly lines, changes in lighting conditions and surface reflection or gloss of parts are two closely related factors that directly and complexly impact the acquisition and processing of visual data. First, instability in lighting conditions causes variations in surface reflection, especially under strong light or overexposure, where reflected light spots and gloss significantly interfere with image quality. This uneven lighting not only blurs surface details but also creates highlight areas in the image. These highlight areas may obscure or distort the true shape of the parts, further affecting the accuracy of defect detection. Simultaneously, the surface materials of the parts (such as metals and glass, which are reflective) exacerbate this effect. The intensity and distribution of reflected light are closely related to changes in lighting. When the light source is unstable, bright spots and artifacts on reflective surfaces become more prominent, leading to unnecessary noise and artifacts in the image, thereby reducing the recognition accuracy of the vision system.

[0003] On the other hand, the high-speed movement and vibration of components in dynamic environments exacerbate the effects of lighting and surface reflection. During assembly, the rapid movement of robotic arms and conveyor belts can cause deviations in the position of components on the assembly line, leading to assembly errors. Even slight offsets in component position are not only a common problem in assembly but also increase the difficulty for vision sensors to track and locate components. As the complexity of dynamic environments increases, components may move at high speeds, making it difficult for vision systems to accurately capture their dynamic states in real time. Furthermore, the interaction between assembly errors and the dynamic environment can create a vicious cycle; errors and positional offsets can lead to blurred or distorted images, making originally clear features difficult to identify, thus affecting the detection of component morphology and defects.

[0004] The interaction of these factors poses significant challenges to the real-time performance and accuracy of visual inspection systems. Due to the interplay of lighting instability and surface reflection issues, the true shape and defect information of parts often cannot be accurately represented. Simultaneously, changes in part position under dynamic conditions further complicate image acquisition. These factors collectively lead to distortion and instability in visual data, significantly reducing the accuracy of recognition, localization, and defect detection. This, in turn, affects the overall efficiency of the production line and product quality, and may even result in undetected defects, causing unnecessary production stoppages or rework. Therefore, addressing the combined effects of lighting, reflection, assembly errors, and dynamic environments is crucial for improving the reliability of intelligent visual inspection methods in automotive parts assembly lines. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide an intelligent vision inspection system and method for automotive parts assembly lines to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: The intelligent vision inspection system for automotive parts assembly lines includes a data synchronization acquisition module, an image preprocessing and illumination optimization module, a motion compensation and viewpoint alignment module, a deep learning detection and feature extraction module, a defect detection and information output module, a precision instability risk assessment module, and a correction instruction execution and adjustment module. The data synchronization acquisition module is used to synchronously acquire timestamped images, positions, attitudes, and ambient lighting data of components through high-precision cameras, laser sensors, and light sensors, generating raw multimodal data streams; The image preprocessing and illumination optimization module is used to perform noise reduction, automatic exposure adjustment and gain control on the image data in the multimodal data stream; at the same time, it uses polarization filtering and multi-light source synthesis to suppress reflected light spots and highlight areas in the image, so as to obtain a preprocessed and illuminated image data stream. The motion compensation and viewpoint alignment module is used to perform motion compensation on the image data using the inertial measurement unit and the conveyor belt motion data, correct the positional deviation and distortion caused by the movement of the components, and perform viewpoint alignment through the multi-camera system to obtain image data after dynamic compensation and spatial alignment. The deep learning detection and feature extraction module is used to apply deep learning detection algorithms to compensated and aligned image data to extract the spatial position, pose and contour features of parts and generate parts recognition results with confidence. The defect detection and information output module is used to perform defect detection on the surface of the component based on the component identification results, identify cracks and scratches, calculate the specific location and size of the defects, and output defect information including defect type, size and relative spatial coordinates of the component. The accuracy instability risk assessment module is used to assess the accuracy instability risk of visual inspection based on the defect information, illumination change parameters, reflection interference level and component position deviation, and generate corresponding correction instructions. The calibration instruction execution and adjustment module is used to adjust the vision system and operating parameters on the assembly line according to the calibration instructions.

[0007] In a preferred embodiment, timestamped images, position, orientation, and ambient lighting data of components are simultaneously acquired using a high-precision camera, a laser sensor, and a light sensor to generate a raw multimodal data stream, as follows: A unified industrial time synchronization protocol is used to perform time synchronization calibration on high-precision cameras, laser sensors, and illumination sensors. The camera resolution, frame rate, and exposure time are set according to the assembly line cycle time and component type, and the scanning frequency and ranging range of the laser sensor are set. Simultaneously, multi-sensor spatial calibration is performed using known geometry to establish a unified extrinsic parameter matrix to achieve coordinate system unification. When a component enters the inspection area, a photoelectric switch signal from the conveyor belt triggers all sensors to synchronously acquire data. The camera captures two-dimensional image frames, the laser sensor outputs a depth point cloud matrix, and the illumination sensor records the current ambient illumination data. Within each sampling cycle, a unified format timestamp is appended to the data acquired by each sensor, and the inertial measurement unit records the component's attitude, velocity, and acceleration information at different times. Image frames, point cloud matrices, illumination data, and attitude parameters acquired at the same timestamp are time-matched and indexed. Based on multimodal fusion standards, these are uniformly packaged into a single frame of multimodal sampling dataset. The continuous time-series multimodal sampling dataset is output in timestamp order to form a structured raw multimodal data stream.

[0008] In a preferred embodiment, the image data in the multimodal data stream undergoes denoising, automatic exposure adjustment, and gain control; simultaneously, polarization filtering and multi-light source synthesis are used to suppress reflected light spots and highlight areas in the image, resulting in a pre-processed and illuminated image data stream, as detailed below: The original image frames and their ambient lighting data at the corresponding timestamps are extracted from the multimodal data stream, and spatial domain noise reduction is performed on the original images using adaptive median filtering. Based on the real-time illuminance value fed back by the light sensor and the image histogram distribution, the average brightness deviation and saturation pixel ratio are calculated. The PID closed-loop control algorithm is used to dynamically adjust the camera exposure time and analog gain, and output a sequence of image frames with balanced brightness. An electrically adjustable polarizer is applied in the imaging optical path. The image brightness response under different polarization directions is scanned by rotating the angle, and the polarization contrast between the reflection component and the diffuse reflection component is calculated. The image frame with the optimal polarization direction is selected or the dereflection image is reconstructed based on the polarization decomposition model. By alternately lighting up controllable light source arrays distributed at different angles and acquiring multiple frames of images, a fused image is synthesized through brightness difference and weighted fusion. The multi-frame images, after denoising, exposure control, polarization suppression, and light source synthesis, are reordered according to their timestamps and matched with the pose parameters and illumination data in the original multimodal data stream to obtain a preprocessed and illuminated image data stream.

[0009] In a preferred embodiment, motion compensation is performed on the image data using an inertial measurement unit and conveyor belt motion data to correct positional deviations and distortions caused by component movement. Furthermore, a multi-camera system is used for viewpoint alignment to obtain image data that has undergone dynamic compensation and spatial alignment, as detailed below: Extract inertial measurement unit data corresponding to each frame of the multimodal data stream, including linear acceleration. and angular velocity Simultaneously, it acquires motion data from the conveyor belt encoder; Based on the inertial measurement unit and conveyor belt motion data, the time interval of each component in each frame is calculated. Spatial displacement within and attitude changes : ,in For the velocity vector of the component, The linear acceleration vector obtained by the inertial measurement unit. The displacement increment obtained by the conveyor belt encoder. Let the unit vector be the direction of the conveyor belt's movement. It is the antisymmetric matrix of the angular velocity vector; Motion compensation for image data: Projecting the component points in each frame of the image. Compensation is performed by mapping from the camera pixel coordinate system to the world coordinate system: ,in For the camera intrinsic parameter matrix, These are the pixel coordinates after motion compensation; Multi-camera viewpoint alignment: based on camera extrinsic matrix Projecting the images from each camera onto a unified reference coordinate system yields a spatially aligned image: ,in Let h be the intrinsic parameter matrix of the h-th camera. These are the extrinsic rotation matrix and translation vector of the h-th camera, respectively. The three-dimensional coordinates of the parts in the world coordinate system. These are the aligned pixel coordinates; The image frames aligned with the viewpoints of each camera are integrated with the corresponding timestamp data and motion compensation parameters to obtain image data after dynamic compensation and spatial alignment.

[0010] In a preferred embodiment, a deep learning detection algorithm is applied to the compensated and aligned image data to extract the spatial position, pose, and contour features of the components, generating a component recognition result with confidence, as follows: The image data stream, after dynamic compensation and spatial alignment, is used as input to the deep learning algorithm. The input image is then normalized to the required fixed size for the model. ,in These are the original image pixel values. The mean of the image pixel values. The standard deviation of the image pixel values. These are the normalized image pixel values; The deep learning detection algorithm's model architecture employs a pre-trained convolutional neural network. Standardized image data is input, and the network performs forward propagation. The convolutional neural network extracts deep features from the image through multiple convolutional, pooling, and fully connected layers. The bounding box positions of the components are obtained using the regression output of the detection network. ; The orientation angles of the components are calculated by using the regression layer of the detection network. ; Image processing techniques are used to extract the contour features of components from images; The extracted contour features are input into a pre-trained classification network for classification to determine the type of parts, and the Softmax function is used to output the probability of each category. By combining the bounding box from object detection, pose estimation, and classification probability, the final component recognition result is generated, including the following: The bounding box position of the identified components ; The attitude angle of the component ; Confidence score.

[0011] In a preferred embodiment, based on the defect information, illumination variation parameters, degree of reflection interference, and component position deviation, the risk of accuracy instability in visual inspection is assessed, and corresponding correction instructions are generated, as follows: The defect impact factor is obtained based on the defect information, using the following formula: ,in As a defect impact factor, Let be the area of ​​the defect. The total area of ​​the image. These are the length and width of the defect, respectively; The illumination variation parameters include the illumination variation factor, and the specific formula is as follows: ,in The light variation factor, This represents the change in ambient light intensity. This represents the current light intensity. The degree of reflection interference includes the reflection interference factor, the specific formula of which is as follows: ,in As a reflection interference factor, The area of ​​the reflected light spot. The total area of ​​the image; The positional deviation of components includes a positional deviation factor, the specific formula of which is as follows: ,in For position deviation factor, This represents the maximum deviation of the component. This represents the maximum acceptable deviation for the assembly line. The weighted summation of the defect impact factor, illumination change factor, reflection interference factor, and position deviation factor yields the following formula for calculating the comprehensive instability risk index: ,in To form a comprehensive instability risk index, As a defect impact factor, The light variation factor, As a reflection interference factor, For position deviation factor, These are the preset weighting coefficients for the defect impact factor, illumination change factor, reflection interference factor, and position deviation factor, respectively. All are greater than 0; The comprehensive instability risk index is compared with the preset comprehensive instability risk threshold. If the comprehensive instability risk index is greater than the comprehensive instability risk threshold, a correction needs to be performed; otherwise, the current setting is maintained. The correction commands include illumination adjustment commands, position adjustment commands, and reflection interference correction commands.

[0012] In a preferred embodiment, the intelligent vision inspection method for automotive parts assembly lines includes the following steps: The system synchronously acquires timestamped images, positions, orientations, and ambient lighting data of components using high-precision cameras, laser sensors, and light sensors, generating a raw multimodal data stream. The image data in the multimodal data stream is subjected to denoising, automatic exposure adjustment, and gain control; at the same time, polarization filtering and multi-light source synthesis are used to suppress reflected light spots and highlight areas in the image, resulting in a pre-processed and optimized image data stream. The image data is motion-compensated using an inertial measurement unit and conveyor belt motion data to correct positional deviations and distortions caused by component movement. The viewpoint is aligned using a multi-camera system to obtain image data after dynamic compensation and spatial alignment. Deep learning detection algorithms are applied to compensated and aligned image data to extract the spatial position, pose and contour features of parts, and generate parts recognition results with confidence. Based on the component identification results, defect detection is performed on the component surface to identify cracks and scratches, calculate the specific location and size of the defects, and output defect information including defect type, size and relative spatial coordinates of the component. Based on the defect information, illumination variation parameters, reflection interference level, and component position deviation, assess the risk of accuracy instability in visual inspection and generate corresponding correction instructions; Adjust the vision system and operating parameters on the assembly line according to the calibration instructions.

[0013] The technical effects and advantages of this invention are as follows: 1. This invention introduces a full-process visual inspection mechanism in automotive parts assembly lines, incorporating multimodal synchronous perception, dynamic compensation, intelligent recognition, and adaptive correction. This mechanism enables high-precision detection and real-time optimization of parts' states under complex lighting environments, reflection interference, and high-speed motion scenarios. The method achieves temporal synchronization of images, positions, postures, and lighting through the collaborative acquisition of high-precision cameras, lasers, and lighting sensors. This ensures consistency of multi-source data in both time and space, reducing information misalignment caused by acquisition delays. Lighting optimization processing, including polarization filtering and multi-source synthesis, effectively suppresses reflected light spots and high-light interference, significantly improving the visibility of surface details and imaging stability of parts. Through coupling compensation between the inertial measurement unit and conveyor belt motion data, the system maintains image spatial alignment and posture consistency in high-speed movement and vibration environments, solving the problems of image blurring and deformation under dynamic conditions in traditional visual inspection. Deep learning detection algorithms further achieve high-precision recognition of the spatial position, posture, and contour of parts. Combined with the defect detection process, the system outputs defect type, size, and spatial coordinate information with confidence, enabling not only defect identification but also precise quantification of their affected area. Finally, through a comprehensive evaluation mechanism considering lighting variation parameters, reflection interference levels, and positional deviations, the system can generate correction commands in real time and automatically adjust vision system parameters, achieving self-correction of detection accuracy and long-term stable operation. This effectively overcomes the shortcomings of traditional assembly line visual inspection, which is greatly affected by lighting fluctuations, reflection interference, and motion distortion. It significantly improves detection accuracy, environmental adaptability, and production line stability, thereby ensuring consistency between product quality and assembly efficiency. Attached Figure Description

[0014] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings; Figure 1 This is a flowchart of the system in Embodiment 1 of the present invention; Figure 2 This is a flowchart of the method in Embodiment 2 of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] Example 1: Figure 1The present invention provides an intelligent vision inspection system for automotive parts assembly lines, comprising a data synchronization acquisition module, an image preprocessing and illumination optimization module, a motion compensation and viewpoint alignment module, a deep learning detection and feature extraction module, a defect detection and information output module, a precision instability risk assessment module, and a correction instruction execution and adjustment module. The data synchronization acquisition module is used to synchronously acquire timestamped images, positions, attitudes, and ambient lighting data of components through high-precision cameras, laser sensors, and light sensors, generating raw multimodal data streams; The image preprocessing and illumination optimization module is used to perform noise reduction, automatic exposure adjustment and gain control on the image data in the multimodal data stream; at the same time, it uses polarization filtering and multi-light source synthesis to suppress reflected light spots and highlight areas in the image, so as to obtain a preprocessed and illuminated image data stream. The motion compensation and viewpoint alignment module is used to perform motion compensation on the image data using the inertial measurement unit and the conveyor belt motion data, correct the positional deviation and distortion caused by the movement of the components, and perform viewpoint alignment through the multi-camera system to obtain image data after dynamic compensation and spatial alignment. The deep learning detection and feature extraction module is used to apply deep learning detection algorithms to compensated and aligned image data to extract the spatial position, pose and contour features of parts and generate parts recognition results with confidence. The defect detection and information output module is used to perform defect detection on the surface of the component based on the component identification results, identify cracks and scratches, calculate the specific location and size of the defects, and output defect information including defect type, size and relative spatial coordinates of the component. The accuracy instability risk assessment module is used to assess the accuracy instability risk of visual inspection based on the defect information, illumination change parameters, reflection interference level and component position deviation, and generate corresponding correction instructions. The calibration instruction execution and adjustment module is used to adjust the vision system and operating parameters on the assembly line according to the calibration instructions.

[0017] The system synchronously acquires timestamped images, positions, orientations, and ambient lighting data of components using high-precision cameras, laser sensors, and light sensors, generating a raw multimodal data stream. In this embodiment of the invention, a unified industrial time synchronization protocol (such as IEEE 1588 PTP or hardware trigger signals) is used to perform time synchronization calibration on high-precision cameras, laser sensors, and illumination sensors to ensure that each sensor triggers sampling under the same time reference. The resolution, frame rate, and exposure time of the camera are set according to the assembly line cycle and component type, and the scanning frequency and ranging range of the laser sensor are set. Simultaneously, multi-sensor spatial calibration is performed using known geometry to establish a unified extrinsic parameter matrix to achieve coordinate system unification. When a component enters the inspection area, a photoelectric switch signal from the conveyor belt triggers all sensors to synchronously acquire data. The camera captures two-dimensional image frames, the laser sensor outputs a depth point cloud matrix, and the illumination sensor records the current ambient illumination data. Within each sampling cycle, a unified format is appended to the data acquired by each sensor. The system uses timestamps to record the attitude, velocity, and acceleration information of components at different times via inertial measurement units, achieving spatiotemporal binding of visual and motion data. Image frames, point cloud matrices, illumination data, and attitude parameters acquired at the same timestamp are time-matched and indexed. Based on multimodal fusion standards, these are uniformly encapsulated into a single frame of multimodal sampling dataset, ensuring a one-to-one correspondence between image information, spatial geometry information, and illumination features. The continuous time-series multimodal sampling datasets are output in timestamp order, forming a structured raw multimodal data stream. The data format includes image frame data, depth point cloud data, illumination data, attitude parameters, and synchronization identifier information, used for subsequent image preprocessing and dynamic compensation analysis.

[0018] The image data in the multimodal data stream is subjected to denoising, automatic exposure adjustment, and gain control; at the same time, polarization filtering and multi-light source synthesis are used to suppress reflected light spots and highlight areas in the image, resulting in a pre-processed and optimized image data stream. In this embodiment of the invention, the original image frame and its ambient lighting data (including illuminance value, color temperature distribution and spectral reflectance characteristics) under the corresponding timestamp are extracted from the multimodal data stream to provide a lighting feature benchmark for subsequent image preprocessing; the original image is spatially denoised by adaptive median filtering to remove high-frequency noise caused by sensor thermal noise, random light flicker and motion blur, while maintaining the edge structure without distortion.

[0019] Based on the real-time illuminance value fed back by the light sensor and the image histogram distribution, the average brightness deviation and saturation pixel ratio are calculated. The PID closed-loop control algorithm is used to dynamically adjust the camera exposure time and analog gain so that the pixel grayscale value is maintained within the preset dynamic range and the image frame sequence with balanced brightness is output. The average brightness deviation and saturation pixel ratio are calculated based on the real-time illuminance value fed back by the light sensor and the image histogram distribution. A PID closed-loop control algorithm is then used to dynamically adjust the camera exposure time and analog gain, as detailed below: Based on the real-time illuminance values ​​collected by the light sensor Calculate the average brightness deviation using the grayscale histogram of the current frame image. : ,in The target brightness reference value, Total number of pixels Let i be the gray value corresponding to the i-th gray level in the image. is the normalized frequency of the gray-level histogram at the i-th gray level; The statistical grayscale value exceeds the saturation threshold. The pixel ratio is used as the saturation pixel ratio. ; Average brightness deviation Compared to saturated pixels Exposure error signal synthesized by weight : ,in These are the average brightness deviations. Compared to saturated pixels The preset proportional coefficient, and All are greater than 0; It should be noted that, The settings should be tailored to the specific circumstances. For example, an expert-empowered approach could be adopted, where experts in relevant fields are invited to determine the pre-defined proportions for each indicator through professional opinion surveys and comprehensive evaluations. The initial value can be 0.5, 0.5; Calculate the exposure time adjustment amount within each control cycle. : ,in These are the proportional, integral, and derivative control coefficients, respectively, and then the exposure time ( Updated to And adjust the analog gain simultaneously. To keep the image brightness within the preset dynamic range Finally, the adjusted exposure time and analog gain are fed back to the camera driver layer through the camera system control unit.

[0020] An electrically adjustable polarizer is applied to the imaging optical path. The image brightness response under different polarization directions is scanned by rotating the angle, and the polarization contrast between the reflection and diffuse reflection components is calculated. The image frame with the optimal polarization direction is selected, or the de-reflection image is reconstructed based on the polarization decomposition model to eliminate surface specular reflection interference, as detailed below: An electrically adjustable polarizer is integrated into the imaging optical path and synchronized with the camera trigger signal. Let the polarization angle be... The polarizer is controlled to rotate at equal intervals (e.g., every 15°) to acquire multiple frames of image sequences with different polarization directions. Calculate the brightness variation curve of each pixel at different polarization angles. A model of reflected light is established based on Malus's law: ,in For diffuse reflection component, For the specular reflection component, The main polarization direction; For each pixel at different polarization angles Brightness value Extract maximum brightness With minimum brightness Calculate polarization contrast : This value reflects the strength of the specular reflection component of a pixel; a high polarization contrast indicates a more pronounced specular reflection. According to polarization contrast With preset reflection threshold Perform binarization to generate a mask for the specular reflection area. : ,in Indicates the area of ​​specular reflection. Indicates the diffuse reflection area; Using a mask Specular reflection component From the original image Separate the image from the middle to obtain the diffuse dominant image. : ; If the reflection area in the detection scene is limited, the frame with the smallest brightness gradient corresponding to the polarization direction is selected as the optimal polarization image; if reflection interference is significant, a polarization decomposition model is used to perform least-squares reconstruction on multiple frames. ; Obtain a high-fidelity image with de-reflection .

[0021] By alternately lighting up controllable light source arrays distributed at different angles and acquiring multiple frames of images, a fused image without saturated highlights is synthesized through brightness difference and weighted fusion. The multi-frame images, after denoising, exposure control, polarization suppression, and light source synthesis, are reordered according to their timestamps and matched with the pose parameters and illumination data in the original multimodal data stream to obtain a preprocessed and illuminated image data stream.

[0022] It should be noted that the above formulas are all dimensionless calculations. Commonly used methods for removing dimensions include Min-Max normalization and Z-Score standardization, which will not be elaborated here. The image data is motion-compensated using an inertial measurement unit and conveyor belt motion data to correct positional deviations and distortions caused by component movement. The viewpoint is aligned using a multi-camera system to obtain image data after dynamic compensation and spatial alignment. In this embodiment of the invention, inertial measurement unit (IMU) data corresponding to each frame of image is extracted from the multimodal data stream, including linear acceleration. and angular velocity Simultaneously, it acquires motion data from the conveyor belt encoder to calculate the displacement of components within the acquisition time interval. Based on the inertial measurement unit and conveyor belt motion data, the time interval of each component in each frame is calculated. Spatial displacement within and attitude changes : ,in For the velocity vector of the component, The linear acceleration vector obtained by the inertial measurement unit. The displacement increment obtained by the conveyor belt encoder. Let the unit vector be the direction of the conveyor belt's movement. It is the antisymmetric matrix of the angular velocity vector; Motion compensation for image data: Projecting the component points in each frame of the image. Compensation is performed by mapping from the camera pixel coordinate system to the world coordinate system: ,in For the camera intrinsic parameter matrix, These are the pixel coordinates after motion compensation; Multi-camera viewpoint alignment: based on camera extrinsic matrix Projecting the images from each camera onto a unified reference coordinate system yields a spatially aligned image: ,in Let h be the intrinsic parameter matrix of the h-th camera. These are the extrinsic rotation matrix and translation vector of the h-th camera, respectively. The three-dimensional coordinates of the parts in the world coordinate system. These are the aligned pixel coordinates; The image frames aligned with the viewpoints of each camera are integrated with the corresponding timestamp data and motion compensation parameters to obtain image data after dynamic compensation and spatial alignment.

[0023] It should be noted that the above formulas are all dimensionless calculations. Commonly used methods for removing dimensions include Min-Max normalization and Z-Score standardization, which will not be elaborated here. Deep learning detection algorithms are applied to compensated and aligned image data to extract the spatial position, pose and contour features of parts, and generate parts recognition results with confidence. In this embodiment of the invention, the image data stream after dynamic compensation and spatial alignment is used as the input to the deep learning algorithm. The input image is standardized in size, converted into a fixed size required by the model (such as 224×224 or 512×512), and then normalized. ,in These are the original image pixel values. The mean of the image pixel values. The standard deviation of the image pixel values. These are the normalized image pixel values; The deep learning detection algorithm's model architecture employs pre-trained convolutional neural networks (CNNs), such as ResNet, YOLO, and Faster R-CNN, for component detection. Standardized image data is input, and the network performs forward propagation. The CNN extracts deep-level features from the image through multiple layers of convolution, pooling, and fully connected layers. During this process, the network learns features such as the component's shape, texture, and location.

[0024] The bounding box positions of the components are obtained using the regression output of the detection network. This refers to the coordinate range of the component in the image.

[0025] The orientation angles of the components are calculated by using the regression layer of the detection network. This refers to the rotation angle relative to the camera, usually represented using Euler angles or quaternions: ,in The coordinates of the feature points extracted from the image of the component; Use image processing techniques (such as Canny edge detection or contour extraction algorithms) to extract the contour features of components from the image; The extracted contour features are input into a pre-trained classification network for classification to determine the type of parts, and the Softmax function is used to output the probability of each category. By combining the bounding box from object detection, pose estimation, and classification probability, the final component recognition result is generated, including the following: The bounding box position of the identified components ; The attitude angle of the component ; Confidence score (classification probability).

[0026] It should be noted that the above formulas are all dimensionless calculations. Commonly used methods for removing dimensions include Min-Max normalization and Z-Score standardization, which will not be elaborated here. Based on the component identification results, defect detection is performed on the component surface to identify cracks and scratches, calculate the specific location and size of the defects, and output defect information including defect type, size and relative spatial coordinates of the component. In this embodiment of the invention, to improve the detectability of defects, the identified component images are enhanced. Common methods include contrast enhancement and edge enhancement. Edge features of the image are enhanced using the Laplacian or Sobel operators to highlight minor defects such as cracks or scratches. An adaptive thresholding algorithm (such as the Otsu method) is used to convert the image into a binary image, where cracks and scratches appear as significant black-and-white differences. Connected component analysis algorithms (such as labeling or morphological operations) are used to extract crack or scratch regions from the image. Morphological operations (erosion, dilation, opening, closing) are used to remove noise and extract valid defect regions. Regions that do not meet the standards (e.g., noise or incomplete defects) are filtered out based on the area or shape characteristics of the defects (such as aspect ratio, roundness, etc.). Further feature extraction is performed on the extracted defect regions to classify different types of defects. Commonly used features include shape features (aspect ratio, roundness, etc.) and texture features (gray-level co-occurrence matrix, etc.). The extracted features are trained using machine learning algorithms (such as support vector machines, random forests, deep learning classifiers, etc.) to classify different types of defects (such as cracks, scratches, etc.). For each defect region, calculate its centroid coordinates. Location of the defect: ,in Defect area The number of pixels within the cell; The output of each defect is formatted and includes its type, dimensions (length, width, area), location coordinates, and spatial coordinates relative to the component. ,in Defect type (such as cracks, scratches, etc.) These are the length and width of the defect, respectively. Let be the area of ​​the defect. These are the coordinates of the defect's location. The coordinates of the defect relative to the three-dimensional space of the component.

[0027] Based on the defect information, illumination variation parameters, reflection interference level, and component position deviation, assess the risk of accuracy instability in visual inspection and generate corresponding correction instructions; In this embodiment of the invention, the defect impact factor is obtained based on defect information, and the specific formula is as follows: ,in As a defect impact factor, Let be the area of ​​the defect. The total area of ​​the image. These are the length and width of the defect, respectively; The impact of defects on test results is assessed based on their location and size. Larger defects have a greater impact.

[0028] The illumination variation parameters include the illumination variation factor, and the specific formula is as follows: ,in The light variation factor, This represents the change in ambient light intensity. This represents the current light intensity. The impact of illumination changes on detection accuracy is assessed based on the magnitude of the illumination variation.

[0029] The degree of reflection interference includes the reflection interference factor, the specific formula of which is as follows: ,in As a reflection interference factor, The area of ​​the reflected light spot. The total area of ​​the image; The degree of reflection interference is assessed by the brightness or area of ​​the reflected light spot in the image; the larger the reflected light spot, the stronger the interference.

[0030] The positional deviation of components includes a positional deviation factor, the specific formula of which is as follows: ,in For position deviation factor, This represents the maximum deviation of the component (which can be obtained through positional deviation after motion compensation). The maximum acceptable deviation for the assembly line (such as the assembly line tolerance standard). The comprehensive instability risk index is obtained by weighting and summing the defect impact factor, illumination change factor, reflection interference factor, and position deviation factor. The weights of the comprehensive instability risk index can be set according to the actual application scenario and the importance of each factor (e.g., if illumination change has a greater impact on accuracy, it can be assigned a higher weight). The formula for calculating the comprehensive instability risk index is as follows: ,in To form a comprehensive instability risk index, As a defect impact factor, The light variation factor, As a reflection interference factor, For position deviation factor, These are the preset weighting coefficients for the defect impact factor, illumination change factor, reflection interference factor, and position deviation factor, respectively. All are greater than 0; It should be noted that the above formulas are all dimensionless calculations. Commonly used methods for removing dimensions include Min-Max normalization and Z-Score standardization, which will not be elaborated here. The settings should be tailored to the specific circumstances. For example, an expert weighting method can be used, which involves inviting experts in relevant fields to determine the pre-defined weighting coefficients for each indicator through professional opinion surveys and comprehensive evaluations. The initial value can be 0.25, 0.25, 0.25, 0.25.

[0031] The comprehensive instability risk index is compared with the preset comprehensive instability risk threshold. If the comprehensive instability risk index is greater than the comprehensive instability risk threshold, a correction needs to be performed; otherwise, the current setting is maintained. The correction commands include illumination adjustment commands, used to adjust the camera's exposure time or gain control to stabilize illumination conditions; position adjustment commands, used to adjust the assembly position of components or the movement speed of the conveyor belt to reduce positional deviations; and reflection interference correction commands, used to adjust the polarization filter or light source angle to reduce interference from reflected light spots.

[0032] Adjust the vision system and operating parameters on the assembly line according to the calibration instructions.

[0033] In this embodiment of the invention, after receiving a calibration command, the content of the calibration command will be classified according to various operating parameters in order to perform the corresponding adjustment operation; Lighting adjustment includes adjusting exposure time and adjusting gain control; The adjustment of exposure time refers to adjusting the camera's exposure time according to the light adjustment command. If the light change factor is large, the exposure time is increased to adapt to changes in ambient light. The adjustment gain control means that if the illumination change factor is large, the vision system adjusts the image gain according to the correction command to ensure the image quality under changing illumination conditions. Position adjustment includes position deviation correction; The position deviation correction is achieved by the system adjusting the assembly position of the parts according to the position adjustment command. If the position deviation factor is large, the control system reduces the error by correcting the position or speed of the conveyor belt or the robotic arm. Correction of reflection interference includes adjusting the angle of the polarizing filter and adjusting the angle of the light source; The adjustment of the polarization filter angle is based on the reflection interference correction command, which adjusts the angle of the polarization filter installed in front of the camera. The polarization filter can effectively reduce reflected light spots and highlight areas in the image, and improve the image quality of visual inspection. The adjustment of the light source angle is to adjust the position or angle of the light source if the reflection interference factor is high, so as to avoid direct or reflected light affecting the image acquisition of the vision system.

[0034] This invention introduces a full-process visual inspection mechanism in automotive parts assembly lines, incorporating multimodal synchronous perception, dynamic compensation, intelligent recognition, and adaptive correction. This mechanism enables high-precision detection and real-time optimization of parts' states under complex lighting environments, reflection interference, and high-speed motion scenarios. The method achieves temporal synchronization of images, positions, postures, and lighting through the collaborative acquisition of high-precision cameras, lasers, and lighting sensors. This ensures consistency of multi-source data in both time and space, reducing information misalignment caused by acquisition delays. Lighting optimization processing, including polarization filtering and multi-source synthesis, effectively suppresses reflected light spots and high-light interference, significantly improving the visibility of surface details and imaging stability of parts. Through coupling compensation between the inertial measurement unit and conveyor belt motion data, the system maintains image spatial alignment and posture consistency in high-speed movement and vibration environments, solving the problems of image blurring and deformation in traditional visual inspection under dynamic conditions. Deep learning detection algorithms further achieve high-precision recognition of the spatial position, posture, and contour of parts. Combined with the defect detection process, the system outputs defect type, size, and spatial coordinate information with confidence, enabling not only defect identification but also precise quantification of their affected area. Finally, through a comprehensive evaluation mechanism considering lighting variation parameters, reflection interference levels, and positional deviations, the system can generate correction commands in real time and automatically adjust vision system parameters, achieving self-correction of detection accuracy and long-term stable operation. This effectively overcomes the shortcomings of traditional assembly line visual inspection, which is greatly affected by lighting fluctuations, reflection interference, and motion distortion. It significantly improves detection accuracy, environmental adaptability, and production line stability, thereby ensuring consistency between product quality and assembly efficiency.

[0035] Example 2: This example introduces an intelligent vision inspection method for automotive parts assembly lines, such as... Figure 2 As shown, it includes the following steps: The system synchronously acquires timestamped images, positions, orientations, and ambient lighting data of components using high-precision cameras, laser sensors, and light sensors, generating a raw multimodal data stream. The image data in the multimodal data stream is subjected to denoising, automatic exposure adjustment, and gain control; at the same time, polarization filtering and multi-light source synthesis are used to suppress reflected light spots and highlight areas in the image, resulting in a pre-processed and optimized image data stream. The image data is motion-compensated using an inertial measurement unit and conveyor belt motion data to correct positional deviations and distortions caused by component movement. The viewpoint is aligned using a multi-camera system to obtain image data after dynamic compensation and spatial alignment. Deep learning detection algorithms are applied to compensated and aligned image data to extract the spatial position, pose and contour features of parts, and generate parts recognition results with confidence. Based on the component identification results, defect detection is performed on the component surface to identify cracks and scratches, calculate the specific location and size of the defects, and output defect information including defect type, size and relative spatial coordinates of the component. Based on the defect information, illumination variation parameters, reflection interference level, and component position deviation, assess the risk of accuracy instability in visual inspection and generate corresponding correction instructions; Adjust the vision system and operating parameters on the assembly line according to the calibration instructions.

[0036] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0037] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0038] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0039] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system and method described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0040] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways.

[0041] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An intelligent vision inspection system for automotive parts assembly lines, characterized in that: It includes a data synchronization acquisition module, an image preprocessing and illumination optimization module, a motion compensation and viewpoint alignment module, a deep learning detection and feature extraction module, a defect detection and information output module, a precision instability risk assessment module, and a correction instruction execution and adjustment module; The data synchronization acquisition module is used to synchronously acquire timestamped images, positions, attitudes, and ambient lighting data of components through high-precision cameras, laser sensors, and light sensors, generating raw multimodal data streams; The image preprocessing and illumination optimization module is used to perform noise reduction, automatic exposure adjustment and gain control on the image data in the multimodal data stream; at the same time, it uses polarization filtering and multi-light source synthesis to suppress reflected light spots and highlight areas in the image, so as to obtain a preprocessed and illuminated image data stream. The motion compensation and viewpoint alignment module is used to perform motion compensation on the image data using the inertial measurement unit and the conveyor belt motion data, correct the positional deviation and distortion caused by the movement of the components, and perform viewpoint alignment through the multi-camera system to obtain image data after dynamic compensation and spatial alignment. The deep learning detection and feature extraction module is used to apply deep learning detection algorithms to compensated and aligned image data to extract the spatial position, pose and contour features of parts and generate parts recognition results with confidence. The defect detection and information output module is used to perform defect detection on the surface of the component based on the component identification results, identify cracks and scratches, calculate the specific location and size of the defects, and output defect information including defect type, size and relative spatial coordinates of the component. The accuracy instability risk assessment module is used to assess the accuracy instability risk of visual inspection based on the defect information, illumination change parameters, reflection interference level and component position deviation, and generate corresponding correction instructions. The calibration instruction execution and adjustment module is used to adjust the vision system and operating parameters on the assembly line according to the calibration instructions.

2. The intelligent vision inspection system for automotive parts assembly lines according to claim 1, characterized in that: High-precision cameras, laser sensors, and illumination sensors are used to simultaneously acquire timestamped images, positions, orientations, and ambient lighting data of components, generating raw multimodal data streams, as follows: A unified industrial time synchronization protocol is used to perform time synchronization calibration on high-precision cameras, laser sensors, and illumination sensors. The camera resolution, frame rate, and exposure time are set according to the assembly line cycle time and component type, and the scanning frequency and ranging range of the laser sensor are set. Simultaneously, multi-sensor spatial calibration is performed using known geometry to establish a unified extrinsic parameter matrix to achieve coordinate system unity. When a component enters the inspection area, a photoelectric switch signal from the conveyor belt triggers all sensors to synchronously acquire data. The camera captures two-dimensional image frames, the laser sensor outputs a depth point cloud matrix, and the illumination sensor records the current ambient illumination data. Within each sampling cycle, a unified format timestamp is appended to the data acquired by each sensor, and the attitude, velocity, and acceleration information of the component at different times is recorded by an inertial measurement unit. Image frames, point cloud matrices, lighting data, and pose parameters acquired at the same timestamp will be matched and indexed together by time. Based on the multimodal fusion standard, it is uniformly packaged into a single frame of multimodal sampling dataset; the continuous time series multimodal sampling dataset is output in timestamp order to form a structured original multimodal data stream.

3. The intelligent vision inspection system for automotive parts assembly lines according to claim 2, characterized in that: The image data in the multimodal data stream is subjected to denoising, automatic exposure adjustment, and gain control; simultaneously, polarization filtering and multi-light source synthesis are used to suppress reflected light spots and highlight areas in the image, resulting in a pre-processed and optimized image data stream, as detailed below: The original image frames and their ambient lighting data at the corresponding timestamps are extracted from the multimodal data stream, and spatial domain noise reduction is performed on the original images using adaptive median filtering. Based on the real-time illuminance value fed back by the light sensor and the image histogram distribution, the average brightness deviation and saturation pixel ratio are calculated. The camera exposure time and analog gain are dynamically adjusted using a PID closed-loop control algorithm to output a sequence of image frames with balanced brightness. An electrically adjustable polarizer is applied in the imaging optical path. The image brightness response under different polarization directions is scanned by rotating the angle, and the polarization contrast between the reflection component and the diffuse reflection component is calculated. The image frame with the optimal polarization direction is selected or the dereflection image is reconstructed based on the polarization decomposition model. By alternately lighting up controllable light source arrays distributed at different angles and acquiring multiple frames of images, a fused image is synthesized through brightness difference and weighted fusion. The multi-frame images, after denoising, exposure control, polarization suppression, and light source synthesis, are reordered according to timestamps and matched with the attitude parameters and illumination data in the original multimodal data stream to obtain a preprocessed and illuminated image data stream.

4. The intelligent vision inspection system for automotive parts assembly lines according to claim 3, characterized in that: The image data is motion-compensated using an inertial measurement unit and conveyor belt motion data to correct positional deviations and distortions caused by component movement. A multi-camera system is then used for viewpoint alignment to obtain dynamically compensated and spatially aligned image data, as detailed below: Extract inertial measurement unit data corresponding to each frame of the multimodal data stream, including linear acceleration. and angular velocity Simultaneously, it acquires motion data from the conveyor belt encoder; Based on the inertial measurement unit and conveyor belt motion data, the time interval of each component in each frame is calculated. Spatial displacement within and attitude changes : ,in For the velocity vector of the component, The linear acceleration vector obtained by the inertial measurement unit. The displacement increment obtained by the conveyor belt encoder. Let the unit vector be the direction of the conveyor belt's movement. It is the antisymmetric matrix of the angular velocity vector; Motion compensation for image data: Projecting the component points in each frame of the image. Compensation is performed by mapping from the camera pixel coordinate system to the world coordinate system: ,in For the camera intrinsic parameter matrix, These are the pixel coordinates after motion compensation; Multi-camera viewpoint alignment: based on camera extrinsic matrix Projecting the images from each camera onto a unified reference coordinate system yields a spatially aligned image: ,in Let h be the intrinsic parameter matrix of the h-th camera. These are the extrinsic rotation matrix and translation vector of the h-th camera, respectively. The three-dimensional coordinates of the components in the world coordinate system. These are the aligned pixel coordinates; The image frames aligned with the viewpoints of each camera are integrated with the corresponding timestamp data and motion compensation parameters to obtain image data after dynamic compensation and spatial alignment.

5. The intelligent vision inspection system for automotive parts assembly lines according to claim 4, characterized in that: Deep learning detection algorithms are applied to the compensated and aligned image data to extract the spatial position, pose, and contour features of the components, generating confident component recognition results, as follows: The image data stream, after dynamic compensation and spatial alignment, is used as input to the deep learning algorithm. The input image is then normalized to the required fixed size for the model. ,in These are the original image pixel values. The mean of the image pixel values. The standard deviation of the image pixel values. These are the normalized image pixel values; The deep learning detection algorithm's model architecture employs a pre-trained convolutional neural network. Standardized image data is input, and the network performs forward propagation. The convolutional neural network extracts deep features from the image through multiple convolutional, pooling, and fully connected layers. The bounding box positions of the components are obtained using the regression output of the detection network. ; The orientation angles of the components are calculated by using the regression layer of the detection network. ; Image processing techniques are used to extract the contour features of components from images; The extracted contour features are input into a pre-trained classification network for classification to determine the type of parts, and the Softmax function is used to output the probability of each category. By combining the bounding box from object detection, pose estimation, and classification probability, the final component recognition result is generated, including the following: The bounding box position of the identified components ; The attitude angle of the component ; Confidence score.

6. The intelligent vision inspection system for automotive parts assembly lines according to claim 1, characterized in that: Based on the defect information, illumination variation parameters, reflection interference level, and component position deviation, the risk of accuracy instability in visual inspection is assessed, and corresponding correction instructions are generated, as follows: The defect impact factor is obtained based on the defect information, using the following formula: ,in As a defect impact factor, Let be the area of ​​the defect. The total area of ​​the image. These are the length and width of the defect, respectively; The illumination variation parameters include the illumination variation factor, and the specific formula is as follows: ,in The light variation factor, This represents the change in ambient light intensity. This represents the current light intensity. The degree of reflection interference includes the reflection interference factor, the specific formula of which is as follows: ,in As a reflection interference factor, The area of ​​the reflected light spot. The total area of ​​the image; The positional deviation of components includes a positional deviation factor, the specific formula of which is as follows: ,in For position deviation factor, This represents the maximum deviation of the component. This represents the maximum acceptable deviation for the assembly line. The weighted summation of the defect impact factor, illumination change factor, reflection interference factor, and position deviation factor yields the following formula for calculating the comprehensive instability risk index: ,in To form a comprehensive instability risk index, As a defect impact factor, The light variation factor, As a reflection interference factor, For position deviation factor, These are the preset weighting coefficients for the defect impact factor, illumination change factor, reflection interference factor, and position deviation factor, respectively. All are greater than 0; The comprehensive instability risk index is compared with the preset comprehensive instability risk threshold. If the comprehensive instability risk index is greater than the comprehensive instability risk threshold, a correction needs to be performed; otherwise, the current setting is maintained. The correction commands include illumination adjustment commands, position adjustment commands, and reflection interference correction commands.

7. An intelligent vision inspection method for automotive parts assembly lines, used to implement the intelligent vision inspection system for automotive parts assembly lines as described in any one of claims 1-6, characterized in that: Includes the following steps: The system synchronously acquires timestamped images, positions, orientations, and ambient lighting data of components using high-precision cameras, laser sensors, and light sensors, generating a raw multimodal data stream. The image data in the multimodal data stream is subjected to denoising, automatic exposure adjustment, and gain control; at the same time, polarization filtering and multi-light source synthesis are used to suppress reflected light spots and highlight areas in the image, resulting in a pre-processed and optimized image data stream. The image data is motion-compensated using an inertial measurement unit and conveyor belt motion data to correct positional deviations and distortions caused by component movement. The viewpoint is aligned using a multi-camera system to obtain image data after dynamic compensation and spatial alignment. Deep learning detection algorithms are applied to compensated and aligned image data to extract the spatial position, pose and contour features of parts, and generate parts recognition results with confidence. Based on the component identification results, defect detection is performed on the component surface to identify cracks and scratches, calculate the specific location and size of the defects, and output defect information including defect type, size and relative spatial coordinates of the component. Based on the defect information, illumination variation parameters, reflection interference level, and component position deviation, assess the risk of accuracy instability in visual inspection and generate corresponding correction instructions; Adjust the vision system and operating parameters on the assembly line according to the calibration instructions.