Hardware surface defect online detection method based on multi-modal sensing fusion

By employing a multimodal sensor fusion method, combining optical cameras, line laser scanners, and thermal imaging cameras, efficient and reliable detection of surface defects in hardware parts has been achieved. This addresses the limitations of single sensors in complex environments and improves detection accuracy and adaptability.

CN121746286AInactive Publication Date: 2026-03-27DONGGUAN DONGJIHUA HARDWARE ELECTRONIC TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-03-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing methods for detecting surface defects in hardware parts rely on a single sensor, which suffers from problems such as high reflectivity, interference from complex textures, low optical image contrast, and large variations in ambient lighting. This results in insufficient detection accuracy and robustness, and the lack of an adaptive mechanism makes it impossible to adapt to changes in different hardware parts types and production conditions.

Method used

A multimodal sensor fusion method is adopted, including a high-resolution optical camera, a line laser scanner and a thermal imaging camera. Color image data, 3D point cloud data and temperature distribution data are acquired through synchronous triggers. Reliability assessment and adaptive weighted fusion are performed by combining fuzzy logic method, and defect classification and localization are performed by support vector machine classifier.

Benefits of technology

It improves the accuracy and robustness of surface defect detection for hardware parts, can efficiently identify minute defects under different lighting conditions, adapts to complex production environments, reduces false detections and missed detections, and improves the overall efficiency and adaptability of the production line.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746286A_ABST
    Figure CN121746286A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of surface defect detection, in particular to a hardware surface defect online detection method based on multi-modal sensing fusion, which comprises the following steps: step 1, arranging multi-modal sensor groups above and on the side of a conveyor belt of a hardware production line; 2, preprocessing the multi-modal sensor data acquired in the step 1; 3, extracting features from the multi-modal sensor data preprocessed in the step 2; 4, performing reliability evaluation on the multi-modal features extracted in the step 3; 5, generating a fusion feature vector, wherein adaptive weighted fusion adopts a weighted feature splicing method; and step 6, outputting defect types and position information to a production line control system. The combined use of the multi-modal sensor can comprehensively detect the surface of the hardware from multiple dimensions, and the accuracy and stability of the system are further improved through reliability evaluation and weighted fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of surface defect detection technology, and in particular to an online detection method for surface defects in hardware parts based on multimodal sensor fusion. Background Technology

[0002] Hardware components, as fundamental elements in industrial manufacturing, are widely used in automobiles, aerospace, electronic equipment, and other fields. Their surface quality directly affects product performance and lifespan. During hardware production, surface defects such as scratches, dents, rust, and cracks are difficult to avoid, thus requiring efficient online inspection methods to ensure quality. In existing technologies, surface defect detection in hardware components mainly relies on single-sensor methods, especially optical vision-based methods, often classified under G01N21 / 88, which refers to detecting surface defects through optical means. These methods typically use high-resolution industrial cameras to acquire images of the hardware component surface and then identify defects using image processing algorithms (such as edge detection and texture analysis). However, single-optical inspection methods have significant limitations: First, hardware surfaces often have high reflectivity or complex textures, leading to severe noise interference in the image and easily masking defect features; second, certain defect types (such as microcracks or shallow dents) have low contrast in optical images, making them difficult to detect reliably; finally, in online inspection environments, high production line speeds and large variations in ambient lighting further reduce the accuracy and robustness of the inspection.

[0003] To overcome the limitations of single sensors, multi-sensor fusion methods have emerged in existing technologies, such as combining vision sensors and laser sensors. However, these methods often employ simple feature stitching or weighted averaging when fusing multimodal data, failing to fully consider the complementarity and reliability differences of different sensors in specific defect detection. For example, vision sensors are sensitive to color and texture changes but not to depth information; laser sensors can provide depth data but are greatly affected by surface materials. Existing fusion methods lack adaptive mechanisms and cannot dynamically adjust fusion strategies according to real-time detection conditions, resulting in poor fusion performance in complex production line environments, with still high false positive and false negative rates. Furthermore, existing methods often rely on preset parameters or empirical thresholds in feature extraction and decision-making processes, lacking a self-evaluation process for sensor data quality, making it difficult to adapt to changes in different hardware types and production conditions.

[0004] Therefore, there is an urgent need for a highly innovative multimodal sensing fusion method that can achieve efficient and reliable online detection of surface defects in hardware parts through the linkage and combination of technical features. Summary of the Invention

[0005] To achieve the above objectives, this invention provides an online detection method for surface defects in hardware parts based on multimodal sensor fusion, comprising the following steps:

[0006] Step 1: Arrange a multimodal sensor group above and to the side of the conveyor belt in the hardware production line. The multimodal sensor group includes a high-resolution optical camera, a line laser scanner and a thermal imaging camera. The high-resolution optical camera, line laser scanner and thermal imaging camera are controlled by a synchronous trigger to collect multimodal sensor data on the surface of the hardware at the same time. The multimodal sensor data includes color image data, three-dimensional point cloud data and temperature distribution data.

[0007] Step 2: Preprocess the multimodal sensor data acquired in Step 1. The preprocessing includes grayscale conversion and histogram equalization of color image data, denoising and registration of 3D point cloud data, and normalization of temperature distribution data.

[0008] Step 3: Extract features from the multimodal sensor data preprocessed in Step 2. Feature extraction includes extracting texture and color features from color image data, extracting geometric features from 3D point cloud data, and extracting thermal features from temperature distribution data. All features are converted into feature vectors of a uniform dimension.

[0009] Step 4: Perform a reliability assessment on the multimodal features extracted in Step 3. The reliability assessment is based on the quality indicators of the sensor data, including the image signal-to-noise ratio and contrast value of the high-resolution optical camera, the point cloud density and integrity value of the line laser scanner, and the temperature stability and resolution value of the thermal imaging camera. The reliability weight of each sensor is calculated using the fuzzy logic method.

[0010] Step 5: Adaptively weightedly fuse the multimodal features extracted in Step 3 with the reliability weights obtained in Step 4 to generate a fused feature vector. The adaptive weighted fusion adopts a weighted feature concatenation method.

[0011] Step 6: Based on the fused feature vector obtained in Step 5, use a support vector machine classifier to classify and locate defects, and output the defect type and location information to the production line control system.

[0012] Preferably, the specific process of multimodal sensor data acquisition in step 1 includes:

[0013] A high-resolution optical camera acquires color image data of the surface of the hardware parts. The high-resolution optical camera is equipped with an adaptive lighting system, which dynamically adjusts the intensity and angle of the light source according to the ambient lighting conditions to minimize reflection interference.

[0014] Line laser scanners collect three-dimensional point cloud data of the surface of hardware parts. The line laser scanner obtains the three-dimensional coordinates of the surface through the principle of laser triangulation. The scanning frequency of the line laser scanner is synchronized with the speed of the production line conveyor belt to ensure that the data covers the entire surface of the hardware parts.

[0015] The thermal imaging camera collects temperature distribution data on the surface of the hardware parts. The thermal imaging camera captures surface thermal radiation through an infrared sensor. The sampling rate of the thermal imaging camera is set higher than the operating frequency of the production line in order to capture transient temperature changes.

[0016] The synchronization trigger uses a hardware-level synchronization signal, which is generated by the main controller of the production line to ensure that the high-resolution optical camera, line laser scanner and thermal imaging camera simultaneously trigger data acquisition when the hardware passes through the inspection area.

[0017] After data acquisition, the multimodal sensor data is transmitted to the processing system in real time through a high-speed data interface, including Gigabit Ethernet or CameraLink interface.

[0018] Preferably, the specific process of multimodal sensor data preprocessing in step 2 includes:

[0019] The color image data is converted to grayscale, transforming the color image from the RGB color space to a grayscale image. The grayscale conversion uses a weighted average method, which assigns weights to the red, green, and blue channels based on human eye sensitivity.

[0020] Histogram equalization is applied to grayscale images. Histogram equalization enhances contrast by redistributing the pixel intensity values ​​of the image. Histogram equalization uses a cumulative distribution function to map pixel values.

[0021] Gaussian filtering is applied to the grayscale image. The Gaussian filtering uses a two-dimensional Gaussian convolution kernel. The size of the two-dimensional Gaussian convolution kernel is adaptively selected according to the image noise level to smooth the image and preserve edge information.

[0022] The 3D point cloud data is denoised using a statistical outlier removal algorithm. This algorithm calculates the average distance between each point in the point cloud and its neighbors and removes outliers whose distance exceeds a threshold.

[0023] The 3D point cloud data is registered using the iterative nearest point algorithm, which aligns the current point cloud with the reference point cloud to correct the pose error caused by the movement of the hardware.

[0024] The temperature distribution data is normalized. The normalization process linearly maps the original temperature values ​​to the range of zero to one. The normalization process is based on the maximum and minimum values ​​of historical temperature data.

[0025] Median filtering is applied to the temperature distribution data using a sliding window method. The size of the sliding window is adjusted according to the rate of temperature change to eliminate impulse noise.

[0026] Preferably, the specific process of multimodal sensor feature extraction in step 3 includes:

[0027] Texture features are extracted from color image data. The texture features are calculated using the Local Binary Pattern Algorithm (LCA). The LCA divides the image into multiple regions and calculates the LCA histogram for each region to obtain the texture complexity value.

[0028] Color features are extracted from color image data by converting the image to the HSV color space and calculating histogram statistics of the saturation components, including the mean, variance, and skewness.

[0029] Geometric features are extracted from 3D point cloud data. These features include surface curvature values ​​and height variance. Surface curvature values ​​are obtained by calculating the rate of change of the normal vector of each point in the point cloud, and height variance is obtained by calculating the standard deviation of the point cloud in the vertical direction.

[0030] Thermal features are extracted from temperature distribution data. These features include temperature gradient and regional average temperature. The temperature gradient is obtained by calculating the temperature difference between adjacent pixels, and the regional average temperature is obtained by dividing the temperature image into grids and calculating the average value of each grid.

[0031] All extracted features are normalized to the same scale. The normalization uses the min-max scaling method, which maps the feature values ​​to the range of zero to one.

[0032] The normalized feature combination is a feature vector of uniform dimension. The dimension of the feature vector is equal to the sum of the number of all features. The feature vector serves as the input for the subsequent fusion step.

[0033] Preferably, the specific process of multimodal sensor reliability assessment in step 4 includes:

[0034] For high-resolution optical cameras, the image signal-to-noise ratio is obtained by calculating the variance ratio of pixel values ​​in smooth regions and edge regions in the image. Smooth regions are selected from uniform areas on the surface of the hardware, and edge regions are selected from areas where defects may occur. The contrast value is obtained by calculating the standard deviation of the image gray levels.

[0035] For line laser scanners, point cloud density is obtained by calculating the number of points in a unit area, which is determined based on the surface size of the hardware. Point cloud integrity is obtained by calculating the proportion of data missing areas, which are identified by a hole detection algorithm in the point cloud.

[0036] For thermal imaging cameras, temperature stability is obtained by calculating the standard deviation of temperature changes in consecutive frames. The number of consecutive frames is set according to the production line speed. Temperature resolution is obtained by the minimum resolvable temperature difference of the thermal imaging camera, which is determined by the sensor specifications of the thermal imaging camera.

[0037] The quality indicators are normalized to the range of zero to one. The normalization uses a linear function, which is based on the maximum and minimum values ​​of historical data.

[0038] The reliability weight is calculated using a fuzzy logic method. The fuzzy logic method includes defining input variables and output variables. The input variables are quality indicators, and the output variables are reliability weights. The fuzzy logic method sets up a fuzzy rule base, which defines the mapping relationship between quality indicators and reliability weights based on expert experience.

[0039] The defuzzification of the fuzzy logic method adopts the centroid method. The centroid method calculates the weighted average of the output variables to obtain the reliability weight value of each sensor. The reliability weight value ranges from zero to one.

[0040] Preferably, the specific process of adaptive weighted fusion of multimodal features in step 5 includes:

[0041] The weighted feature concatenation method first multiplies the feature vector of each sensor by the corresponding reliability weight to obtain a weighted feature vector. The weighted feature vector is calculated by multiplying each element of the feature vector by the reliability weight.

[0042] The weighted feature vectors are concatenated into a fused feature vector. The concatenation operation is performed according to the sensor order, which is high-resolution optical camera, line laser scanner and thermal imaging camera.

[0043] The dimension of the fused feature vector is equal to the sum of the dimensions of all weighted feature vectors, and the fused feature vector serves as the input to the defect decision-making step.

[0044] During the splicing process, if the feature vector dimensions are inconsistent, the zero-padding method is used to align the dimensions. The zero-padding method adds zero values ​​to the end of the shorter feature vector.

[0045] Adaptive weighted fusion also includes a fusion validity check, which is performed by calculating the variance of the fused feature vector. If the variance is lower than a threshold, the reliability assessment step is re-executed. The threshold is set based on historical fusion data.

[0046] The fused feature vectors are stored in a cache, the size of which is set according to the production line processing speed to ensure real-time processing.

[0047] Preferably, the specific process of defect decision-making and output in step 6 includes:

[0048] The support vector machine classifier is trained offline using historical data, which includes labeled surface defect data of hardware parts. The training process uses a kernel function, which is a radial basis function.

[0049] During the online phase, the support vector machine classifier classifies the fused feature vectors and outputs the defect type, which includes scratches, dents, rust, and cracks.

[0050] For samples classified as defects, defect localization is performed. Defect localization is achieved by back-mapping the fused feature vector to the original sensor data to locate the specific location of the defect on the surface of the hardware. The back-mapping adopts a coordinate transformation method based on the sensor calibration parameters.

[0051] The defect location information is output to the production line control system, which triggers the sorting mechanism to remove the defective part. The sorting mechanism includes pneumatic push rods or robotic arms.

[0052] The test results are recorded in real time to the database, which stores the defect type, location and timestamp for quality traceability and process optimization.

[0053] The support vector machine classifier is updated periodically, with the update cycle set according to the production line output. During the update process, the classifier is retrained using newly collected labeled data.

[0054] Preferably, the specific implementation of the synchronous trigger in step 1 includes:

[0055] The synchronous trigger receives pulse signals from the encoder of the production line. The pulse signals are synchronized with the movement distance of the conveyor belt, and the frequency of the pulse signals is adaptively adjusted according to the speed of the conveyor belt.

[0056] The synchronization trigger generates a synchronization signal, which is simultaneously sent to the high-resolution optical camera, the line laser scanner, and the thermal imaging camera. The synchronization signal is triggered by the rising edge.

[0057] Upon receiving the synchronization signal, the high-resolution optical camera, line laser scanner, and thermal imaging camera immediately begin data acquisition, and the data acquisition timestamp is recorded in the metadata.

[0058] The synchronous trigger also monitors the timing consistency of data acquisition. If any sensor fails to respond within a predetermined time, the synchronous trigger issues a warning signal, which notifies the operator to check the sensor status.

[0059] The hardware circuit of the synchronous trigger is implemented using an FPGA, which is programmed and configured to use a multi-output synchronous mode to ensure that signal delay is minimized.

[0060] Preferably, the preprocessing in step 2 also includes a data quality assessment and correction process:

[0061] Data quality assessment calculates a quality score for each sensor's data. The quality score is based on signal-to-noise ratio, integrity, and consistency metrics. The signal-to-noise ratio is calculated by the power ratio of the signal to the noise, the integrity is calculated by the data missing rate, and the consistency is calculated by the spatiotemporal alignment error between data from multiple sensors.

[0062] If the quality score is below the threshold, data correction is performed. Data correction uses interpolation or reconstruction methods. For color image data, bilinear interpolation is used to fill in missing areas; for 3D point cloud data, surface fitting algorithm is used to recover missing points; for temperature distribution data, nearest neighbor interpolation is used to compensate for outliers.

[0063] Data quality assessment and correction are performed in real time during the preprocessing step, and the corrected data replaces the original data for subsequent feature extraction.

[0064] The quality score is recorded in a log file for subsequent analysis and system optimization.

[0065] Preferably, the defect decision in step 6 further includes a multi-level verification process:

[0066] The multi-level verification process first uses a support vector machine classifier for preliminary classification, and then uses a convolutional neural network for fine classification. The structure of the convolutional neural network includes convolutional layers, pooling layers and fully connected layers. The convolutional neural network is trained offline.

[0067] For samples whose classification results from the support vector machine classifier and the convolutional neural network are inconsistent, manual review is performed. The manual review displays the defect image and location through the operator interface, and the operator inputs the final label.

[0068] The multi-level verification process also includes defect severity assessment, which is based on defect size and location. Defect size is calculated by the number of pixels or point cloud area, and defect location is assessed by the distance from the edge of the hardware.

[0069] The severity assessment results, along with the defect type, are output to the production line control system for priority sorting.

[0070] The data flow of the multi-level verification process is executed in parallel with the main process to ensure real-time performance. The parallel execution adopts multi-threading technology.

[0071] The beneficial effects of this invention are:

[0072] 1. This invention combines visual sensors with laser sensors and other different types of sensors through a multimodal sensor fusion method, overcoming the shortcomings of optical inspection in detecting high reflectivity, complex textures, and minute defects. The laser sensor provides depth information, effectively supplementing the visual sensor's lack of depth information, enhancing the detection capability for various surface defects, and ensuring high accuracy and robustness under different lighting conditions.

[0073] 2. The multimodal sensor fusion method proposed in this invention has an adaptive mechanism, enabling it to dynamically adjust the fusion strategy based on real-time detection conditions. By intelligently analyzing the data characteristics of different sensors and combining their advantages in specific defect detection, the fusion process is optimized, improving overall detection accuracy. In changing production environments, the system can automatically adjust, reducing false detections and missed detections, and improving the system's adaptability and stability.

[0074] 3. The multimodal fusion method in this invention can achieve self-assessment of sensor data quality and intelligently adjust parameters and decision-making processes based on the assessment results. By introducing an adaptive adjustment mechanism, the system can automatically optimize the detection process according to changes in different hardware types and production environments, avoiding the limitations of traditional methods that rely on fixed thresholds and empirical parameters, enabling the system to operate stably under various complex production conditions.

[0075] 4. This invention employs efficient multi-sensor data fusion and processing technology, enabling real-time and efficient defect detection on high-speed production lines. Through parallel processing and intelligent optimization, the detection system ensures real-time performance and accuracy even during high-speed production line operation, significantly improving the overall efficiency and adaptability of the production line.

[0076] 5. This invention achieves multi-dimensional and comprehensive detection of surface defects in hardware parts by fusing data from different types of sensors. Each sensor can provide complementary information on different types of defects, thereby comprehensively improving the coverage and accuracy of defect detection and ensuring that even difficult-to-detect defects such as micro-cracks and shallow pits can be accurately identified. Attached Figure Description

[0077] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0078] Figure 1 This is a flowchart of the steps of the method of the present invention;

[0079] Figure 2This is a flowchart illustrating the specific steps of multimodal sensor data acquisition in step 1 of the method of the present invention.

[0080] Figure 3 This is a flowchart illustrating the specific steps of multimodal sensor feature extraction in step 3 of the method of the present invention. Detailed Implementation

[0081] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should also be noted that, to make the embodiments more comprehensive, the following embodiments are the best and preferred embodiments, and those skilled in the art can use other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0082] Please see Figures 1-3 This invention provides an online detection method for surface defects in hardware parts based on multimodal sensor fusion. In step 1, a multimodal sensor group is first arranged above and to the side of the conveyor belt in the hardware production line. This sensor group includes a high-resolution optical camera, a line laser scanner, and a thermal imaging camera. These three sensors are responsible for collecting different types of surface data: the optical camera acquires color image data, the line laser scanner acquires three-dimensional point cloud data of the hardware surface, and the thermal imaging camera provides temperature distribution data. Through synchronous trigger control, these sensors can collect data simultaneously, ensuring data consistency and avoiding errors caused by time differences in acquisition. By using multiple sensors in combination, surface information of the hardware parts can be obtained from multiple dimensions, thus providing more reference information for subsequent defect detection.

[0083] In step 2, after data acquisition, preprocessing is performed on different types of sensor data. Color image data first undergoes grayscale conversion and histogram equalization, which helps improve image contrast, making defects more visually prominent and reducing the impact of lighting variations. 3D point cloud data undergoes denoising and registration to eliminate noise and ensure spatial consistency across different sensor data, thereby improving data reliability. Temperature distribution data is normalized to ensure a uniform scale, facilitating subsequent feature extraction and fusion. This preprocessing step ensures data quality and improves the accuracy and stability of subsequent analysis.

[0084] Step 3, extracting features from the preprocessed data, is a crucial step in this method. First, texture and color features are extracted from the color image. Texture features reflect changes in the surface's microstructure, while color features capture anomalous color variations. Second, geometric features, such as surface curvature and shape variations, are extracted from the 3D point cloud data. These features accurately describe surface irregularities. Finally, thermal features are extracted from the thermal imaging data. The heatmap reflects temperature differences on the surface, further revealing defect areas. All these features are uniformly converted into feature vectors of the same dimension for subsequent processing. Through multi-dimensional feature extraction, the system can comprehensively capture the characteristics of different defects, enhancing the comprehensiveness and accuracy of defect detection.

[0085] In step 4, to improve the reliability of detection, this step performs a reliability assessment on the extracted multimodal features. First, the reliability of each sensor is evaluated based on its data quality metrics (such as the image signal-to-noise ratio of a high-resolution optical camera, the point cloud density of a line laser scanner, and the temperature stability of a thermal imaging camera). Using fuzzy logic, a reliability weight for each sensor is calculated based on the quality of the sensor data. This evaluation mechanism dynamically adjusts the role of different sensors in defect detection, enhancing the system's adaptability to different data sources. In this way, the system can automatically adjust and prioritize sensors with higher data quality, improving the overall accuracy of detection.

[0086] Following the evaluation in step 4, step 5 performs adaptive weighted fusion by combining the reliability weights of each sensor. This weighted concatenation of features from different sensors generates a fused feature vector. This weighted fusion method dynamically adjusts the fusion approach based on the quality of the sensor data, ensuring a reasonable and balanced contribution from different sensors. This process not only effectively integrates the advantages of different sensors but also minimizes errors caused by poor data quality from certain sensors, improving the robustness and accuracy of defect detection.

[0087] In step 6, a Support Vector Machine (SVM) classifier is used to classify and locate defects based on the fused feature vectors. SVM can efficiently process high-dimensional feature data and find the optimal boundary during classification, thus achieving high-precision defect classification. The system outputs the defect type and its location, feeding it back to the production line control system in real time for timely handling. It can accurately identify different types of defects and pinpoint their precise locations, providing a reliable basis for subsequent quality control and automated processing.

[0088] By introducing multiple sensors and an adaptive data fusion mechanism, the limitations of traditional detection methods are overcome, providing an efficient, accurate, and robust defect detection solution. The combined use of multimodal sensors enables comprehensive inspection of hardware surfaces from multiple dimensions. Reliability assessment and weighted fusion further enhance the system's accuracy and stability, enabling it to adapt to defect detection needs under different production environments and conditions.

[0089] In one possible implementation, a high-resolution optical camera first acquires color image data of the metal part's surface. To avoid image quality fluctuations due to changes in ambient lighting conditions, the optical camera is equipped with an adaptive illumination system. This system dynamically adjusts the light source based on real-time ambient light intensity and angle to minimize reflections and interference. This dynamic adjustment mechanism ensures high-quality image data is obtained under different lighting conditions, avoiding common problems such as light spots, shadows, or reflections, thereby improving the accuracy of subsequent defect identification.

[0090] Line laser scanners acquire 3D point cloud data of hardware parts' surfaces using the principle of laser triangulation. Laser triangulation utilizes the geometric relationships of triangles formed when a laser beam illuminates the object's surface to calculate the 3D coordinates of each point on the surface. To ensure complete coverage of the hardware parts' surfaces, the line laser scanner's scanning frequency is synchronized with the production line conveyor belt speed. This means that regardless of changes in conveyor belt speed, the scanner can accurately capture complete surface data for each hardware part, avoiding missed defects due to incomplete scanning.

[0091] The thermal imaging camera is responsible for collecting temperature distribution data on the surface of the hardware parts. It captures thermal radiation from the object's surface using an infrared sensor to generate a thermal image. To capture potential transient temperature changes, the sampling rate of the thermal imaging camera is set higher than the frequency of the production line operation. This high sampling rate ensures that thermal changes are accurately captured as the hardware parts pass through the inspection area, allowing for timely detection of temperature anomalies caused by material defects or processing problems, thereby further improving the sensitivity of defect detection.

[0092] All sensor data acquisition is coordinated by a synchronization trigger, which uses a hardware-level synchronization signal generated by the production line main controller. This ensures that the high-resolution optical camera, line laser scanner, and thermal imaging camera can simultaneously trigger data acquisition as the hardware passes through the inspection area. The use of hardware-level synchronization signals effectively eliminates time skew between data from different sensors, ensuring consistency and synchronization of data acquired from different modal sensors, thereby avoiding potential problems caused by data acquisition time differences.

[0093] After data acquisition, all multimodal sensor data is transmitted to the processing system in real time via a high-speed data interface. To ensure high-speed, high-volume data transmission, the system employs Gigabit Ethernet or CameraLink interfaces. These interfaces feature high bandwidth and low latency, supporting real-time data transmission, ensuring efficient processing and timely feedback of the entire data stream, and preventing data transmission delays from affecting defect detection efficiency.

[0094] By combining high-resolution images, 3D point cloud data, and temperature distribution data, multi-dimensional and comprehensive surface information of hardware parts can be provided. The adaptive lighting system and high sampling rate design enable each sensor to maintain efficient and accurate data acquisition capabilities under different operating conditions. Hardware-level synchronous triggering and high-speed data transmission interfaces ensure data real-time performance and consistency, further enhancing the reliability and response speed of the detection system.

[0095] In one possible implementation, the color image data is first converted to grayscale during data preprocessing. This step converts the original RGB color image into a grayscale image to simplify image processing and improve computational efficiency. A weighted average method is used during the conversion, assigning different weights to the red, green, and blue channels based on the human eye's sensitivity to different colors. This ensures that the proportions of red, green, and blue reflect the human eye's perception of these colors, making the converted grayscale image more visually accurate and facilitating subsequent image processing and defect detection.

[0096] Next, histogram equalization is performed on the grayscale image. This method redistributes the pixel intensity values, thereby enhancing image contrast. This process uses a cumulative distribution function to remap pixel values, making the differences between dark and bright areas of the image more pronounced and revealing details more clearly. This step is crucial for improving defect detection capabilities under uneven lighting or low-contrast conditions, helping the system identify more potential surface defects.

[0097] To remove noise from the image while preserving crucial edge information, the next step is to apply Gaussian filtering to the grayscale image. Gaussian filtering smooths the image using a two-dimensional Gaussian convolution kernel, reducing noise. The size of the filter kernel adaptively adjusts according to the noise level of the image to ensure effective noise reduction without affecting important edge details. This helps enhance image quality and improve the accuracy of subsequent defect detection.

[0098] For 3D point cloud data, denoising is first performed using a statistical outlier removal algorithm. This algorithm calculates the average distance between each point and its neighbors, removing outliers whose distance exceeds a set threshold, thus eliminating abnormal values ​​or noise in the point cloud data and ensuring data accuracy. Next, the iterative nearest-point algorithm is used to register the point cloud data. This algorithm aligns the currently acquired point cloud with a reference point cloud, correcting pose errors caused by changes in the position of the hardware components. Point cloud data registration ensures the consistency and accuracy of data acquired at each time step, facilitating subsequent defect analysis.

[0099] When preprocessing temperature distribution data, normalization is performed first, linearly mapping the original temperature values ​​to the range of 0 to 1. This eliminates data bias caused by variations in temperature across different devices or the environment. Normalization is based on the maximum and minimum values ​​of historical temperature data, ensuring comparability between temperature data from different devices. Next, median filtering is used to further process the temperature data, smoothing the temperature curve using a sliding window approach to remove outliers caused by impulse noise. The size of the sliding window is dynamically adjusted according to the rate of temperature change, thus preserving the true characteristics of temperature variation.

[0100] Grayscale conversion and histogram equalization help improve image quality, making defects more obvious; Gaussian filtering effectively removes noise while preserving the image's edge features; denoising and registration of point cloud data ensure the accuracy and consistency of 3D data; and normalization and median filtering of temperature data further enhance the accurate capture of temperature changes. These steps work together to enable the entire multimodal sensor fusion system to better cope with complex production environments, thereby improving the accuracy and robustness of surface defect detection in hardware parts.

[0101] In one possible implementation, texture features are first extracted from the color image data. These features are calculated using the Local Binary Pattern (LBP) algorithm. This algorithm segments the image into multiple small regions and calculates a local binary pattern histogram for each region. The LBP algorithm generates a binary value by comparing the brightness values ​​of each pixel with those of its surrounding neighboring pixels and calculates its histogram distribution, thereby obtaining the texture complexity value of the image. These texture features can effectively describe the surface structure information of the image and are very helpful in identifying surface defects in hardware parts (such as scratches, dents, etc.).

[0102] Next, color features are extracted by converting the color image to the HSV (Hue, Saturation, Lightness) color space. The HSV color space is more consistent with human visual perception than the RGB color space, making it suitable for color information extraction. The extraction process includes calculating the histogram statistical features of the saturation component, with statistical values ​​including mean, variance, and skewness. These statistical values ​​help quantify the color distribution characteristics of the image, highlighting surface color variations in certain situations and aiding in the detection of color unevenness or blemishes caused by material defects.

[0103] For 3D point cloud data, extracting geometric features is crucial. Geometric features include surface curvature values ​​and height variance. Surface curvature is obtained by calculating the rate of change of the normal vector at each point in the point cloud; it describes the degree of surface bending and helps identify geometric defects such as protrusions or depressions. Height variance is obtained by calculating the standard deviation of the point cloud in the vertical direction; it reflects the degree of variation in surface height and is important for detecting surface unevenness, fluctuations, and other defects.

[0104] Thermal features extracted from temperature distribution data include temperature gradient and regional average temperature. The temperature gradient, obtained by calculating the temperature difference between adjacent pixels, reflects surface temperature variations; temperature unevenness is often associated with surface defects (such as cracks or porosity). The regional average temperature is obtained by dividing the temperature image into multiple grids and calculating the average temperature of each grid. This helps detect temperature anomalies in localized areas of the surface and further identifies defective regions.

[0105] After extracting all the features described above, all features need to be normalized. The min-max scaling method is used to map the values ​​of different features to a uniform range of 0 to 1. This process avoids bias during model training caused by differences in feature scales, ensuring that each feature has the same importance in subsequent analysis.

[0106] Finally, all normalized features are combined into a single feature vector with a dimension equal to the sum of the number of extracted features. This feature vector will serve as input for subsequent data fusion steps, supporting defect identification and classification in machine learning or deep learning models.

[0107] This invention comprehensively captures different features of the surface and temperature distribution of hardware parts, enabling the system to accurately analyze surface defect information from multiple dimensions (texture, color, geometry, and temperature). By combining data from different types of sensors, the accuracy and robustness of defect detection can be improved, avoiding the limitations that may arise from relying on data from a single sensor. Feature normalization and vectorization not only improve data processing efficiency but also make subsequent feature fusion and machine learning analysis more efficient, further enhancing the practicality and real-time performance of the online inspection system.

[0108] In one possible implementation, for a high-resolution optical camera, image quality is first evaluated using the image signal-to-noise ratio (SNR). The SNR is calculated by comparing the variance of pixel values ​​in smooth regions and edge regions of the image. Smooth regions refer to areas with uniform, defect-free surfaces, while edge regions are areas that may contain defects. By calculating the ratio of the variance of pixel values ​​in these two regions, the image's sharpness and noise level can be determined. Secondly, the contrast ratio is evaluated by calculating the standard deviation of the image's gray levels, which helps quantify differences in image brightness, thus reflecting the discernibility of surface details.

[0109] For line laser scanners, the first step is to evaluate the point cloud density. Point cloud density refers to the number of points per unit area, and this metric reflects the level of detail in the scan. The size of the unit area is determined based on the actual dimensions of the hardware surface. Secondly, the integrity of the point cloud is assessed by calculating the proportion of missing data areas, which are identified using a hole detection algorithm within the point cloud. The presence of holes in the point cloud may indicate data loss or incomplete scanning during the process; therefore, the integrity of the point cloud reflects the reliability of the scanning equipment.

[0110] For thermal imaging cameras, key evaluation metrics include temperature stability and temperature resolution. Temperature stability is achieved by calculating the standard deviation of temperature changes across consecutive frames; a smaller standard deviation indicates more stable temperature changes and more reliable detection results. The number of consecutive frames is set according to the production line speed to ensure real-time data acquisition. Temperature resolution is evaluated using the minimum resolvable temperature difference of the thermal imaging camera, which is determined by the camera's sensor specifications and reflects the device's ability to detect minute temperature changes.

[0111] All quality assessment metrics are normalized, transforming them to a range of 0 to 1. Normalization uses a linear function based on the maximum and minimum values ​​from historical data, ensuring that quality assessment values ​​from different sensors are standardized to the same benchmark, facilitating subsequent processing and comparison.

[0112] After normalizing all quality metrics, a fuzzy logic method is used to calculate the reliability weight of each sensor. The fuzzy logic method first defines input and output variables, where the input variables are the various quality metrics, and the output variable is the reliability weight of each sensor. A fuzzy rule base defines the mapping relationship between quality metrics and reliability weights based on expert experience and historical data. Using these rules, a reliability weight value can be obtained based on the quality assessment results of each sensor, with the weight value ranging from 0 to 1.

[0113] In fuzzy logic methods, the defuzzification process uses the centroid method. The centroid method calculates a weighted average of the outputs of all fuzzy rules to derive a reliability weight value for each sensor. These weight values ​​reflect the reliability of each sensor in defect detection; higher reliability indicates better data quality from the sensor, while lower reliability suggests potential data distortion or noise interference.

[0114] This multimodal sensor reliability assessment method assigns a reasonable weight to each sensor, giving higher-quality sensors greater influence in subsequent data fusion and defect detection processes. This avoids inaccurate or misjudged results due to poor data quality from individual sensors, improving the overall reliability and accuracy of the detection system. Furthermore, the application of fuzzy logic methods makes the entire assessment process more flexible, enabling it to cope with uncertainties and changes in different production environments, thus improving the system's adaptability and robustness.

[0115] In one possible implementation, firstly, for each sensor (including a high-resolution optical camera, a line laser scanner, and a thermal imaging camera), the feature vector is weighted according to its corresponding reliability weight. Each sensor's feature vector is represented by a set of numbers reflecting the surface features detected by that sensor. Each element of these feature vectors is multiplied by its corresponding reliability weight to obtain a weighted feature vector. The weighted feature vector better reflects the quality of the sensors and the reliability of the data, ensuring that the importance of different sensors is reasonably reflected in the subsequent fusion process.

[0116] Then, all weighted feature vectors are concatenated according to the sensor order: high-resolution optical camera, line laser scanner, and thermal imaging camera. The weighted feature vectors from each sensor are concatenated to form a fused feature vector. The dimension of this fused feature vector is equal to the sum of the dimensions of all weighted feature vectors, representing all feature information from different sensors. This fused feature vector serves as the input to subsequent defect decision-making steps, directly affecting the accuracy of defect detection and the effectiveness of the decision.

[0117] During the stitching process, there may be inconsistencies in the dimensions of feature vectors from different sensors. To ensure smooth fusion, a zero-padding method is used to align shorter feature vectors. Specifically, zero values ​​are added to the end of the feature vectors with lower dimensions to make all feature vectors have the same dimensions. This method is simple and efficient, and can avoid stitching errors caused by inconsistent dimensions.

[0118] To ensure the quality of feature fusion, a fusion validity check is implemented. This check is performed by calculating the variance of the fused feature vector. If the variance of the fused feature vector is lower than a set threshold, it indicates a significant inconsistency problem between the data from different sensors, possibly due to anomalies or poor quality in the output data of a particular sensor. In this case, the system will re-execute the reliability assessment step, readjusting the weights of each sensor until the fusion result meets the expected validity standard. The threshold is adjusted based on historical fusion data to ensure the adaptability and flexibility of the detection process.

[0119] The fused feature vectors are stored in a buffer. The size of the buffer is set according to the processing speed of the production line to ensure that each frame of image data can be stored and processed in a timely manner, guaranteeing real-time performance. The production line typically processes data at a high speed, so the buffer needs to have sufficient storage space to hold multiple frames of data and prevent data loss.

[0120] The embodiments of the present invention can adaptively adjust the weight of the sensors to ensure that the surface defect detection system can maintain high efficiency and high accuracy under different production environments and different defect conditions, thereby greatly improving the robustness and reliability of automated detection.

[0121] In one possible implementation, firstly, in the offline phase, a Support Vector Machine (SVM) classification model is built. The system collects a large amount of historical sample data with defect labels, derived from the fused features of optical cameras, line laser scanners, and thermal imaging cameras. Each sample corresponds to an actual defect type, such as scratches, dents, rust, or cracks. Accurate category labels are assigned to each sample through manual or automatic annotation. During training, the system uses radial basis functions as kernel functions to enhance the classifier's discriminative ability in the nonlinear feature space, thereby effectively distinguishing defects that are similar in appearance but different in nature. After training, the generated SVM model is saved for use in the online detection phase.

[0122] During the online inspection phase, the system inputs the fused feature vectors of the hardware parts acquired in real time into an SVM classifier. Based on the decision boundaries obtained from offline training, the classifier classifies the input samples and outputs the corresponding defect type. When a sample is identified as a defect, the system further performs defect localization. The localization process achieves precise spatial positioning of the defect on the hardware part's surface by back-mapping the fused feature vectors to the original sensor coordinate system. The back-mapping employs a coordinate transformation method based on sensor calibration parameters, ensuring that data acquired by different sensors can correspond to specific physical locations within the same coordinate system, achieving millimeter-level spatial matching accuracy.

[0123] Once the location is determined, the system sends the defect location information to the production line control system. The control system automatically triggers the sorting device based on the location coordinates. Common actuators include pneumatic push rods and robotic arms. Pneumatic push rods are suitable for sorting lightweight parts on high-speed production lines, offering fast response and low cost; robotic arms are suitable for precise gripping and rejection of complex structures or large-sized hardware parts.

[0124] Meanwhile, all inspection results are recorded in real time to the database, including defect type, location coordinates, inspection time, and product number. The database is used not only for quality traceability but also provides data support for subsequent process parameter optimization and model retraining. Based on production line output and the amount of inspection data, the system periodically triggers the SVM model's self-learning update mechanism, retraining the model using newly collected and labeled defect samples to ensure the model's accuracy and adaptability in long-term operation.

[0125] By combining offline training with online detection, the system can stably and quickly identify multiple types of defects, achieving precise spatial localization and automatic sorting. Reverse mapping and calibration coordinate transformation ensure the traceability and reliability of the detection results. Real-time database recording and dynamic model update mechanisms further enhance the system's adaptability, enabling continuous optimization of detection results as processes change. This method effectively improves the automation level of production lines and product quality control capabilities, reduces manual inspection errors and response delays, and achieves efficient defect management in an intelligent manufacturing environment.

[0126] In one possible implementation, the synchronization trigger first receives pulse signals from the production line encoder. The frequency of these pulse signals is proportional to the conveyor belt's speed, meaning each pulse represents a fixed distance the conveyor belt travels. The synchronization trigger uses these pulse signals to monitor the conveyor belt's movement in real time and automatically adjusts the pulse signal frequency based on the conveyor belt's real-time speed. This adaptive adjustment mechanism ensures that the system's synchronization accuracy remains high even when the production line speed changes, adapting to the needs of different production environments.

[0127] Upon receiving a pulse signal, the synchronization signal generated by the synchronization trigger is immediately sent to various sensor devices, including high-resolution optical cameras, line laser scanners, and thermal imaging cameras. These devices accurately begin data acquisition when triggered by the rising edge of the synchronization signal. This rising edge triggering method ensures instantaneous signal response, avoiding false triggering caused by changes in the signal waveform, thereby improving the accuracy and consistency of data acquisition. Simultaneously, all acquired data is timestamped and recorded in metadata for subsequent defect detection and data analysis.

[0128] To ensure the timing consistency of data acquisition, the synchronization trigger also monitors the response time of each sensor. If a sensor fails to respond to the synchronization signal within a specified time window, the system will notify the operator to check the sensor's status via a warning signal. This mechanism effectively avoids detection problems caused by sensor malfunctions or data delays.

[0129] From a hardware implementation perspective, the synchronous trigger is implemented using FPGA (Field-Programmable Gate Array) technology. FPGA provides high-speed parallel processing capabilities, minimizing the generation and transmission delays of the synchronization signal. By being programmed into a multi-output synchronization mode, the FPGA can simultaneously control the trigger signals of multiple sensors, ensuring that multiple devices complete data acquisition in a very short time and reducing signal latency within the system.

[0130] First, the adaptive pulse signal adjustment based on the production line encoder can accurately match the conveyor belt's speed, ensuring synchronized data acquisition from all sensors and avoiding detection errors caused by data asynchrony between different sensors. Second, the application of FPGA technology ensures high-speed and low-latency signal transmission, improving real-time performance and responsiveness. Finally, the system can monitor sensor status and issue timely warnings when anomalies occur, reducing manual intervention and improving the system's automation level and reliability.

[0131] In one possible implementation, data quality assessment first evaluates the data from each sensor, calculating a comprehensive quality score. This quality score consists of three main metrics:

[0132] Signal-to-noise ratio (SNR): This measure assesses the clarity and quality of a signal by calculating the power ratio of the signal to the noise. A higher SNR indicates a larger proportion of valid signal in the sensor data, less noise, and higher data quality.

[0133] Completeness: The completeness of data is judged by the missing data rate. If too much data is missing in a certain part, the completeness score of that data will be low, which may affect subsequent defect detection. Missing parts are usually supplemented by interpolation methods.

[0134] Consistency: The consistency of data from different sensors is evaluated by calculating the spatiotemporal alignment error between data from multiple sensors. If the alignment error between multiple sensors is too large in time and space, it indicates a problem with data fusion, which may lead to inaccurate feature vectors after fusion, thus affecting the identification of defects.

[0135] When the calculated quality score falls below a preset threshold, the system initiates a data correction process. There are several methods for data correction; the specific method chosen depends on the data type.

[0136] For color image data, bilinear interpolation is used to fill in missing areas. Bilinear interpolation can estimate the pixel value of the missing area by using a weighted average of the surrounding pixels, thus avoiding severe distortion of the image data.

[0137] For 3D point cloud data, a surface fitting algorithm is used to recover missing points. The surface fitting algorithm can generate a smooth surface from the existing point cloud data and estimate the location of missing points, ensuring the spatial coherence and accuracy of the point cloud data.

[0138] For temperature distribution data, the nearest neighbor interpolation method is used to compensate for outliers. Nearest neighbor interpolation fills in the missing data by selecting the value of the nearest known point. It is suitable for continuously changing data types such as temperature distribution and can effectively compensate for abrupt changes or anomalies in temperature data.

[0139] The entire data quality assessment and correction process is performed in real time. During the preprocessing stage, sensor data is evaluated and corrected to ensure that the corrected data can replace the original data for subsequent feature extraction and defect detection. In this way, the system can adjust and optimize data quality in real time during operation, guaranteeing high accuracy of the detection results.

[0140] Finally, the quality score is recorded in a log file, providing a basis for subsequent system optimization and data analysis. By periodically analyzing these log files, the system can understand the trend of data quality changes and adjust sensor configuration, data acquisition methods, or algorithms as needed, further improving the reliability and efficiency of detection.

[0141] By evaluating the quality of data from each sensor, low-quality data can be effectively identified, preventing it from negatively impacting subsequent processing. Low-quality data can be corrected using interpolation or reconstruction methods to ensure data integrity and consistency, thereby improving the accuracy of defect detection. At the same time, real-time quality evaluation and correction enable the system to adapt to environmental changes, improving its application effect in dynamic production lines and ensuring that the production line's inspection system always maintains efficient and stable operation.

[0142] In one possible implementation, a Support Vector Machine (SVM) classifier is first used to perform preliminary classification of surface defects on hardware parts. SVM is a powerful supervised learning algorithm that distinguishes different categories of defect data by constructing a hyperplane. Due to its good classification performance, SVM can effectively handle both linear and nonlinear problems and has good generalization ability.

[0143] Next, a Convolutional Neural Network (CNN) is used to refine the initial classification results of the SVM. A CNN is a deep learning model typically composed of multiple convolutional layers, pooling layers, and fully connected layers. Convolutional layers are responsible for extracting local features of the image, pooling layers help reduce computational complexity and enhance the model's robustness, and fully connected layers are used for final classification. In the offline stage, the CNN is trained on a large amount of labeled data to achieve more accurate defect identification.

[0144] For samples where the results from the SVM classifier and the CNN classifier are inconsistent, the system initiates a manual review process. Operators view the defective image and its corresponding location information through an operator interface and input the final label based on their judgment. Manual review provides the system with an opportunity to correct errors, especially in edge cases that are difficult to identify using automated models; human intervention can effectively improve the system's accuracy.

[0145] Besides defect type identification, defect severity assessment is also a crucial component of the multi-level verification process. The assessment primarily relies on the size and location of the defect. Defect size is assessed by calculating the number of pixels in the defect area or the area of ​​the point cloud data; a larger size indicates a more severe defect. Defect location is assessed by calculating the distance between the defect and the edge of the hardware component; a closer distance suggests a greater potential impact. The severity assessment results, along with the defect type, are output to the production line control system for priority sorting, determining the processing order of hardware components based on defect severity.

[0146] To ensure the real-time nature of the defect decision-making process, the data flow executes in parallel with the main workflow. This means that defect decision-making, classification, review, and evaluation processes are synchronized with the production line operation, without causing delays to the production line's real-time performance. To achieve this, the system employs multi-threading technology, allowing each task to execute in an independent thread, avoiding bottlenecks caused by serial processing and ensuring efficient and timely decision-making.

[0147] This multi-level verification process not only effectively improves the accuracy of testing, but also ensures the efficient operation of the production line and the rational use of resources, demonstrating significant technical advantages and application value.

[0148] Example:

[0149] This embodiment uses surface defect detection in an automotive bolt production line as an application scenario. In this scenario, the bolt surfaces are highly reflective, have complex textures, and the defect types include scratches, dents, corrosion, and cracks. The production line conveyor belt speed is 0.5 meters per second, and the detection system needs to process 10 bolts per second to ensure real-time performance.

[0150] Step 1: Multimodal sensor data acquisition;

[0151] A multimodal sensor array, including a high-resolution optical camera, a line laser scanner, and a thermal imaging camera, is arranged above and to the side of the conveyor belt in the bolt production line. The high-resolution optical camera uses a 20-megapixel industrial camera equipped with an adaptive lighting system. This system dynamically adjusts the intensity and angle of the LED light source based on real-time ambient light intensity detected by an ambient light sensor: when the ambient light intensity is below 500 lux, the LED light source intensity increases to 1000 lumens, and the angle is adjusted to a 45-degree incident angle to minimize glare interference on the bolt surface; when the ambient light intensity is above 500 lux, the LED light source intensity decreases to 500 lumens, and the angle is adjusted to a 30-degree incident angle. The line laser scanner uses a scanner based on the principle of laser triangulation, with a scanning frequency set to 1000 Hz, synchronized with the conveyor belt speed to ensure that each bolt surface is completely scanned; the laser wavelength of the line laser scanner is 650 nanometers, and the power is 10 milliwatts to avoid damaging the bolt surface. The thermal imaging camera is an uncooled microbolometer-type thermal imaging camera with a sampling rate set to 200 Hz, which is higher than the production line operating frequency (50 Hz) to capture transient temperature changes on the bolt surface; the spectral range of the thermal imaging camera is 8-14 micrometers, and the temperature resolution is 0.1 degrees Celsius.

[0152] The synchronization trigger employs a hardware-level synchronization signal, generated by the production line's main controller. The main controller integrates an encoder, which emits pulse signals. The frequency of these pulse signals is synchronized with the conveyor belt's movement distance: one pulse is emitted for every 1 millimeter the conveyor belt moves. Upon receiving the pulse signal, the synchronization trigger generates a synchronization signal, triggered by a rising edge, which is simultaneously sent to the high-resolution optical camera, line laser scanner, and thermal imaging camera. The hardware circuitry of the synchronization trigger is implemented using an FPGA (Field-Programmable Gate Array), programmed and configured for multi-output synchronization with an output delay of less than 1 microsecond to minimize signal latency. After data acquisition, multimodal sensor data is transmitted in real-time to the processing system via a Gigabit Ethernet interface, using the GigEVision standard to ensure data integrity and real-time performance.

[0153] If any sensor fails to respond to the synchronization signal within a predetermined time (e.g., 10 milliseconds), the synchronization trigger will issue a warning signal. This warning signal will notify the operator to check the sensor status via the digital output module. The warning signal will also trigger an audible and visual alarm and display specific sensor fault information on the operator interface.

[0154] Step 2: Multimodal sensor data preprocessing;

[0155] The multimodal sensor data acquired in step 1 is preprocessed to eliminate noise and enhance useful information.

[0156] First, the color image data acquired by the high-resolution optical camera is preprocessed. The color image data is converted from the RGB color space to grayscale images using a weighted average method. The weighted average method assigns weights of 0.299 to the red channel, 0.587 to the green channel, and 0.114 to the blue channel based on human eye sensitivity. Next, histogram equalization is performed on the grayscale image. Histogram equalization maps pixel values ​​using a cumulative distribution function: the cumulative distribution function calculates the cumulative probability distribution of the image's grayscale levels and maps the pixel values ​​to the new values, resulting in a uniform distribution of grayscale levels in the output image. After histogram equalization, Gaussian filtering is applied to the image using a two-dimensional Gaussian convolution kernel. The size of the two-dimensional Gaussian convolution kernel is adaptively selected based on the image noise level: by calculating the local variance of the image, if the local variance is higher than a threshold of 100 (the threshold is statistically derived from historical image data and indicates a high noise level), the convolution kernel size is set to 5x5 pixels; if the local variance is lower than the threshold of 100, the convolution kernel size is set to 3x3 pixels. The standard deviation of the two-dimensional Gaussian convolution kernel is set to one-sixth of the kernel size.

[0157] Secondly, the 3D point cloud data acquired by the line laser scanner is preprocessed. The 3D point cloud data is first denoised using a statistical outlier removal algorithm: the algorithm calculates the average distance between each point in the point cloud and its 10 nearest neighbors. If the average distance exceeds twice the standard deviation of the global average distance (the standard deviation is based on the entire point cloud), the point is considered an outlier and removed. Then, the point cloud is registered using an iterative nearest-neighbor algorithm: the iterative nearest-neighbor algorithm aligns the current point cloud with a reference point cloud (the reference point cloud comes from a defect-free standard bolt model). By minimizing the point-to-point distance error, the rotation matrix and translation vector are iteratively optimized. The number of iterations is set to 100, and the error threshold is set to 0.01 mm.

[0158] Finally, the temperature distribution data acquired by the thermal imaging camera is preprocessed. The temperature distribution data undergoes normalization, which linearly maps the original temperature values ​​to a range of 0 to 1: the mapping function is T. norm =(TT) min ) / (T max-T min ), where T is the original temperature value, T min and T max Based on historical temperature data, T in this embodiment is obtained. min At 20 degrees Celsius (ambient temperature), T max The temperature is set to 50 degrees Celsius (the highest temperature on the bolt surface). Then, median filtering is applied to the normalized temperature data. Median filtering uses a sliding window method, and the size of the sliding window is adjusted according to the rate of temperature change: the rate of temperature change of consecutive frames is calculated. If the rate of change is higher than 5 degrees Celsius per second, the window size is set to 5x5 pixels; if the rate of change is lower than 5 degrees Celsius per second, the window size is set to 3x3 pixels.

[0159] Preprocessing also includes data quality assessment and correction. Data quality assessment calculates a quality score for each sensor's data, based on signal-to-noise ratio (SNR), integrity, and consistency metrics. SNR is calculated as the power ratio of signal to noise: for image data, signal power is the square of the average pixel value of the region of interest (the bolt surface area), and noise power is the variance of the background area; for point cloud data, signal power is a function of point cloud density, and noise power is the proportion of outliers; for temperature data, signal power is the variance of temperature values, and noise power is the standard deviation of temperature fluctuations. Integrity is calculated using the missing data rate: the ratio of actual collected data points to expected data points. Consistency is calculated using the spatiotemporal alignment error between multi-sensor data: the alignment error is the maximum deviation of sensor data in time and space. The quality score is calculated as a weighted sum of SNR, integrity, and consistency metrics, with weights of 0.5, 0.3, and 0.2 (weights set based on expert experience). If the quality score is below a threshold of 0.7 (the threshold is based on historical data statistics and represents the minimum acceptable quality), data correction is performed. Data correction employs interpolation or reconstruction methods: For color image data, bilinear interpolation is used to fill in missing regions, calculating the missing pixel value based on the values ​​of the four nearest neighbors; for 3D point cloud data, surface fitting algorithms are used to recover missing points, employing a quadratic surface model and fitting a surface based on neighboring point cloud data; for temperature distribution data, nearest neighbor interpolation is used to compensate for outliers, replacing outliers with the temperature value of the nearest pixel. The corrected data replaces the original data for subsequent feature extraction. Quality scores are recorded in a log file in CSV format, including a timestamp, sensor type, and quality score value.

[0160] Step 3: Feature extraction from multimodal sensors;

[0161] Features related to surface defects are extracted from the multimodal sensor data after preprocessing in step 2.

[0162] Texture and color features are extracted from color image data. Texture features are calculated using the Local Binary Pattern Algorithm (LCA): The LCA divides the image into multiple 8x8 pixel regions and calculates a LCA histogram for each region. The LCA compares the grayscale values ​​of each pixel with its eight neighboring pixels. If a neighboring pixel's value is greater than the center pixel's value, it is marked as 1; otherwise, it is marked as 0, forming an 8-bit binary number, which is then converted to a decimal value as the LCA value for that point. Next, the LCA histogram for each region is calculated. The histogram has 256 buckets (corresponding to an 8-bit binary range). Texture complexity is extracted from the histogram; the texture complexity value is the entropy of the histogram, calculated using the formula: Where p i This represents the frequency of the i-th bucket. Color features are calculated by converting the image to the HSV color space and then calculating the histogram statistics of the saturation components: the saturation component histogram statistics include the mean, variance, and skewness. The formula for calculating the mean is... Where s j This is the saturation value, where N is the number of pixels; the variance calculation formula is... Where μ is the mean; the formula for calculating skewness is: Where σ is the standard deviation.

[0163] Geometric features are extracted from 3D point cloud data, including surface curvature values ​​and height variance. Surface curvature values ​​are obtained by calculating the rate of change of the normal vector at each point in the point cloud: first, principal component analysis is used to calculate the normal vector at each point; the normal vector is the eigenvector corresponding to the smallest eigenvalue of the point cloud covariance matrix; then, the standard deviation of the angle between each point and the normal vectors of its neighbors is calculated as the curvature value. Height variance is obtained by calculating the standard deviation of the point cloud in the vertical direction: the height variance is the variance of the z-coordinate values ​​of the point cloud.

[0164] Thermal features are extracted from the temperature distribution data, including temperature gradient and regional average temperature. The temperature gradient is obtained by calculating the temperature difference between adjacent pixels: the temperature gradient is the maximum temperature difference between each pixel in the image and its four nearest neighbors (top, bottom, left, and right). The regional average temperature is calculated by dividing the temperature image into a 5x5 pixel grid and calculating the average temperature value for each grid cell.

[0165] All extracted features are normalized to the same scale using a min-max scaling method: the min-max scaling method maps feature values ​​to the range of 0 to 1, with the mapping function being x. norm =(xx) min ) / (x max -x min ), where \(x\) are the original eigenvalues, x min and x maxThis is obtained based on statistical analysis of the training data. The normalized feature combination forms a feature vector of uniform dimension, where the dimension of the feature vector is equal to the sum of the number of all features. In this embodiment, the texture feature has 1 value (texture complexity), the color feature has 3 values ​​(mean saturation, variance, skewness), the geometric feature has 2 values ​​(surface curvature, height variance), and the thermal feature has 2 values ​​(temperature gradient, average temperature of the region). Therefore, the feature vector dimension is 8.

[0166] Step 4: Multimodal sensor reliability assessment;

[0167] The reliability of the multimodal features extracted in step 3 is evaluated to determine the reliability weight of each sensor under the current detection conditions.

[0168] Reliability assessment is based on quality metrics derived from sensor data. For high-resolution optical cameras, the image signal-to-noise ratio (SNR) is obtained by calculating the ratio of the variance of pixel values ​​in smooth regions to those in edge regions: the smooth region is selected as a 10x10 pixel area in the center of the bolt surface, and the edge region is selected as a 10x10 pixel area near the bolt outline; the SNR is the ratio of the variance of the smooth region to the variance of the edge region. The contrast ratio is obtained by calculating the standard deviation of image gray levels: the standard deviation of gray levels is the standard deviation of all gray values ​​in the image.

[0169] For line laser scanners, point cloud density is obtained by calculating the number of points per unit area: the unit area is set to 1 square millimeter, and the number of points is obtained by counting the points within that area. Point cloud integrity is obtained by calculating the proportion of missing data regions: missing data regions are identified using a hole detection algorithm in the point cloud. The hole detection algorithm is based on point cloud meshing; if there is no point cloud data in a mesh, it is considered a missing region. The integrity ratio is the ratio of the number of missing meshes to the total number of meshes.

[0170] For the thermal imaging camera, temperature stability is obtained by calculating the standard deviation of temperature changes across consecutive frames: the number of consecutive frames is set to 10 frames (based on production line speed, covering 0.1 seconds), and the standard deviation is calculated as the standard deviation of the temperature difference between consecutive frames. Temperature resolution is obtained by the minimum resolvable temperature difference of the thermal imaging camera: the minimum resolvable temperature difference is determined by the sensor specifications of the thermal imaging camera, and in this embodiment it is 0.1 degrees Celsius.

[0171] The quality index is normalized to a range of 0 to 1. Normalization uses a linear function: the linear function is based on the maximum and minimum values ​​of historical data. For example, if the historical maximum signal-to-noise ratio (SNR) is 50 and the minimum is 5, then the normalized SNR is...

[0172] Reliability weights are calculated using fuzzy logic. The fuzzy logic method defines input and output variables: input variables are quality indicators (signal-to-noise ratio, contrast ratio, point cloud density, point cloud integrity, temperature stability, and temperature resolution), and output variables are reliability weights. The fuzzy sets of the input and output variables are defined as "low," "medium," and "high," with a trigonometric function for membership. For example, the "low" set for signal-to-noise ratio is [0, 0.3], the "medium" set is [0.2, 0.8], and the "high" set is [0.7, 1]. A fuzzy rule base defines the mapping relationship between quality indicators and reliability weights based on expert experience. The rule form is "if the signal-to-noise ratio is high and the contrast ratio is high, then the reliability weight is high." The rule base contains 20 rules. Defuzzification using the fuzzy logic method employs the centroid method: the centroid method calculates the weighted average of the output variables to obtain the reliability weight value for each sensor, with the reliability weight value ranging from 0 to 1.

[0173] Step 5: Adaptive weighted fusion of multimodal features;

[0174] The multimodal features extracted in step 3 are adaptively weighted and fused with the reliability weights obtained in step 4 to generate a fused feature vector.

[0175] Adaptive weighted fusion employs a weighted feature concatenation method. This method first multiplies the feature vector of each sensor by its corresponding reliability weight to obtain a weighted feature vector. The weighted feature vector is calculated as the product of each element in the feature vector and its reliability weight. For example, if the feature vector of a high-resolution optical camera is [0.1, 0.2, 0.3] and the reliability weight is 0.8, then the weighted feature vector is [0.08, 0.16, 0.24].

[0176] The weighted feature vectors are concatenated into a fused feature vector. The concatenation operation is performed according to the sensor order: high-resolution optical camera, line laser scanner, and thermal imaging camera. The dimension of the fused feature vector is equal to the sum of the dimensions of all weighted feature vectors. In this embodiment, each sensor feature vector has a dimension of 8, but the actual number of features varies. Therefore, a zero-padding method is used to align the dimensions: the zero-padding method adds zero values ​​to the end of the shorter feature vectors to make all feature vectors have a consistent dimension of 8.

[0177] Adaptive weighted fusion also includes a fusion validity check, which is performed by calculating the variance of the fused feature vector: the variance calculation formula is... Where x i These are the elements of the fused feature vector, where μ is the mean and N is the dimension. If the variance is below the threshold of 0.01 (the threshold is set based on historical fused data, indicating that the feature variation is too small), the reliability assessment step is re-executed.

[0178] The fused feature vectors are stored in a cache, the size of which is set according to the production line processing speed: the cache is a circular buffer with a size of 100 fused feature vectors to ensure real-time processing.

[0179] Step 6: Defect Decision and Output;

[0180] Based on the fused feature vector obtained in step 5, defect classification and localization are performed.

[0181] Defect decision-making uses a Support Vector Machine (SVM) classifier. The SVM classifier is trained offline using historical data: the historical data includes 1000 labeled bolt surface defects, with labels including scratches, dents, rust, cracks, and no defects. The training process uses a kernel function, specifically a radial basis function (RBF) with the formula K(x). i ,x j )=exp(-γ||x i -x j || 2 The parameter γ is determined through grid search optimization, and in this embodiment, γ is 0.1. The penalty parameter C of the support vector machine classifier is set to 1.0.

[0182] In the online phase, the support vector machine classifier classifies the fused feature vectors and outputs the defect type. For samples classified as defects, defect localization is performed by back-mapping the fused feature vectors back to the original sensor data. This back-mapping uses a coordinate transformation method based on sensor calibration parameters, including camera intrinsic and extrinsic parameters, obtained through the Zhang Zhengyou calibration method. The specific coordinates of the defect location on the bolt surface are calculated through projection.

[0183] The defect location information is output to the production line control system, which triggers a sorting mechanism to remove the defective part. This mechanism includes a pneumatic pusher that moves within 0.1 seconds of receiving a signal, pushing the defective bolt into the scrap bin. The detection results are recorded in real-time to a database (an SQL database) storing the defect type, location, and timestamp.

[0184] The support vector machine classifier is updated regularly, with the update cycle set according to the production line output: every 10,000 bolts produced, the classifier is retrained using newly collected labeled data, which is obtained through manual sampling.

[0185] The defect decision-making process also includes a multi-level verification process. This process first uses a Support Vector Machine (SVM) classifier for initial classification, followed by a Convolutional Neural Network (CNN) for finer classification. The CNN structure consists of convolutional layers, pooling layers, and fully connected layers: the convolutional layers have 32 3x3 kernels using the ReLU activation function; the pooling layers use 2x2 max pooling; the fully connected layers have 128 neurons; and the output layer has 5 neurons corresponding to 5 categories (4 types of defects and no defect). The CNN is trained offline using the same training data as the SVM, with the Adam optimizer and a learning rate of 0.001. For samples where the SVM and CNN classification results are inconsistent, manual verification is performed: this verification is done through an operator interface displaying the defect image and location, with the operator inputting the final label. The multi-level verification process also includes defect severity assessment, based on defect size and location: defect size is calculated using the number of pixels or point cloud area, for example, the size of a scratch defect is the number of pixels in the scratch area; defect location is assessed by its distance from the bolt edge, with a distance less than 1 mm considered severe. Severity assessment results, along with defect type, are output to the production line control system for priority sorting: severe defects are sorted first. The data flow of the multi-level verification process is executed in parallel with the main process to ensure real-time performance. Parallel execution employs multi-threading technology: the main thread handles support vector machine classification, while sub-threads handle convolutional neural network classification and manual review.

[0186] This embodiment details the implementation of all technical features, ensuring the feasibility and effectiveness of the method. Through multimodal sensor fusion and adaptive mechanisms, the accuracy and robustness of bolt surface defect detection are improved, meeting the needs of online inspection on production lines.

[0187] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0188] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An online detection method for surface defects in hardware parts based on multimodal sensor fusion, characterized in that, Includes the following steps: Step 1: Arrange a multimodal sensor group above and to the side of the conveyor belt in the hardware production line. The multimodal sensor group includes a high-resolution optical camera, a line laser scanner and a thermal imaging camera. The high-resolution optical camera, line laser scanner and thermal imaging camera are controlled by a synchronous trigger to collect multimodal sensor data on the surface of the hardware at the same time. The multimodal sensor data includes color image data, three-dimensional point cloud data and temperature distribution data. Step 2: Preprocess the multimodal sensor data acquired in Step 1. The preprocessing includes grayscale conversion and histogram equalization of color image data, denoising and registration of 3D point cloud data, and normalization of temperature distribution data. Step 3: Extract features from the multimodal sensor data preprocessed in Step 2. Feature extraction includes extracting texture and color features from color image data, extracting geometric features from 3D point cloud data, and extracting thermal features from temperature distribution data. All features are converted into feature vectors of a uniform dimension. Step 4: Perform a reliability assessment on the multimodal features extracted in Step 3. The reliability assessment is based on the quality indicators of the sensor data, including the image signal-to-noise ratio and contrast value of the high-resolution optical camera, the point cloud density and integrity value of the line laser scanner, and the temperature stability and resolution value of the thermal imaging camera. The reliability weight of each sensor is calculated using the fuzzy logic method. Step 5: Adaptively weightedly fuse the multimodal features extracted in Step 3 with the reliability weights obtained in Step 4 to generate a fused feature vector. The adaptive weighted fusion adopts a weighted feature concatenation method. Step 6: Based on the fused feature vector obtained in Step 5, use a support vector machine classifier to classify and locate defects, and output the defect type and location information to the production line control system.

2. The online detection method for surface defects of hardware parts based on multimodal sensor fusion according to claim 1, characterized in that, The specific process of multimodal sensor data acquisition in step 1 includes: A high-resolution optical camera acquires color image data of the surface of the hardware parts. The high-resolution optical camera is equipped with an adaptive lighting system, which dynamically adjusts the intensity and angle of the light source according to the ambient lighting conditions to minimize reflection interference. Line laser scanners collect three-dimensional point cloud data of the surface of hardware parts. The line laser scanner obtains the three-dimensional coordinates of the surface through the principle of laser triangulation. The scanning frequency of the line laser scanner is synchronized with the speed of the production line conveyor belt to ensure that the data covers the entire surface of the hardware parts. The thermal imaging camera collects temperature distribution data on the surface of the hardware parts. The thermal imaging camera captures surface thermal radiation through an infrared sensor. The sampling rate of the thermal imaging camera is set higher than the operating frequency of the production line in order to capture transient temperature changes. The synchronization trigger uses a hardware-level synchronization signal, which is generated by the main controller of the production line to ensure that the high-resolution optical camera, line laser scanner and thermal imaging camera simultaneously trigger data acquisition when the hardware passes through the inspection area. After data acquisition, the multimodal sensor data is transmitted to the processing system in real time through a high-speed data interface, including Gigabit Ethernet or CameraLink interface.

3. The online detection method for surface defects of hardware parts based on multimodal sensor fusion according to claim 1, characterized in that, The specific process of multimodal sensor data preprocessing in step 2 includes: The color image data is converted to grayscale, transforming the color image from the RGB color space to a grayscale image. The grayscale conversion uses a weighted average method, which assigns weights to the red, green, and blue channels based on human eye sensitivity. Histogram equalization is applied to grayscale images. Histogram equalization enhances contrast by redistributing the pixel intensity values ​​of the image. Histogram equalization uses a cumulative distribution function to map pixel values. Gaussian filtering is applied to the grayscale image. The Gaussian filtering uses a two-dimensional Gaussian convolution kernel. The size of the two-dimensional Gaussian convolution kernel is adaptively selected according to the image noise level to smooth the image and preserve edge information. The 3D point cloud data is denoised using a statistical outlier removal algorithm. This algorithm calculates the average distance between each point in the point cloud and its neighbors and removes outliers whose distance exceeds a threshold. The 3D point cloud data is registered using the iterative nearest point algorithm, which aligns the current point cloud with the reference point cloud to correct the pose error caused by the movement of the hardware. The temperature distribution data is normalized. The normalization process linearly maps the original temperature values ​​to the range of zero to one. The normalization process is based on the maximum and minimum values ​​of historical temperature data. Median filtering is applied to the temperature distribution data using a sliding window method. The size of the sliding window is adjusted according to the rate of temperature change to eliminate impulse noise.

4. The online detection method for surface defects of hardware parts based on multimodal sensor fusion according to claim 1, characterized in that, The specific process of multimodal sensor feature extraction in step 3 includes: Texture features are extracted from color image data. The texture features are calculated using the Local Binary Pattern Algorithm (LCA). The LCA divides the image into multiple regions and calculates the LCA histogram for each region to obtain the texture complexity value. Color features are extracted from color image data by converting the image to the HSV color space and calculating histogram statistics of the saturation components, including the mean, variance, and skewness. Geometric features are extracted from 3D point cloud data. These features include surface curvature values ​​and height variance. Surface curvature values ​​are obtained by calculating the rate of change of the normal vector of each point in the point cloud, and height variance is obtained by calculating the standard deviation of the point cloud in the vertical direction. Thermal features are extracted from temperature distribution data. These features include temperature gradient and regional average temperature. The temperature gradient is obtained by calculating the temperature difference between adjacent pixels, and the regional average temperature is obtained by dividing the temperature image into grids and calculating the average value of each grid. All extracted features are normalized to the same scale. The normalization uses the min-max scaling method, which maps the feature values ​​to the range of zero to one. The normalized feature combination is a feature vector of uniform dimension. The dimension of the feature vector is equal to the sum of the number of all features. The feature vector serves as the input for the subsequent fusion step.

5. The online detection method for surface defects of hardware parts based on multimodal sensor fusion according to claim 1, characterized in that, The specific process of multimodal sensor reliability assessment in step 4 includes: For high-resolution optical cameras, the image signal-to-noise ratio is obtained by calculating the variance ratio of pixel values ​​in smooth regions and edge regions in the image. Smooth regions are selected from uniform areas on the surface of the hardware, and edge regions are selected from areas where defects may occur. The contrast value is obtained by calculating the standard deviation of the image gray levels. For line laser scanners, point cloud density is obtained by calculating the number of points in a unit area, which is determined based on the surface size of the hardware. Point cloud integrity is obtained by calculating the proportion of data missing areas, which are identified by a hole detection algorithm in the point cloud. For thermal imaging cameras, temperature stability is obtained by calculating the standard deviation of temperature changes in consecutive frames. The number of consecutive frames is set according to the production line speed. Temperature resolution is obtained by the minimum resolvable temperature difference of the thermal imaging camera, which is determined by the sensor specifications of the thermal imaging camera. The quality indicators are normalized to the range of zero to one. The normalization uses a linear function, which is based on the maximum and minimum values ​​of historical data. The reliability weight is calculated using a fuzzy logic method. The fuzzy logic method includes defining input variables and output variables. The input variables are quality indicators, and the output variables are reliability weights. The fuzzy logic method sets up a fuzzy rule base, which defines the mapping relationship between quality indicators and reliability weights based on expert experience. The defuzzification of the fuzzy logic method adopts the centroid method. The centroid method calculates the weighted average of the output variables to obtain the reliability weight value of each sensor. The reliability weight value ranges from zero to one.

6. The online detection method for surface defects of hardware parts based on multimodal sensor fusion according to claim 1, characterized in that, The specific process of adaptive weighted fusion of multimodal features in step 5 includes: The weighted feature concatenation method first multiplies the feature vector of each sensor by the corresponding reliability weight to obtain a weighted feature vector. The weighted feature vector is calculated by multiplying each element of the feature vector by the reliability weight. The weighted feature vectors are concatenated into a fused feature vector. The concatenation operation is performed according to the sensor order, which is high-resolution optical camera, line laser scanner and thermal imaging camera. The dimension of the fused feature vector is equal to the sum of the dimensions of all weighted feature vectors, and the fused feature vector serves as the input to the defect decision-making step. During the splicing process, if the feature vector dimensions are inconsistent, the zero-padding method is used to align the dimensions. The zero-padding method adds zero values ​​to the end of the shorter feature vector. Adaptive weighted fusion also includes a fusion validity check, which is performed by calculating the variance of the fused feature vector. If the variance is lower than a threshold, the reliability assessment step is re-executed. The threshold is set based on historical fusion data. The fused feature vectors are stored in a cache, the size of which is set according to the production line processing speed to ensure real-time processing.

7. The online detection method for surface defects of hardware parts based on multimodal sensor fusion according to claim 1, characterized in that, The specific process of defect decision-making and output in step 6 includes: The support vector machine classifier is trained offline using historical data, which includes labeled surface defect data of hardware parts. The training process uses a kernel function, which is a radial basis function. During the online phase, the support vector machine classifier classifies the fused feature vectors and outputs the defect type, which includes scratches, dents, rust, and cracks. For samples classified as defects, defect localization is performed. Defect localization is achieved by back-mapping the fused feature vector to the original sensor data to locate the specific location of the defect on the surface of the hardware. The back-mapping adopts a coordinate transformation method based on the sensor calibration parameters. The defect location information is output to the production line control system, which triggers the sorting mechanism to remove the defective part. The sorting mechanism includes pneumatic push rods or robotic arms. The test results are recorded in real time to the database, which stores the defect type, location and timestamp for quality traceability and process optimization. The support vector machine classifier is updated periodically, with the update cycle set according to the production line output. During the update process, the classifier is retrained using newly collected labeled data.

8. The online detection method for surface defects of hardware parts based on multimodal sensor fusion according to claim 1, characterized in that, The specific implementation methods of the synchronous trigger in step 1 include: The synchronous trigger receives pulse signals from the encoder of the production line. The pulse signals are synchronized with the movement distance of the conveyor belt, and the frequency of the pulse signals is adaptively adjusted according to the speed of the conveyor belt. The synchronization trigger generates a synchronization signal, which is simultaneously sent to the high-resolution optical camera, the line laser scanner, and the thermal imaging camera. The synchronization signal is triggered by the rising edge. Upon receiving the synchronization signal, the high-resolution optical camera, line laser scanner, and thermal imaging camera immediately begin data acquisition, and the data acquisition timestamp is recorded in the metadata. The synchronous trigger also monitors the timing consistency of data acquisition. If any sensor fails to respond within a predetermined time, the synchronous trigger issues a warning signal, which notifies the operator to check the sensor status. The hardware circuit of the synchronous trigger is implemented using an FPGA, which is programmed and configured to use a multi-output synchronous mode to ensure that signal delay is minimized.

9. The online detection method for surface defects of hardware parts based on multimodal sensor fusion according to claim 1, characterized in that, Step 2 of the preprocessing also includes data quality assessment and correction processes: Data quality assessment calculates a quality score for each sensor's data. The quality score is based on signal-to-noise ratio, integrity, and consistency metrics. The signal-to-noise ratio is calculated by the power ratio of the signal to the noise, the integrity is calculated by the data missing rate, and the consistency is calculated by the spatiotemporal alignment error between data from multiple sensors. If the quality score is below the threshold, data correction is performed. Data correction uses interpolation or reconstruction methods. For color image data, bilinear interpolation is used to fill in missing areas; for 3D point cloud data, surface fitting algorithm is used to recover missing points; for temperature distribution data, nearest neighbor interpolation is used to compensate for outliers. Data quality assessment and correction are performed in real time during the preprocessing step, and the corrected data replaces the original data for subsequent feature extraction. The quality score is recorded in a log file for subsequent analysis and system optimization.

10. The online detection method for surface defects of hardware parts based on multimodal sensor fusion according to claim 1, characterized in that, Step 6, the defect decision-making process, also includes a multi-level verification process: The multi-level verification process first uses a support vector machine classifier for preliminary classification, and then uses a convolutional neural network for fine classification. The structure of the convolutional neural network includes convolutional layers, pooling layers and fully connected layers. The convolutional neural network is trained offline. For samples whose classification results from the support vector machine classifier and the convolutional neural network are inconsistent, manual review is performed. The manual review displays the defect image and location through the operator interface, and the operator inputs the final label. The multi-level verification process also includes defect severity assessment, which is based on defect size and location. Defect size is calculated by the number of pixels or point cloud area, and defect location is assessed by the distance from the edge of the hardware. The severity assessment results, along with the defect type, are output to the production line control system for priority sorting. The data flow of the multi-level verification process is executed in parallel with the main process to ensure real-time performance. The parallel execution adopts multi-threading technology.

Citation Information

Cited By

  • A defect detection and process localization method

    CN122196938A

  • 一种缺陷检测与工艺定位方法

    CN122196938B

  • Automotive part defect detection method, apparatus, and medium

    CN122312623A