A defect detection and process localization method
By performing temporal and spatial alignment of multimodal data for automotive parts, and combining vibration detection and dynamic weight fusion, the accuracy and adaptability issues of multimodal data in online inspection are solved. This enables the correlation between defect detection results and process steps, improving inspection accuracy and production process traceability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG ZHIMOU ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-05-14
- Publication Date
- 2026-06-12
AI Technical Summary
In existing technologies for online inspection of automotive parts, it is difficult to reliably align multimodal data in time and space. Environmental factors have a significant impact, and the correlation between defect detection results and process steps is difficult to determine, resulting in insufficient detection accuracy and adaptability.
By synchronously collecting multimodal data of automotive parts in online conveying, a unified time reference is generated, and time and space alignment is performed. The coordinate transformation parameters are corrected based on vibration detection information, and dynamic weights are determined by combining defect type, modal reliability, and characteristic response degree. This enables the weighted combination of multimodal data and the correlation between defect detection results and process steps.
It improves the spatial overlap accuracy of multimodal data and the adaptability of fusion results, enabling accurate detection of defect types, locations, and levels, and timely identification of abnormal processes, thus providing support for process optimization.
Smart Images

Figure CN122196938A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of intelligent manufacturing, machine vision and multi-sensor fusion detection technology, specifically involving a defect detection and process localization method. Background Technology
[0002] Automotive parts, especially body panels, structural components, and stamped parts, are prone to various defects during production, such as scratches, dents, bumps, wrinkles, cracks, excessive springback, misaligned holes, indentations, and abnormal material thickness. These defects not only affect the appearance quality and assembly precision of the parts but may also further impact the consistency, reliability, and safety of the entire vehicle. Therefore, timely and accurate defect detection is of paramount importance in the automotive parts manufacturing process.
[0003] Existing technologies for defect detection in automotive parts mainly include manual visual inspection, offline sampling inspection, and automated inspection based on single visual information. Manual visual inspection relies on the experience and judgment of inspectors, and is greatly affected by subjective factors. It is prone to missed or false detections of defects such as minor scratches, micro-cracks, and localized thermal anomalies, and is difficult to adapt to high-speed online production cycles. While offline sampling inspection can improve inspection accuracy to some extent, it typically cannot achieve real-time inspection of all workpieces, resulting in a lag in inspection results and difficulty in timely feedback to the production process. Automated inspection based on single visual information improves inspection efficiency, but because it only utilizes single modal data, its stability and adaptability remain limited when facing complex scenarios where multiple defects coexist, such as surface texture variations, morphological anomalies, and thermal anomalies.
[0004] To improve detection accuracy, existing technologies employ multi-sensor or multi-modal data for defect detection. For example, visible light images are used to acquire surface texture information, three-dimensional contour data is used to acquire surface undulation information, and temperature distribution data is used to acquire thermal anomaly information, based on which defects are identified. However, existing multi-modal detection schemes typically still suffer from the following problems: Different modal acquisition devices differ in sampling frequency, triggering method, data format, and spatial coordinate system, making it difficult to achieve a reliable one-to-one correspondence between multimodal data. This problem is even more pronounced when the inspected automotive parts are in an online transport state, easily causing inconsistencies in the corresponding locations of the same defect in different modalities, thus affecting the accuracy of subsequent fusion results.
[0005] Existing solutions typically focus on calibration or registration under static conditions. However, in actual production line environments, factors such as rack vibration, support offset, and environmental fluctuations can cause changes in the relative positions of sensors, resulting in a drift in the pre-established spatial mapping relationship and thus reducing the spatial overlap accuracy of multimodal data.
[0006] Existing multimodal fusion methods mostly adopt fixed-weight fusion, simple splicing fusion, or unified global fusion methods. They do not fully consider the different dependence of different defect types on different modal information, nor do they make targeted adjustments based on the current acquisition quality of the modality and the characteristic response of the defect area. Therefore, in complex industrial environments, the adaptability and stability of the fusion results still need to be improved.
[0007] Existing detection solutions typically limit their output to whether a defect exists, its type, or its location. They struggle to link defect detection results with process parameters and production processes, making it difficult to effectively trace the source of defects and providing direct support for subsequent process optimization and anomaly handling.
[0008] Therefore, how to achieve reliable spatiotemporal alignment of multimodal data in the online transportation scenario of automotive parts, reduce the impact of environmental factors such as vibration on the detection results, achieve more reasonable multimodal fusion according to different defect types and modal states, and further determine the correlation between defect detection results and process links has become a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0009] This application provides a defect detection and process location method to solve one of the above-mentioned technical problems.
[0010] The technical solution adopted in this application is as follows: This application provides a defect detection and process location method, including: First multimodal data is synchronously collected for automotive parts in online transport, and time stamps under a unified time reference are generated for each modal data in the first multimodal data. Based on the time stamps, the time of each modal data in the first multimodal data is time-aligned to obtain second multimodal data. The coordinate system of the target modal data in the second multimodal data is used as the preset reference coordinate system, and other modal data in the second multimodal data are mapped to the preset reference coordinate system to obtain the third multimodal data. The third multimodal data is used to characterize the spatial correspondence of the same detection area under different modes, and the spatial correspondence satisfies the preset mapping accuracy requirements determined based on vibration detection information used to characterize the vibration state of the production line frame and / or sensor mounting bracket. Defect detection features are extracted from each modality of the third multimodal data, and dynamic weights corresponding to the defect detection features of each modality are determined based on at least two of the first, second, and third information. Based on the dynamic weights, the defect detection features corresponding to the same detection area in each modal data are weighted and combined to obtain fused features; Based on the fusion features, the defect detection results of the automotive parts are determined, and based on the correlation between the defect detection results and process parameters, the process step information corresponding to the defect detection results is determined.
[0011] According to one embodiment of this application, the first multimodal data includes visible light image data, three-dimensional contour data, and temperature distribution data; The visible light image data is used to characterize the texture information and / or grayscale variation information of the automotive parts surface, the three-dimensional contour data is used to characterize the height variation information and / or topographic undulation information of the automotive parts surface, and the temperature distribution data is used to characterize the local thermal anomaly information of the automotive parts surface.
[0012] According to one embodiment of this application, the target modal data is one modal data in the second multimodal data, and the other modal data are the remaining modal data in the second multimodal data other than the target modal data; The target modal data is used to provide the preset reference coordinate system, and the other modal data is used to map to the preset reference coordinate system to form the third multimodal data.
[0013] According to one embodiment of this application, the first information is information used to characterize the degree of dependence of the type of defect to be detected on different modal features; The first information is obtained through at least one of the following methods: Candidate defect types are determined based on the pre-classification results corresponding to at least one modality of data; The degree of dependence of different modes on the candidate defect types is determined based on the pre-defined correspondence between defect types and modal characterization capabilities. The degree of dependence of different modalities on the candidate defect type is determined based on the contribution of different modalities output by the training model to the current defect determination.
[0014] The second piece of information is used to characterize the reliability of the current data for each modality; The second information is obtained based on the operating status and data quality indicators of the acquisition devices corresponding to each mode; The operating status includes at least one of the following: device online status, frame loss status, abnormal alarm status, and operating temperature status; the data quality indicators include at least one of the following: image clarity, point cloud density, thermal image signal-to-noise ratio, and effective area coverage. The third information is information used to characterize the degree of response of the defect detection features of each modality to the background; The third information is obtained based on the response intensity of the defect detection features of each modality in the defect region; The response intensity is determined based on at least one of the following: response map, gradient distribution, activation intensity, and degree of local difference.
[0015] According to one embodiment of this application, the step of time-aligning the modal data in the first multimodal data based on the time stamp to obtain the second multimodal data includes: A unified master clock is used to synchronize the time of the acquisition devices corresponding to each mode; The acquisition devices corresponding to each mode are activated based on the trigger signal to acquire data. Based on the time stamps corresponding to each modal data, time matching is performed on different modal data to determine the multimodal data corresponding to the same detection time; The modal data with the highest frame rate is used as the baseline time series. For the other modal data, the corresponding baseline time nodes are matched according to the nearest time marker. When no matching data exists for the target time point, interpolation is performed on the corresponding modal data to obtain second multimodal data that corresponds one-to-one in time.
[0016] According to one embodiment of this application, mapping other modal data in the second multimodal data to the preset reference coordinate system to obtain the third multimodal data includes: Based on the external parameter transformation relationship between each modal corresponding acquisition device and the preset reference coordinate system, the other modal data are transformed into the preset reference coordinate system; The three-dimensional contour data is represented as a point cloud, contour line, or height matrix in the coordinate system of the corresponding acquisition device, and mapped to the preset reference coordinate system according to the external parameter transformation relationship. The temperature distribution data is represented as a temperature matrix in the coordinate system of the corresponding acquisition device, and projected onto the preset reference coordinate system according to the external parameter transformation relationship; Different modal data are made to correspond to the same detection area under the preset reference coordinate system.
[0017] According to one embodiment of this application, the vibration detection information includes vibration displacement information and / or vibration acceleration information of the production line frame and / or sensor mounting bracket; The step of determining the preset mapping accuracy requirement based on vibration detection information used to characterize the vibration state of production line racks and / or sensor mounting brackets, and ensuring that the spatial correspondence meets the preset mapping accuracy requirement, includes: Based on the vibration detection information, the rotation and / or translation parameters in the coordinate transformation parameters used to map the other modal data to the preset reference coordinate system are corrected; When the vibration displacement reaches a preset threshold, dynamic recalibration is triggered to update the coordinate transformation parameters; Based on the updated coordinate transformation parameters, the spatial correspondence between each modality data in the second multimodal data and the same detection area is re-established.
[0018] According to one embodiment of this application, the step of determining dynamic weights corresponding to the defect detection features of each modality data based on at least two of the following: first information characterizing the dependence of the defect type to be detected on different modal features; second information characterizing the reliability of the current data of each modality; and third information characterizing the degree of response of the defect detection features of each modality to the background; and then weighting and combining the defect detection features corresponding to the same detection region in each modality data based on the dynamic weights to obtain fused features, including: Determine the defect type dependency coefficient corresponding to each mode based on the first information; Determine the modality confidence level corresponding to each modality based on the second information; The feature saliency corresponding to each mode is determined based on the third information; Based on at least two of the defect type dependency coefficient, modality confidence, and feature saliency, determine the dynamic weights corresponding to the defect detection features of each modality data; The defect detection features corresponding to the same detection area in each modal data are converted into a unified feature expression; The unified feature representation is combined with the corresponding dynamic weights to obtain the weighted features for each modality; The weighted features corresponding to each modality are weighted and summed and / or concatenated to obtain the fused features.
[0019] According to one embodiment of this application, the step of determining the defect detection result of the automotive component based on the fusion features, and determining the process step information corresponding to the defect detection result based on the correlation between the defect detection result and process parameters, includes: Based on the fusion features, at least one of the defect type, defect location, and defect level is determined as the defect detection result; Obtain the timing data of process parameters corresponding to the current automotive parts. The timing data of process parameters includes at least one of stamping pressure, mold temperature, conveying speed, process switching time point, and equipment operating status. Based on the correspondence between the defect detection results and the process parameter time series data, the target process step corresponding to the defect detection results is determined; The correspondence includes at least one of the following: the correspondence between defect location and process operation location, the correspondence between defect type and process anomaly type, and the correspondence between defect detection time and abnormal fluctuation period of process parameters.
[0020] A second aspect of this application provides a defect detection and process location system, comprising: The acquisition module is used to synchronously acquire first multimodal data of automotive parts in online transportation, and generate time stamps under a unified time reference for each modal data in the first multimodal data; A time alignment module is used to perform time alignment on each modal data in the first multimodal data based on the time marker to obtain the second multimodal data; The spatial alignment module is used to take the coordinate system of the target modal data in the second multimodal data as a preset reference coordinate system, and map other modal data in the second multimodal data to the preset reference coordinate system to obtain the third multimodal data. The third multimodal data is used to characterize the spatial correspondence of the same detection area under different modes, and the spatial correspondence satisfies the preset mapping accuracy requirements determined based on vibration detection information used to characterize the vibration state of the production line frame and / or sensor mounting bracket. The feature extraction module is used to extract defect detection features from each modality of the third multimodal data; The weight determination module is used to determine the dynamic weights corresponding to the defect detection features of each modality data based on at least two of the following: first information used to characterize the dependence of the defect type to be detected on different modal features, second information used to characterize the reliability of the current data of each modality, and third information used to characterize the degree of defect detection features of each modality relative to the background response. The fusion module is used to perform weighted combination of defect detection features corresponding to the same detection area in each modal data based on the dynamic weights, so as to obtain fused features; The determination and positioning module is used to determine the defect detection result of the automotive component based on the fusion features, and to determine the process step information corresponding to the defect detection result according to the correlation between the defect detection result and the process parameters.
[0021] Due to the adoption of the above technical solution, the beneficial effects achieved by this application are as follows: This application achieves time alignment by synchronously acquiring multimodal data of automotive parts during online transportation and using time stamps under a unified time reference. This reduces data misalignment caused by different sampling frequencies and trigger times of different acquisition devices, enabling different modal data to more reliably correspond to the same part area at the same detection time, thus providing a more accurate data foundation for subsequent defect detection.
[0022] This application maps non-reference modal data to a preset reference coordinate system and corrects the coordinate transformation parameters by combining vibration detection information. This reduces the impact of factors such as production line vibration and support offset on spatial mapping relationships, thereby improving the spatial overlap accuracy between different modal data and enhancing the correspondence and consistency between multimodal features.
[0023] This application does not use fixed weights or simple splicing for fusion. Instead, it determines the dynamic weights of each modal feature based on at least two of the following: defect type information, modal confidence information, and feature saliency information. This allows for more targeted adjustments to the fusion process for different defect scenarios, modal states, and feature responses, thereby improving the adaptability of the fusion results to complex industrial environments and various types of defects.
[0024] This application can further divide the area to be detected into multiple local feature regions, and perform weight enhancement on key pixels or key feature points corresponding to key defect locations, so that the fusion process is more focused on defect-related areas, reducing the interference of background information and irrelevant features on the detection results, thereby improving the stability of defect detection results.
[0025] This application can not only output detection results such as defect type, defect location and defect level, but also determine the corresponding process step information by combining process parameter time series data, thereby realizing the extension from defect detection to defect formation source location, which is conducive to timely identification and traceability of abnormal links in the production process, and provides support for process optimization, quality control and production decision-making.
[0026] This application is applicable to online inspection scenarios for automotive parts. While ensuring the accuracy of multimodal inspection, it also takes into account actual industrial conditions such as online transportation, environmental vibration, and the coexistence of multiple types of defects, and has good engineering application value and promotion value. Attached Figure Description
[0027] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1A flowchart illustrating a defect detection and process location method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0028] Figure label: 810, Processor; 820, Communication interface; 830, Memory; 840, Communication bus. Detailed Implementation
[0029] To more clearly illustrate the overall concept of this application, a detailed explanation is provided below with reference to the accompanying drawings.
[0030] Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of this application is not limited to the specific embodiments disclosed below. It should be noted that, unless otherwise specified, the embodiments of this application and the features thereof can be combined with each other.
[0031] In this application, unless otherwise expressly specified and limited, the "above" or "below" of the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. In the description of this specification, references to terms such as "an embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples.
[0032] Example 1 like Figure 1 As shown, a defect detection and process location method includes: S1. Synchronously collect first multimodal data of automotive parts in online transportation, and generate time stamps under a unified time reference for each modal data in the first multimodal data. Based on the time stamps, time-align each modal data in the first multimodal data to obtain second multimodal data.
[0033] Specifically, the objects to be inspected are automotive parts that move continuously along a conveying path. These automotive parts can be automotive body panels, stamped parts, structural parts, or other parts requiring surface and morphology quality inspection. Since the objects to be inspected are in an online conveying state, their position changes continuously over time. Therefore, during the data acquisition phase, it is necessary not only to acquire different types of data that reflect the defect state, but also to ensure that these different types of data correspond to each other in the time dimension. To this end, this embodiment performs synchronous acquisition of first multimodal data on the automotive parts. The first multimodal data is the raw multimodal data directly acquired during the acquisition phase, and may include at least two of the following: visible light image data, three-dimensional contour data, and temperature distribution data. Visible light image data is used to characterize the texture information and / or grayscale variation information of the automotive part surface; three-dimensional contour data is used to characterize the height variation information and / or morphological undulation information of the automotive part surface; and temperature distribution data is used to characterize the local thermal anomaly information of the automotive part surface. Different modal data characterize the state of the automotive parts from different physical perspectives, thereby providing a complementary data foundation for subsequent defect detection.
[0034] The synchronous acquisition described here does not require that each modal acquisition device completes sampling at an absolutely simultaneous physical moment. Rather, it means that the modal data is acquired around the same workpiece, the same detection process, and the same time reference, and can be mapped to the same detection moment in subsequent processing. To achieve this, a time stamp under a unified time reference is generated for each modal data in the first multimodal data during the acquisition process. The unified time reference can be provided by a unified master clock, which provides a unified timing signal to the acquisition devices corresponding to each modality, making the time stamps generated by different modal acquisition devices comparable. The time stamp can be attached to each frame of image data, each frame of contour data, or each frame of temperature matrix to characterize the acquisition time of the corresponding data.
[0035] When automotive parts enter the inspection station, a trigger source outputs a trigger signal to activate the corresponding acquisition devices for each modality to collect data. The trigger source can be a photoelectric sensor installed on the conveyor path, or an encoder, proximity switch, or other detection element capable of characterizing the workpiece's arrival status. The trigger signal can be sent directly to each modality acquisition device using a hard trigger method, or it can be sent first to an intermediate control module, which then distributes it to each modality acquisition device. Through these methods, different modal data can be collected around the same arrival time of the same workpiece, thereby ensuring the consistency of the first multimodal data in terms of time reference.
[0036] After completing the first multimodal data acquisition, the various modal data in the first multimodal data are time-aligned based on the timestamps to obtain the second multimodal data. The purpose of the time alignment is to ensure that the different modal data participating in subsequent defect detection correspond to the same detection time, rather than selecting data from their respective independent times and combining them. Since the sampling frequency, frame rate, and output rhythm of different modal acquisition devices may differ, it is necessary to establish a time correspondence between different modalities based on the timestamps carried by each modal data. Specifically, within a preset time tolerance range, the data with the closest times in different modalities can be grouped into data groups corresponding to the same detection time; or, a time node of one modality can be used as a reference, and the data frame with the smallest time difference can be found from other modalities as the matching result. Through the above time matching, the originally acquired first multimodal data can be converted into second multimodal data that corresponds one-to-one in the time dimension.
[0037] The modality with the highest frame rate can be selected as the baseline time series. For the remaining modal data, the corresponding baseline time nodes are matched according to the nearest time marker. When there is no directly matching data for a certain target time node, interpolation can be performed on the corresponding modal data to obtain second multimodal data that corresponds one-to-one in time. In this way, the impact of inconsistent sampling frequencies of different modalities on subsequent defect detection can be reduced, the temporal consistency between different modal data can be improved, and a reliable data foundation can be provided for subsequent spatial alignment, feature extraction, and multimodal fusion.
[0038] S2. The coordinate system of the target modal data in the second multimodal data is used as a preset reference coordinate system, and other modal data in the second multimodal data are mapped to the preset reference coordinate system to obtain third multimodal data. The third multimodal data is used to characterize the spatial correspondence of the same detection area under different modes, and the spatial correspondence satisfies the preset mapping accuracy requirements determined based on vibration detection information used to characterize the vibration state of the production line frame and / or sensor mounting bracket.
[0039] Specifically, the second multimodal data is time-aligned multimodal data. Although the time-aligned modal data can correspond to the same detection time, due to differences in the installation position, imaging method, and coordinate representation of the acquisition devices corresponding to different modalities, there are still situations where the different modal data cannot directly correspond to the same area of the automotive parts in the spatial dimension. Therefore, in this step, the second multimodal data is spatially aligned to obtain the third multimodal data, so that the different modal data can correspond not only in the time dimension but also in the spatial dimension to the same detection area.
[0040] The target modal data is one modal data point in the second multimodal data, and its coordinate system is selected as a preset reference coordinate system. The other modal data are the remaining modal data in the second multimodal data excluding the target modal data. The target modal data provides a unified spatial reference, and the other modal data are mapped to this unified spatial reference. In this way, multimodal data that were originally located in different coordinate systems can be unified into a single coordinate system for expression, thereby establishing a one-to-one correspondence between the same detection area and different modalities.
[0041] The target modal data is preferably visible light image data, and correspondingly, the coordinate system of the visible light image data is used as the preset reference coordinate system. The reason for preferably using the coordinate system corresponding to the visible light image data is that visible light images typically have high resolution and facilitate subsequent display, annotation, and manual verification of defect locations. Of course, in other embodiments, the coordinate system corresponding to the three-dimensional contour data or other unified reference coordinate system can be selected as the preset reference coordinate system based on the actual system structure, detection requirements, or modal characteristics.
[0042] After determining the preset reference coordinate system, the other modal data are mapped to the preset reference coordinate system. Here, "mapping" refers to transforming each non-reference modal data from its original coordinate system to the preset reference coordinate system based on the extrinsic transformation relationship of each modal acquisition device relative to the preset reference coordinate system. The extrinsic transformation relationship can be obtained through a calibration process and typically includes rotation and translation relationships. For 3D contour data, the acquisition results can first be represented as point clouds, contour lines, or height matrices in the local coordinate system based on the imaging model of the 3D contour acquisition device itself, and then transformed to the preset reference coordinate system according to the extrinsic transformation relationship. For temperature distribution data, the pixels in the temperature matrix can first be associated with their local imaging coordinates, and then the temperature information can be projected onto the corresponding position in the preset reference coordinate system through the extrinsic transformation relationship and pixel mapping relationship.
[0043] In practical implementation, the above mapping result can be expressed in at least one of the following forms: projecting a three-dimensional point cloud onto a two-dimensional image plane to obtain contour or depth information consistent with the target modal data position; resampling temperature pixels in the temperature distribution data into the pixel grid corresponding to the target modal data to obtain a temperature distribution map consistent with the target modal data position; or, uniformly expressing different modal data as a feature matrix under the same grid size. Through the above mapping, data in different modalities that were originally located in different coordinate systems can describe the same detection area on automotive parts under the same spatial reference, thereby obtaining the third multimodal data. The third multimodal data is not simply the result of splicing raw data, but a set of multimodal data with established spatial correspondence.
[0044] The phrase "the third multimodal data is used to characterize the spatial correspondence of the same detection area in different modalities" can be understood as follows: in the third multimodal data, data units from different modalities can all point to the same physical area on the automotive component. For example, the texture anomaly location of the same defect area in the visible light image, the height anomaly location in the three-dimensional contour data, and the thermal anomaly location in the temperature distribution data can be established as correspondences in a unified coordinate system after spatial alignment. Thus, during subsequent defect detection feature extraction and multimodal fusion, the feature information in each modality can be jointly analyzed around the same detection area, thereby reducing the impact of spatial misalignment between different modalities on the detection results.
[0045] Considering that this application is applied in an online production line environment, the production line frame, sensor mounting bracket, and conveying mechanism may experience mechanical vibrations during operation, causing changes in the relative positions of the acquisition equipment. If the coordinate transformation relationship obtained from the initial calibration is still used, the spatial mapping relationship may drift, reducing the spatial correspondence accuracy between different modal data. Therefore, this application further introduces vibration detection information to characterize the vibration state of the production line frame and / or sensor mounting bracket during the spatial alignment process, ensuring that the spatial correspondence meets the preset mapping accuracy requirements determined based on the vibration detection information. In other words, this step not only requires completing the coordinate transformation between modalities but also requires that the established spatial correspondence still meets the spatial mapping accuracy required for subsequent defect detection even in the presence of vibration disturbances.
[0046] Specifically, the vibration detection information can be acquired by a vibration detection unit installed on the production line frame and / or sensor mounting bracket. The vibration detection information may include vibration displacement information and / or vibration acceleration information. After acquiring the vibration detection information, the system can correct the coordinate transformation parameters used in the mapping process in real time based on a pre-established error model. Let the coordinate transformation parameters from the initial calibration non-reference mode to the preset reference coordinate system be T0=[R0, t0], where R0 is the initial rotation matrix and t0 is the initial translation vector. Let the displacement change corresponding to the vibration detection information acquired at the current moment be Δpk=[Δxk, Δyk, Δzk]^T, and the attitude change be Δθk=[Δαk, Δβk, Δγk]^T. Then, the initial coordinate transformation parameters can be corrected based on the displacement change and attitude change to obtain the coordinate transformation parameters Tk=[Rk, tk] at the current moment, where tk=t0+Kp·Δpk, Rk=ΔRk·R0, and ΔRk is determined by the incremental rotation matrix constructed from the attitude change. For any spatial point Pm in the non-reference mode, it can be mapped to a preset reference coordinate system according to Pb=Rk·Pm+tk. If the preset reference coordinate system is an image coordinate system, Pb can be further projected to the pixel coordinate position corresponding to the target mode through a camera projection model.
[0047] When the vibration amplitude is large, real-time correction alone may not be sufficient to maintain the preset mapping accuracy requirements. Therefore, dynamic recalibration triggering conditions can be set. Specifically, dynamic recalibration is triggered when the displacement change satisfies ||Δpk||2>Dth, and / or the attitude change satisfies ||Δθk||2>θth, and / or the mapping error of the reference feature point satisfies Ek>Eth; where Ek=(1 / N)·Σ||qi-q'i||^2, qi is the actual position of the reference feature point in the target mode, and q'i is the predicted position obtained based on the current coordinate transformation parameters. During dynamic recalibration, the updated extrinsic parameter matrix Tnew can be resolved by re-acquiring calibration component data, identifying fixed reference feature points, or using the standard positioning structure preset by the production line, employing at least one of the least squares method, PnP algorithm, and iterative nearest point algorithm, so that the updated spatial correspondence meets the preset mapping accuracy requirements. The above method can balance real-time performance and mapping accuracy under online operating conditions, maintaining spatial consistency of different modal data for the same detection area.
[0048] By performing this step, the second multimodal data can be converted into the third multimodal data, so that different modal data can represent the same detection area on automotive parts under a unified spatial reference, and still maintain the spatial correspondence that meets the preset mapping accuracy requirements under vibration environment, thereby providing a reliable spatial data foundation for subsequent defect detection feature extraction, dynamic weight determination and fusion feature generation.
[0049] S3. Extract defect detection features from each modal data in the third multimodal data, and determine the dynamic weights corresponding to the defect detection features of each modal data based on at least two of the first information, the second information, and the third information.
[0050] Specifically, the third multimodal data is multimodal data that has undergone time and spatial alignment. Since each modality in the third multimodal data corresponds to the same detection time in the time dimension and the same detection area on the automotive component in the spatial dimension, defect detection features can be extracted from each modality under a unified spatiotemporal reference, and further joint analysis can be performed on the defect detection features of different modalities. The defect detection features mentioned here are not the original image data, original contour data, or the original temperature matrix itself, but rather feature information extracted from the original data that can characterize the presence, type, location, and severity of defects.
[0051] Different types of defect detection features can be extracted for different modal data. For visible light image data, features such as texture variation, grayscale discontinuity, edge anomaly, local contrast, scratch response, and surface spot can be extracted to characterize visual anomalies on the surface of automotive parts. For three-dimensional contour data, features such as height variation, surface curvature variation, local morphological abrupt changes, contour deviation, and spatial geometric anomalies can be extracted to characterize morphological anomalies such as pits, protrusions, wrinkles, springback, or hole deviations on the surface of automotive parts. For temperature distribution data, features such as local temperature anomalies, temperature gradients, thermal distribution discontinuities, and thermal anomaly region range can be extracted to characterize defects related to thermal effects or process thermal anomalies. These features can be extracted using preset feature calculation rules, trained feature extraction networks, or a combination of rule extraction and network extraction. Regardless of the method used, the core of this step is to extract features that are significant for defect discrimination for different modalities and retain the correspondence between these features and the same detection area in the third multimodal data.
[0052] After extracting defect detection features from each modality, instead of directly applying fixed weights to these features, dynamic weights are determined based on at least two types of information within the current detection scenario. These at least two types of information include: first information characterizing the dependence of the detected defect type on different modal features; second information characterizing the reliability of the current data for each modality; and third information characterizing the degree of response of each modality's defect detection features relative to the background. By incorporating this information, different modalities can achieve different participation ratios under the current detection conditions, thus avoiding the problem that fixed weights cannot adapt to different defect types, modal states, and feature response conditions.
[0053] The first information is used to characterize the degree of dependence of the detected defect type on different modal features. In other words, the first information reflects which modality or modalities are more effective at characterizing the type of defect in the current defect scenario. For example, for surface scratches, visible light image data is usually better at characterizing texture changes and grayscale anomalies, so the corresponding visible light modality has a high degree of dependence on this type of defect; for morphological defects such as wrinkles, bumps, or springback, three-dimensional contour data is usually better at reflecting height changes and surface undulations, so the corresponding three-dimensional contour modality has a high degree of dependence on this type of defect; for defects related to thermal anomalies or process heat effects, temperature distribution data can more directly characterize the distribution of thermal anomalies, so the corresponding temperature modality has a high degree of dependence on this type of defect. The system can determine the degree of dependence of different modalities on the current defect based on the candidate category of the defect to be judged, pre-classification results, prior rules, or training model output, and form the first information accordingly.
[0054] The second information is used to characterize the reliability of the current data for each modality. Specifically, the second information reflects whether a certain modality is suitable for participating in subsequent defect determination at the current detection moment. For example, even if a certain modality theoretically has a strong characterization ability for a certain type of defect, if its corresponding acquisition device is offline, experiencing continuous frame loss, abnormal alarms, or abnormal operating temperature at the current moment, or if its current acquired data has quality problems such as insufficient clarity, low point cloud density, poor thermal image signal-to-noise ratio, or insufficient effective area coverage, then the reliability of the current data for that modality will decrease, and correspondingly, the participation of that modality in the current detection should also be reduced. Therefore, the system can determine the second information based on the operating status and data quality indicators of the acquisition devices corresponding to each modality. The higher the second information, the more reliable the current data of the corresponding modality is, and the more suitable it is for participating in subsequent fusion judgment.
[0055] The third information is used to characterize the response degree of each modality's defect detection features relative to the background. This third information reflects whether a particular modality produces a sufficiently significant feature response within the current defect region. If the response map, gradient distribution, activation intensity, or local difference of a particular modality in the defect region is significantly higher than that in the background region, it indicates that the modality has strong discriminative ability in the current defect determination, and its corresponding feature significance is high. Conversely, if the response of a particular modality in the defect region is not significantly different from the background, it indicates that the modality's current discriminative ability for the defect is weak, and its corresponding feature significance is low. Therefore, the system can form the third information based on the response intensity of each modality's defect detection features in the defect region.
[0056] After obtaining at least two of the first, second, and third information, the system determines the dynamic weights corresponding to the defect detection features of each modality based on the at least two pieces of information. The system can determine the defect type dependency coefficient for each modality based on the degree of dependence of the defect type on different modalities; determine the modality confidence level for each modality based on the operating status and data quality indicators of each modality acquisition device; determine the feature saliency for each modality based on the response intensity of the defect region in the feature extraction results; and then determine the dynamic weights corresponding to the defect detection features of each modality based on at least two of the defect type dependency coefficient, the modality confidence level, and the feature saliency. In other words, the dynamic weights are not fixed constants but participation coefficients that dynamically change according to the current detection object, the current modality state, and the current feature response.
[0057] The dynamic weights can be determined through rule-based calculation, parameterized function calculation, or model output calculation. For example, at least two of the defect type dependency coefficient, modality confidence, and feature saliency can be combined using weighted summation, product normalization, piecewise function mapping, or other deterministic calculation methods to obtain the dynamic weights of each modality in the current detection scenario. The final dynamic weights are used to characterize the degree of participation of each modality in the current defect judgment. In this way, modalities that are more sensitive to the current defect, have more reliable current data, and have more significant current responses can be given higher weights, while modalities with weaker representation capabilities, poorer data quality, or less obvious responses can be given lower weights, thereby improving the targeting, stability, and adaptability of subsequent defect judgments.
[0058] S4. Based on the dynamic weights, the defect detection features corresponding to the same detection area in each modal data are weighted and combined to obtain fused features.
[0059] Specifically, after the aforementioned time and spatial alignment processes, the various modal data in the third multimodal data can collectively correspond to the same detection area on the automotive parts. Furthermore, after feature extraction, defect detection features for each modality within the same detection area have been obtained. Based on this, this step does not employ fixed-weight fusion, simple splicing fusion, or uniformly applying the same set of weights to the entire detection area. Instead, it uses the aforementioned determined dynamic weights to weightedly combine the defect detection features corresponding to the same detection area in each modal data to obtain fused features that comprehensively reflect the effective multimodal information in the current defect scenario.
[0060] Weighted combination refers to assigning different participation ratios to the defect detection features of each modality according to the dynamic weights corresponding to each modality. This ensures that modalities with stronger characterization capabilities for the current defect type, higher data reliability, and more obvious responses in the current defect region have a larger proportion in the final fusion result; while modalities with weaker characterization capabilities, poorer data quality, or less obvious responses have a reduced participation. This approach avoids treating different modal features equally during fusion, thereby improving the specificity of the fusion result for the current defect scenario.
[0061] Before weighted combination, the defect detection features corresponding to the same detection region in each modality can be converted into a unified feature representation. A unified feature representation means that the defect detection features of different modalities have a corresponding representation form in the subsequent combination process. For example, visible light image features, three-dimensional contour features, and temperature distribution features can be converted into feature vectors, feature matrices, or intermediate feature tensors of a unified dimension; or, while maintaining the original feature structure of different modalities, the features of each modality can correspond to the same set of local locations within the same detection region. This unified processing provides a foundation for subsequent combination according to dynamic weights.
[0062] After obtaining a unified feature representation, the defect detection features of each modality can be combined with their corresponding dynamic weights to obtain weighted features for each modality. In other words, for each modality, its defect detection features are no longer directly used in subsequent judgments, but are first weighted according to the dynamic weights corresponding to that modality. For modalities with higher weights, their corresponding features retain a greater influence after weighting; for modalities with lower weights, the contribution of their corresponding features to the final fusion result is correspondingly weakened after weighting. This method ensures that the fusion process truly reflects the actual effectiveness of each modality under the current detection conditions.
[0063] The weighted combination can be implemented through at least one of the following methods: weighted summation of the weighted features corresponding to each modality to obtain the fused feature representation; weighted concatenation of the weighted features corresponding to each modality to obtain the fused feature representation; or, inputting the weighted features corresponding to each modality into a subsequent fusion network or intermediate representation layer for joint mapping to output the fused feature. In other words, the fusion can occur at the feature layer or at the intermediate representation layer before judgment, as long as the final fused feature can comprehensively reflect the effective information of different modalities in the current defect scenario, the technical concept of this application can be realized.
[0064] The weighted combination is not applied uniformly across the entire detection area, but rather allows for finer-grained dynamic fusion by combining local regions and key locations. Specifically, the system first divides the detection area into multiple local feature regions and determines the region-level dynamic weight for each local feature region. This allows different modalities to participate in different proportions in different local regions. For example, in one region, 3D contour features more clearly characterize morphological anomalies, so the 3D contour modality has a higher weight in that region; while in another local region, visible light image features more clearly characterize surface texture anomalies, so the visible light modality has a higher weight in that region. This region-level weighted combination method allows different modalities to play a greater role in the local regions where they are better suited to characterize features.
[0065] Beyond region-level weighted combination, the dynamic weights of corresponding modalities can be enhanced and adjusted for key pixels or key feature points at critical defect locations. These key pixels or key feature points can be defect center points, abrupt edge changes, crack endpoints, depth extrema, temperature anomaly peaks, etc. Enhancement adjustments can be made by introducing enhancement coefficients on top of the region-level dynamic weights, or by redistributing the weights of each modality within the neighborhood of the key point. This approach allows the fusion process to not only highlight effective modalities at the region scale but also further strengthen the focus on core defect features at the key point scale, thereby improving the final fused features' ability to represent defects.
[0066] Furthermore, the dynamic weights are not fixed after being determined once, but can be updated when preset conditions are met, and the weighted combination process can be re-executed accordingly. The preset conditions may include at least one of the following: a change in the detected defect type, a change in modality confidence reaching a preset threshold, or a change in feature significance reaching a preset threshold. In other words, when the detected object changes from one defect type to another, or when the effectiveness of a modality changes due to changes in equipment status, data quality, or local response enhancement or weakening, the system can redetermine the dynamic weights corresponding to each modality and recombine the defect detection features of each modality based on the updated dynamic weights to obtain fusion features adapted to the current detection scenario.
[0067] The fused features are not simply a continuation of single-modal features, but rather a joint feature representation that integrates texture, shape, and / or thermal anomaly information of the same detection area under different modalities. These fused features can more comprehensively reflect the defect state of the current detection area, thus providing a more reliable feature foundation for subsequently determining the defect detection results of automotive parts based on these fused features. Compared to relying solely on single-modal features or using fixed weights for fusion, this step can adaptively adjust the participation ratio of different modalities in the fusion process according to the current defect type, current modal state, and current feature response, thereby improving the accuracy, stability, and adaptability to complex industrial environments in defect detection.
[0068] S5. Determine the defect detection result of the automotive component based on the fusion features, and determine the process step information corresponding to the defect detection result according to the correlation between the defect detection result and the process parameters.
[0069] Specifically, the fused feature is a joint feature representation obtained by weighted combination of defect detection features corresponding to the same detection area in different modal data. Since the fused feature comprehensively reflects the texture, shape, and thermal anomaly information of the same detection area under different modalities such as visible light images, three-dimensional contours, and / or temperature distributions, it can comprehensively characterize the defect state of the current detection area. Based on this, the system first determines the defect detection result of the automotive component based on the fused feature.
[0070] The defect detection result is not limited to the presence or absence of a defect, but may include at least one of the following: defect type, defect location, and defect level. The defect type characterizes whether the current defect belongs to a category such as scratch, dent, wrinkle, crack, indentation, hole misalignment, or other defects. The defect location characterizes the spatial position of the defect on the surface, contour surface, or target area of the automotive part. The defect level characterizes whether the defect belongs to a minor defect, a general defect, a severe defect, or other preset level classification. In other words, the defect detection result not only reflects the presence or absence of a defect, but also further reflects what the defect is, where it is located, and its severity.
[0071] The system can output defect types based on the fused features. Specifically, the fused features can be input into a classification network, detection network, segmentation network, or rule-based decision module to obtain the corresponding defect category. For example, when the fused features indicate obvious linear texture abnormalities, significant local gray-scale abrupt changes with small morphological changes, and insignificant thermal anomalies, it can be identified as a scratch-type defect; when the fused features indicate significant local height changes and obvious curvature abnormalities, it can be identified as a pit, bump, or wrinkle-type defect; when the fused features indicate significant local temperature anomalies, it can be identified as a thermal anomaly-related defect. In this way, different defect types can be distinguished based on the fused features.
[0072] The system can also determine the defect location based on the fused features. Here, the defect location can be a specific coordinate position, regional position, edge region, central region, or position corresponding to a preset reference structure on the automotive component. Since the fused features are based on completed temporal and spatial alignment, the features characterizing the same defect region under different modalities have been unified into the same spatiotemporal reference. The system can then locate the defect region and output the corresponding location result. For implementations requiring more precise positioning, the defect location can also be determined using detection boxes, segmented regions, key points, or response peak positions.
[0073] Furthermore, the system can determine the defect level based on the fused features. The defect level can be classified according to defect area, depth, length, temperature difference, contour deviation, or a comprehensive score. For example, when the defect area is small, the response is weak, and it does not affect subsequent assembly or use, it can be determined as a minor defect; when the defect area is large, the feature response is obvious, or it has a significant impact on the quality of the workpiece, it can be determined as a serious defect. Thus, the fused features can further support the system's judgment of defect severity.
[0074] After obtaining the defect detection results, the system does not stop at defect identification itself, but further determines the process step information corresponding to the defect detection results based on the correlation between the defect detection results and process parameters. Here, process parameters refer to parameter data associated with the current automotive part production process; preferably, the process parameters are stored in the form of time-series data, forming time-series process parameter data corresponding to the current automotive part. The time-series process parameter data may include at least one of the following: stamping pressure, die temperature, conveyor speed, pressing position, process switching time point, equipment operating status changes, or other process control parameters. Since these process parameters can reflect the operating status of automotive parts at each process stage during production, they can serve as the basis for locating subsequent process steps.
[0075] Correlation analysis refers to the correlation analysis between the defect detection results and the process parameters to determine which process stage or type of process action the current defect is more likely to have occurred in. In other words, it is not simply about providing process information directly based on defect detection results, but rather about combining defect type, defect location, defect level, and / or detection time with time-series data of process parameters in the production process to analyze the correspondence between different defect manifestations and different process anomalies, thereby identifying the target process step that caused the defect.
[0076] The association can be established through at least one of the following methods: First, matching is performed based on the correspondence between the defect location and the known process step's location. For example, if the defect is located in the stamping contact area or a specific mold action area, the defect can be preferentially associated with the corresponding process step. Second, matching is performed based on the historical correspondence between the defect type and different process anomaly types. For example, regular indentations are more likely to correspond to stamping contact anomalies, local thermal anomalies are more likely to correspond to heat-affected process anomalies, and hole offsets or contour deviations are more likely to correspond to forming or conveying fluctuations. Third, matching is performed based on abnormal fluctuations in process parameters near the detection time. For example, if abnormal fluctuations occur in stamping pressure, mold temperature, conveying speed, or equipment status during the production period corresponding to the defect detection, the process step corresponding to the abnormal fluctuation can be used as a candidate target process step. Fourth, the fused features and the temporal features of the process parameters can also be input together based on the trained association model to output the target process step corresponding to the defect detection result.
[0077] For example, if the system detects regular indentations on the surface of a workpiece, and the mold contact pressure abnormally increases during the corresponding production period, the defect can be associated with the stamping contact process. If the system detects localized thermal anomalies accompanied by abnormal fluctuations in mold or workpiece temperature parameters, the defect can be associated with the corresponding heat-affected zone. If the system detects hole misalignment and contour deviation, and the corresponding conveyor speed fluctuates or forming parameters are abnormal, the defect can be associated with the corresponding forming or transport process. In this way, the system can not only identify the defect itself but also trace its source.
[0078] In some embodiments of this application, the first multimodal data includes visible light image data, three-dimensional contour data, and temperature distribution data; The visible light image data is used to characterize the texture information and / or grayscale variation information of the automotive parts surface, the three-dimensional contour data is used to characterize the height variation information and / or topographic undulation information of the automotive parts surface, and the temperature distribution data is used to characterize the local thermal anomaly information of the automotive parts surface.
[0079] Specifically, visible light image data is mainly used to reflect the texture features, grayscale differences, and appearance anomalies of automotive parts, such as scratches, spots, or edge abnormalities; three-dimensional contour data is mainly used to reflect the height variations and morphological undulations of automotive parts, such as pits, protrusions, wrinkles, or contour deviations; and temperature distribution data is mainly used to reflect local thermal anomalies on the surface of automotive parts, such as abnormal local temperature rises, uneven heat distribution, or abnormal areas related to process thermal effects. These three types of data allow for the characterization of automotive parts from the perspectives of appearance, morphology, and thermal state, providing a complementary data foundation for subsequent defect detection.
[0080] In some embodiments of this application, the target modal data is one modal data in the second multimodal data, and the other modal data are the remaining modal data in the second multimodal data other than the target modal data; The target modal data is used to provide the preset reference coordinate system, and the other modal data is used to map to the preset reference coordinate system to form the third multimodal data.
[0081] Specifically, the target modal data is the modal data selected as a unified spatial reference from the second multimodal data, and its corresponding coordinate system is used as a preset reference coordinate system; the other modal data are the remaining modal data excluding the target modal data. In subsequent processing, the other modal data are uniformly mapped to this preset reference coordinate system, so that data from different modalities can correspond to the same detection area under the same spatial reference, thereby forming the third multimodal data. In other words, the target modal data mainly serves to provide a reference, while the other modal data mainly serves to "align to the reference".
[0082] In some embodiments of this application, the first information is information used to characterize the degree to which the type of defect to be detected depends on different modal features; The first information is obtained through at least one of the following methods: Candidate defect types are determined based on the pre-classification results corresponding to at least one modality of data; The degree of dependence of different modes on the candidate defect types is determined based on the pre-defined correspondence between defect types and modal characterization capabilities. The degree of dependence of different modalities on the candidate defect type is determined based on the contribution of different modalities output by the training model to the current defect determination.
[0083] The second piece of information is used to characterize the reliability of the current data for each modality; The second information is obtained based on the operating status and data quality indicators of the acquisition devices corresponding to each mode; The operating status includes at least one of the following: device online status, frame loss status, abnormal alarm status, and operating temperature status; the data quality indicators include at least one of the following: image clarity, point cloud density, thermal image signal-to-noise ratio, and effective area coverage. The third information is information used to characterize the degree of response of the defect detection features of each modality to the background; The third information is obtained based on the response intensity of the defect detection features of each modality in the defect region; The response intensity is determined based on at least one of the following: response map, gradient distribution, activation intensity, and degree of local difference.
[0084] Specifically, the first information is used to explain which modal features the current defect to be detected depends more on, which can be obtained based on the pre-classification results, the correspondence between the preset defect type and the modal representation ability, or the modal contribution degree output by the training model; The second information is used to indicate whether the current data of each modality is reliable. It can be obtained based on the operating status of the corresponding acquisition device and data quality indicators, such as whether the device is online, whether there are frame drops, as well as image clarity, point cloud density, thermal image signal-to-noise ratio, etc. The third piece of information is used to describe whether the characteristic response of each modality is obvious in the defect region. It can be obtained based on the response intensity of the defect detection features in the defect region, such as by using response maps, gradient distribution, activation intensity, or degree of local difference.
[0085] In other words, the first piece of information reflects which modality is more suitable for detecting the current defect, the second piece of information reflects which modality is more reliable at present, and the third piece of information reflects which modality responds more obviously to the defect at present, thus providing a basis for determining the subsequent dynamic weights.
[0086] In some embodiments of this application, the step of time-aligning the modal data in the first multimodal data based on the time stamp to obtain the second multimodal data includes: A unified master clock is used to synchronize the time of the acquisition devices corresponding to each mode; The acquisition devices corresponding to each mode are activated based on the trigger signal to acquire data. Based on the time stamps corresponding to each modal data, time matching is performed on different modal data to determine the multimodal data corresponding to the same detection time; The modal data with the highest frame rate is used as the baseline time series. For the other modal data, the corresponding baseline time nodes are matched according to the nearest time marker. When no matching data exists for the target time point, interpolation is performed on the corresponding modal data to obtain second multimodal data that corresponds one-to-one in time.
[0087] Specifically, time alignment refers to ensuring that different modal data correspond to the same detection time under a unified time reference. To achieve this, a unified master clock is first used to synchronize the acquisition devices corresponding to each modality, and then each acquisition device is activated to collect data under a trigger signal, giving the different modal data comparable time markers. Subsequently, based on the time markers corresponding to each modal data, time matching is performed on the different modal data. Specifically, the modal data with the highest frame rate can be used as the reference time series, and the remaining modal data can be matched to the corresponding reference time nodes according to the nearest time marker. When a time node lacks corresponding data, it can be generated through interpolation to complete the data, thereby obtaining second multimodal data with one-to-one temporal correspondence.
[0088] In some embodiments of this application, mapping other modal data in the second multimodal data to the preset reference coordinate system to obtain the third multimodal data includes: Based on the external parameter transformation relationship between each modal corresponding acquisition device and the preset reference coordinate system, the other modal data are transformed into the preset reference coordinate system; The three-dimensional contour data is represented as a point cloud, contour line, or height matrix in the coordinate system of the corresponding acquisition device, and mapped to the preset reference coordinate system according to the external parameter transformation relationship. The temperature distribution data is represented as a temperature matrix in the coordinate system of the corresponding acquisition device, and projected onto the preset reference coordinate system according to the external parameter transformation relationship; Different modal data are made to correspond to the same detection area under the preset reference coordinate system.
[0089] Specifically, based on the extrinsic parameter transformation relationship of each modal acquisition device relative to a preset reference coordinate system, coordinate transformation is first performed on the different modal data. Specifically, three-dimensional contour data can be mapped to the preset reference coordinate system in the form of point clouds, contour lines, or height matrices, and temperature distribution data can be projected onto the preset reference coordinate system in the form of a temperature matrix. Through the above processing, different modal data can correspond to the same detection area on the automotive parts under the same spatial reference, thereby forming the third multimodal data.
[0090] In some embodiments of this application, the vibration detection information includes vibration displacement information and / or vibration acceleration information of the production line frame and / or sensor mounting bracket; The step of determining the preset mapping accuracy requirement based on vibration detection information used to characterize the vibration state of production line racks and / or sensor mounting brackets, and ensuring that the spatial correspondence meets the preset mapping accuracy requirement, includes: Based on the vibration detection information, the rotation and / or translation parameters in the coordinate transformation parameters used to map the other modal data to the preset reference coordinate system are corrected; When the vibration displacement reaches a preset threshold, dynamic recalibration is triggered to update the coordinate transformation parameters; Based on the updated coordinate transformation parameters, the spatial correspondence between each modality data in the second multimodal data and the same detection area is re-established.
[0091] Specifically, the vibration detection information reflects the vibration state of the production line frame and / or sensor mounting bracket during operation, such as displacement or acceleration changes. Based on this vibration detection information, the system corrects the coordinate transformation parameters used when mapping other modal data to a preset reference coordinate system, for example, by adjusting rotation and / or translation parameters to compensate for spatial mapping deviations caused by vibration. When the vibration displacement reaches a preset threshold, dynamic recalibration is further triggered, the coordinate transformation parameters are updated, and the spatial correspondence between each modal data and the same detection area is re-established based on the updated parameters, thereby ensuring spatial alignment accuracy.
[0092] In some embodiments of this application, the step of determining the dynamic weights corresponding to the defect detection features of each modality data based on at least two of the following: first information characterizing the dependence of the defect type to be detected on different modal features; second information characterizing the reliability of the current data of each modality; and third information characterizing the degree of defect detection features of each modality relative to the background response; and then weighting and combining the defect detection features corresponding to the same detection region in each modality data based on the dynamic weights to obtain fused features, includes: Determine the defect type dependency coefficient corresponding to each mode based on the first information; Determine the modality confidence level corresponding to each modality based on the second information; The feature saliency corresponding to each mode is determined based on the third information; Based on at least two of the defect type dependency coefficient, modality confidence, and feature saliency, determine the dynamic weights corresponding to the defect detection features of each modality data; The defect detection features corresponding to the same detection area in each modal data are converted into a unified feature expression; The unified feature representation is combined with the corresponding dynamic weights to obtain the weighted features for each modality; The weighted features corresponding to each modality are weighted and summed and / or concatenated to obtain the fused features.
[0093] Specifically, based on at least two of the first, second, and third pieces of information, the participation level of each modality in the current detection scenario is determined. Then, the defect detection features of each modality within the same detection area are combined according to the participation level. Specifically, the defect type dependency coefficient, modality confidence, and feature saliency corresponding to each modality can be obtained first, and the dynamic weight of each modality can be determined accordingly. Subsequently, the defect detection features of each modality are converted into a corresponding unified feature expression, and then combined with the corresponding dynamic weights to obtain the weighted features of each modality. Finally, a fusion feature is generated through weighted summation and / or weighted concatenation. In other words, the purpose of this step is to give a larger proportion to modalities that are more sensitive to the current defect, have more reliable current data, and have a more obvious current response in the final fusion result, thereby improving the representation effect of the fusion features on the defect.
[0094] In some embodiments of this application, determining the defect detection result of the automotive component based on the fusion features, and determining the process step information corresponding to the defect detection result based on the correlation between the defect detection result and process parameters, includes: Based on the fusion features, at least one of the defect type, defect location, and defect level is determined as the defect detection result; Obtain the timing data of process parameters corresponding to the current automotive parts. The timing data of process parameters includes at least one of stamping pressure, mold temperature, conveying speed, process switching time point, and equipment operating status. Based on the correspondence between the defect detection results and the process parameter time series data, the target process step corresponding to the defect detection results is determined; The correspondence includes at least one of the following: the correspondence between defect location and process operation location, the correspondence between defect type and process anomaly type, and the correspondence between defect detection time and abnormal fluctuation period of process parameters.
[0095] Specifically, the system obtains the defect detection results of the current automotive parts based on fusion features, and then combines this with the time-series data of the corresponding process parameters to trace and locate the source of the defects. The defect detection results can include at least one of the following: defect type, defect location, and defect level. The time-series data of the process parameters can include information such as stamping pressure, die temperature, conveyor speed, process changeover time points, and equipment operating status. The system analyzes the correspondence between the defect detection results and the time-series data of the process parameters—for example, the correspondence between defect location and process application location, defect type and process anomaly type, and the correspondence between defect detection time and periods of abnormal fluctuation in process parameters—to determine the target process step corresponding to the defect detection result, thereby achieving the location of the defect source.
[0096] A second aspect of this application provides a defect detection and process location system, comprising: The acquisition module is used to synchronously acquire first multimodal data of automotive parts in online transportation, and generate time stamps under a unified time reference for each modal data in the first multimodal data; A time alignment module is used to perform time alignment on each modal data in the first multimodal data based on the time marker to obtain the second multimodal data; The spatial alignment module is used to take the coordinate system of the target modal data in the second multimodal data as a preset reference coordinate system, and map other modal data in the second multimodal data to the preset reference coordinate system to obtain the third multimodal data. The third multimodal data is used to characterize the spatial correspondence of the same detection area under different modes, and the spatial correspondence satisfies the preset mapping accuracy requirements determined based on vibration detection information used to characterize the vibration state of the production line frame and / or sensor mounting bracket. The feature extraction module is used to extract defect detection features from each modality of the third multimodal data; The weight determination module is used to determine the dynamic weights corresponding to the defect detection features of each modality data based on at least two of the following: first information used to characterize the dependence of the defect type to be detected on different modal features, second information used to characterize the reliability of the current data of each modality, and third information used to characterize the degree of defect detection features of each modality relative to the background response. The fusion module is used to perform weighted combination of defect detection features corresponding to the same detection area in each modal data based on the dynamic weights, so as to obtain fused features; The determination and positioning module is used to determine the defect detection result of the automotive component based on the fusion features, and to determine the process step information corresponding to the defect detection result according to the correlation between the defect detection result and the process parameters.
[0097] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any of the first aspects above.
[0098] Figure 2 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 2As shown, the electronic device may include a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 may call logical instructions in the memory 830 to execute the method in any of the embodiments of the first aspect described above.
[0099] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0100] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program that can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to perform the methods provided above.
[0101] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the methods provided above.
[0102] Example 2 The objects under inspection are automotive parts being transported online, preferably steel automotive body panels, including the outer panel of the engine hood, the inner panel of the doors, and the outer panel of the trunk lid. Because the workpieces are in continuous motion during transport and the production line has a high cycle time, multimodal data is collected synchronously from the workpieces during the data acquisition phase. The collected multimodal data includes at least two of the following: visible light image data, three-dimensional contour data, and temperature distribution data; preferably, it includes all three types of data. Visible light image data is used to characterize the texture, edges, and grayscale changes of the workpiece surface; three-dimensional contour data is used to characterize the undulations, unevenness, and contour deviations of the workpiece surface; and temperature distribution data is used to characterize thermal anomalies on the workpiece surface. To ensure a one-to-one correspondence between subsequent modal data, a time stamp under a unified time base is generated for each modal data during acquisition.
[0103] After acquiring multimodal data, the time alignment stage begins. The purpose of time alignment is not merely to arrange different modal data according to their acquisition order, but to ensure that different modal data participating in subsequent analysis correspond to the same detection time. In this embodiment, a unified master clock is used to synchronize the time of each modal acquisition device, and a trigger signal activates each modal acquisition device to begin acquisition. After each modal data is formed, time matching is performed based on its corresponding time marker to determine the multimodal data group corresponding to the same detection time. Further, in some implementations, the modal data with the highest frame rate is used as the baseline time series; for the remaining modal data, the corresponding baseline time node is matched according to the nearest time marker; when no matching data exists for the target time node, interpolation is performed on the corresponding modal data to obtain multimodal data with one-to-one temporal correspondence. This method can solve the problem of inconsistent sampling frequencies among different modalities.
[0104] After time alignment, spatial alignment is performed on the multimodal data. The purpose of spatial alignment is to enable different modal data to represent the same area of the workpiece under the same spatial reference. In this embodiment, non-reference modal data are mapped to a preset reference coordinate system. Preferably, the coordinate system corresponding to the visible light image data is used as the preset reference coordinate system; based on the extrinsic parameter transformation relationship of each modal acquisition device relative to the preset reference coordinate system, the three-dimensional contour data and temperature distribution data are mapped to the visible light image coordinate system. In this way, during subsequent feature extraction, data corresponding to the same defect area in different modalities can be processed under a unified coordinate system.
[0105] Considering that vibrations may occur in the production line frame, sensor mounting bracket, and conveying mechanism during operation, causing a shift in the original mapping relationship, this embodiment introduces vibration detection information during the mapping process to correct the coordinate transformation parameters. Specifically, vibration displacement and / or vibration acceleration information of the production line frame and / or sensor mounting bracket are collected; the rotation and / or translation parameters in the coordinate transformation parameters are corrected in real time based on the vibration detection information; when the vibration displacement reaches a preset threshold, dynamic recalibration is triggered to update the coordinate transformation parameters. In this way, the spatial correspondence of multimodal data can be maintained under online operating conditions.
[0106] After obtaining time-aligned and spatially aligned multimodal data, defect detection features are extracted from each modality. For visible light image data, texture variation features, gray-level discontinuity features, edge anomaly features, and local contrast features can be extracted. For 3D contour data, height variation features, curvature variation features, local morphological abrupt changes, and contour deviation features can be extracted. For temperature distribution data, local temperature anomaly features, temperature gradient features, and thermal distribution discontinuity features can be extracted. These features can be extracted using preset rules or through feature extraction networks such as convolutional neural networks.
[0107] In the feature fusion stage, this embodiment does not employ fixed weights or simple concatenation. Instead, it determines the dynamic weights corresponding to each modality's defect detection features based on at least two of the following: defect type information, modal confidence information, and feature saliency information. Fusion is then performed based on these dynamic weights. The defect type dependency coefficient for each modality is determined based on the degree of dependence of the detected defect type on different modalities. The modal confidence of each modality is determined based on the operating status and data quality indicators of each modal acquisition device. The feature saliency of each modality is determined based on the response intensity of the defect region in the feature extraction results. Then, the dynamic weights of each modality are determined based on the above parameters. The region to be detected is divided into multiple local feature regions, and a region-level dynamic weight for each local feature region is determined. The weights of the corresponding modalities are enhanced and adjusted for key pixels or key feature points corresponding to key defect locations. The dynamic weights are updated when the detected defect type, modal confidence, or feature saliency changes by a preset degree.
[0108] After feature fusion is completed, the defect detection results of the automotive parts are determined based on the fused features. The defect detection results include at least one of defect type, defect location, and defect level. This embodiment does not merely output the defect itself, but rather combines the time-series data of the process parameters corresponding to the current automotive part to determine the target process stage corresponding to the defect detection results. The time-series data of the process parameters may include stamping pressure, die temperature, conveyor speed, equipment status, and process changeover time. Based on the correlation between the defect detection results and the process parameters, the system determines that the defect is more likely to form in the stamping contact stage, the die adhesion stage, the conveyor disturbance stage, or other process stages.
[0109] Example 3 The object being inspected is a steel engine hood outer panel for automobiles. After the stamping process, this panel is conveyed to the inspection station by a conveyor belt. The inspection point is located after the stamping process and before the trimming process. A visible light array camera is installed above the inspection station, laser line scanning sensors are installed on both sides, and an infrared thermal imager is installed on the outer side. A photoelectric sensor is installed at the entrance of the inspection station as a trigger source, and vibration sensors are installed on the conveyor belt frame and sensor mounting brackets. The visible light camera is used to acquire images of the workpiece surface, the laser line scanning sensor is used to acquire the three-dimensional contour of the workpiece surface, the infrared thermal imager is used to acquire the surface temperature distribution of the workpiece, the photoelectric sensor is used to detect the workpiece's arrival and trigger acquisition, and the vibration sensor is used to provide vibration detection information.
[0110] The visible light camera is a line scan camera, mounted directly above the conveyor belt. The main camera's optical axis is essentially perpendicular to the workpiece surface, while auxiliary cameras on both sides are arranged at an angle to the workpiece surface to supplement image information in edge areas. Laser line scan sensors are positioned on either side of the main camera, with the laser scanning direction intersecting the conveyor belt's movement direction to obtain the workpiece's lateral contour information. An infrared thermal imager is positioned on one side of the inspection station entrance, covering the main inspection area of the workpiece. After the photoelectric sensor detects the workpiece entering the inspection station, it sends an arrival signal to the synchronization trigger card, which simultaneously triggers the visible light camera, laser line scan sensor, and infrared thermal imager to begin data acquisition. To ensure a unified time reference for the three modal data, an industrial Ethernet clock server is used as the unified master clock in this embodiment, and each acquisition device generates a global timestamp under this unified master clock.
[0111] After data acquisition, the industrial control computer receives the raw data from the three modes. Since the line frequency of the visible light camera is higher than that of the laser line scan sensor and the infrared thermal imager, the timestamp sequence output by the visible light camera is used as the reference time sequence. For both laser and infrared data, the timestamp closest to the reference time node is searched frame by frame. If a corresponding laser or infrared frame does not exist for a given reference time node, linear interpolation is used to complete the contour or temperature data for that time, thus forming a one-to-one temporally corresponding data set. The purpose of this is to ensure that the data entering the subsequent fusion stage all come from the same local moment of the same workpiece, rather than being a mixture of different time slices.
[0112] In the spatial alignment phase, this embodiment uses the visible light main camera coordinate system as the global reference coordinate system. During calibration, a three-dimensional calibration plate adapted to the size of the cover is used. The calibration plate is equipped with checkerboard corner points, cylindrical bosses, and temperature calibration points for visible light, laser, and infrared mode recognition, respectively. First, the intrinsic parameters of the three sensors are calibrated. Then, multiple sets of data are collected from the calibration plate at different positions and orientations, and feature points in each mode are extracted. Using the coordinates of the visible light feature points as a reference, the extrinsic parameter matrices of the laser sensor and the infrared thermal imager relative to the visible light main camera are solved using an iterative nearest-point algorithm. During operation, the laser point cloud is projected onto the visible light image plane, and the infrared temperature pixels are resampled into the visible light image coordinate grid.
[0113] Considering the vibrations present on the production line, this embodiment does not directly use static calibration parameters for extended periods. Instead, it continuously collects displacement and acceleration information from the vibration sensor and inputs this information into the error correction module. If the vibration sensor detects a slight displacement of the mounting bracket, it corrects the translation parameters in the extrinsic parameter matrix in real time; if it detects a change in the bracket's attitude, it corrects the rotation parameters synchronously. When the vibration displacement reaches a preset threshold, the system performs dynamic recalibration, recalculating the extrinsic parameters using pre-reserved standard reference points or quick calibration components within the detection station to restore mapping accuracy.
[0114] Visible light images undergo bilateral filtering, adaptive histogram equalization, workpiece contour ROI cropping, distortion correction, and temperature drift compensation. Laser point cloud data undergoes Gaussian filtering, median filtering, point cloud smoothing, ROI cropping, distortion correction, and temperature drift compensation. Infrared thermal image data undergoes wavelet denoising, pseudo-color enhancement, ROI cropping, distortion correction, and temperature drift compensation. Image frames, point cloud frames, and thermal image frames that do not meet the quality threshold are discarded, and buffer queues are used to fill in the missing frames. After preprocessing, the three-modal data are uniformly converted into floating-point matrices and normalized to a unified range.
[0115] Visible light images are input into a defect feature extraction network to extract surface texture, edge anomalies, and local gray-level abrupt changes. Laser contour matrices are input into a contour feature extraction network to extract height variations, curvature anomalies, and morphological abrupt changes. Infrared temperature matrices are input into a thermal feature extraction network to extract temperature anomalies and temperature gradients. Simultaneously, the time-series data of the current workpiece's process parameters are input into a time-series feature extraction network to extract time-series features related to changes in the stamping process state. This time-series data includes stamping pressure, die temperature, conveyor speed, process time points, and equipment operating status.
[0116] This embodiment employs a dynamic weight fusion method. Specifically, firstly, the defect type dependence coefficients of the three modes are determined based on the candidate defect types. For example, for surface scratches, the visible light mode coefficient is higher, followed by the laser mode, and then the infrared mode. For wrinkles or springback deviations, the laser mode coefficient is higher; for stamping indentations or thermal anomalies with hidden cracks, the infrared mode coefficient is increased. Next, modal confidence is calculated based on indicators such as the current online status of the equipment, image clarity, point cloud density, and thermal image signal-to-noise ratio. The significance of each mode is calculated based on the feature response map of the current defect candidate region. Substituting the above parameters into the dynamic weight calculation formula yields the basic weight for each mode. Subsequently, the workpiece detection area is divided into multiple local feature regions, and region-level weights are calculated for each region. For key locations such as scratch centers, crack endpoints, and temperature anomaly peaks, a key point weight gain is added to the region weights.
[0117] The system integrates feature input for defect classification and localization, outputting defect type, defect location, and defect level. For example, if a linear texture anomaly exceeding a preset threshold is detected in the center of the engine hood outer panel, along with a shallow groove in the corresponding laser profile, and no significant abnormality in the infrared temperature of that area, the system classifies it as a scratch defect. If significant fluctuations in the workpiece edge profile and abnormal curvature are detected, and the corresponding edges are blurred in the visible light image, the system classifies it as a wrinkling or unclear edge defect. If a local depression accompanied by a short-term temperature anomaly and corresponding to a period of abnormal stamping pressure is detected, the system classifies it as a stamping indentation.
[0118] This embodiment does not only output the defect itself, but also determines the defect formation stage by combining process parameter timing data. For example, if the indentation defect coincides with the time period of mold contact pressure increase, the defect is classified as the stamping contact stage; if irregular surface defects correspond to mold temperature fluctuations and material sticking alarm records, the defect is classified as the mold material sticking stage; if hole position offset or springback deviation corresponds to conveyor speed fluctuations and workpiece positioning offset records, the defect is classified as the transmission or forming adjustment stage. The judgment result, along with the workpiece number, is sent to the PLC and industrial control computer. The industrial control computer generates an inspection report, and the PLC controls the audible and visual alarms, rejection, or re-inspection actions. For severely defective parts, the PLC controls the pneumatic pusher to push the workpiece to the non-conforming product channel; for slightly defective parts, they are pushed to the re-inspection channel; if the same type of severe defects occurs continuously, a temporary stop command is triggered.
[0119] Example 4 In this embodiment, multimodal acquisition is not limited to using a visible light array camera, a laser line scan sensor, and an infrared thermal imager simultaneously. As long as at least two of the following can be provided: surface visual information, geometric topography information, and thermal anomaly information, the multimodal acquisition scheme of this application can be constituted.
[0120] For example, a visible light array camera can be replaced by an area array industrial camera, especially in scenarios where the production line cycle is slow or the single piece stays for a long time, the area array camera can also complete the surface image acquisition. The laser line scan sensor can be replaced by a structured light sensor, a stereo vision depth camera, or other three-dimensional contour acquisition device, as long as it can output information on the change in workpiece surface height or contour deviation. Infrared thermal imagers can be replaced by thermal imagers, near-infrared cameras, or other acquisition devices that can characterize the distribution of thermal anomalies. When dealing only with defects in molding dimensions, a dual-modal combination of "visible light + three-dimensional contour" can be used; when dealing only with surface defects of thermal anomalies, a dual-modal combination of "visible light + temperature distribution" can be used.
[0121] Regarding the time synchronization method, this embodiment is not limited to a specific clock server or trigger card model. A unified time reference can be provided by an industrial Ethernet clock server or by a synchronization module built into the edge controller; the trigger signal can be generated by a photoelectric sensor, an encoder, a proximity switch, a position detector, or other components characterizing the workpiece's position. As long as different modal acquisition devices can generate comparable time markers based on the same time reference and perform time alignment accordingly, it falls within the technical scope of this application. For the time alignment algorithm, it is not limited to using nearest-neighbor matching plus linear interpolation. Other variations may also employ spline interpolation, piecewise interpolation, sliding window predictive interpolation, or other methods that achieve one-to-one temporal correspondence.
[0122] Regarding spatial calibration and mapping, this embodiment is not limited to using the visible light image coordinate system as the global reference coordinate system. Although using the visible light image coordinate system as the reference is more convenient for display and manual verification, in other variations, a three-dimensional contour coordinate system can also be used as the reference coordinate system, and the visible light image and temperature distribution can be mapped to this coordinate system; or an independently established workpiece model coordinate system can be used as an intermediate reference, and the three modes can be uniformly mapped to this model coordinate system.
[0123] Similarly, the method for solving extrinsic parameters is not limited to a specific algorithm. In addition to the iterative nearest point algorithm, PnP solving, least squares fitting, feature point matching transformation, or other methods that can establish coordinate transformation relationships between modes can also be used. As long as the data of different modes can ultimately be mapped to the same region of the workpiece under the same spatial reference, the spatial alignment concept of this application can be realized.
[0124] Regarding vibration correction, this embodiment is not limited to requiring the vibration sensor to be installed at a fixed position on the frame and support, nor is it limited to collecting displacement and acceleration. In other variations, attitude angle changes, velocity changes, or support offset trends obtained from multi-sensor fusion can also be collected, and coordinate transformation parameters can be corrected accordingly.
[0125] The modification of coordinate transformation parameters is not limited to directly modifying the rotation matrix and translation vector. In other implementations, the same effect can also be achieved by modifying the projection matrix, compensating for distortion parameters, updating the local region mapping table, or switching the preset calibration parameter group.
[0126] For dynamic recalibration, in addition to automatically triggering when the vibration displacement exceeds the threshold, it can also be set to automatically trigger when the mapping error exceeds the threshold for multiple consecutive frames, the overlap deviation of the defect position exceeds the threshold, or after shift change or equipment maintenance.
[0127] Regarding feature extraction, this embodiment is not limited to using a specific network structure. Feature extraction for the visible light mode, 3D contour mode, and temperature mode can be accomplished by convolutional neural networks, or by a Transformer structure, a hybrid feature extraction network, or a combination of regular features and deep features. Temporal features of process parameters can be extracted using LSTM, GRU, temporal convolutional networks, or other networks capable of characterizing the temporal changes in the process. As long as features meaningful for defect determination and process localization can be extracted from different modes and process timelines, the validity of this application's solution remains unaffected.
[0128] Regarding dynamic fusion, this embodiment does not limit the dynamic weights to using a specific formula. Although in the preferred implementation, the dynamic weights can be calculated by weighted summation based on defect type dependency coefficients, modal confidence, and feature saliency, in other variations, normalized product, piecewise mapping, lookup table mapping, gating mechanisms, or attention weight output methods can also be used to obtain the dynamic weights.
[0129] Similarly, the conditions for weight updates can be adjusted according to the application scenario. For example, in addition to changes in defect type, modal confidence, and feature saliency, changes in workpiece model, production line speed, ambient lighting, ambient temperature, or lens contamination can also be used as triggering factors for weight updates.
[0130] The weighting at the region and keypoint levels is not limited to a fixed grid area plus key pixels. In other implementations, dynamic partitioning can be performed based on candidate defect boxes, segmented regions, CAD functional areas, or process-sensitive areas. As long as the fusion process still reflects the idea of "adjusting the participation ratio of each modality according to the current defect scenario," it is considered an equivalent implementation of the dynamic fusion in this application.
[0131] Regarding process localization, this embodiment is not limited to using rule-based association. In addition to matching based on "defect type + location + abnormal process parameter period", a joint model of defect features and process time sequence features can also be used to output the target process step; alternatively, one or more candidate process steps and their confidence levels can be output.
[0132] The scope of the process steps is not limited to stamping contact, die sticking, and conveyor disturbance. For different production lines, it can also be extended to feeding, positioning, edge pressing, forming, unloading, or post-processing. As long as the essence is still to locate the source of defect formation based on the correlation between defect detection results and process parameters, it does not depart from the technical concept of this application.
[0133] In terms of application scenarios, this embodiment is not limited to automotive steel body panels. For metal stamping parts, 3C electronic structural components, new energy battery casings, aluminum alloy sheets, or other online conveying parts, as long as there is a need for multimodal defect detection and process positioning, it can be adapted by adjusting the sensor resolution, installation distance, detection area range, model input size, weight parameters, and judgment threshold. The corresponding defect types can be expanded from scratches, dents, wrinkles, and hole misalignment to include burrs, pinholes, uneven coating, mounting misalignment, and edge cracking.
[0134] Regarding the execution entity, this embodiment is not limited to being jointly performed by an edge box and an industrial control computer. As long as the synchronous acquisition and control, spatiotemporal alignment, vibration correction, feature extraction, dynamic fusion, defect determination, and process localization processes in this application can be executed, the execution entity can be a single industrial control computer, a GPU server, an edge computing terminal, or a distributed computing system. The program can be stored in software form on a storage medium or embedded in a dedicated board or embedded system.
[0135] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
[0136] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A defect detection and process location method, characterized in that, include: First multimodal data is synchronously collected for automotive parts in online transport, and time stamps under a unified time reference are generated for each modal data in the first multimodal data. Based on the time stamps, the time of each modal data in the first multimodal data is time-aligned to obtain second multimodal data. The coordinate system of the target modal data in the second multimodal data is used as the preset reference coordinate system, and other modal data in the second multimodal data are mapped to the preset reference coordinate system to obtain the third multimodal data. The third multimodal data is used to characterize the spatial correspondence of the same detection area under different modes, and the spatial correspondence satisfies the preset mapping accuracy requirements determined based on vibration detection information used to characterize the vibration state of the production line frame and / or sensor mounting bracket. Defect detection features are extracted from each modality of the third multimodal data, and dynamic weights corresponding to the defect detection features of each modality are determined based on at least two of the first, second, and third information. Based on the dynamic weights, the defect detection features corresponding to the same detection area in each modal data are weighted and combined to obtain fused features; Based on the fusion features, the defect detection results of the automotive parts are determined, and based on the correlation between the defect detection results and process parameters, the process step information corresponding to the defect detection results is determined.
2. The method according to claim 1, characterized in that, The first multimodal data includes visible light image data, three-dimensional contour data, and temperature distribution data; The visible light image data is used to characterize the texture information and / or grayscale variation information of the automotive parts surface, the three-dimensional contour data is used to characterize the height variation information and / or topographic undulation information of the automotive parts surface, and the temperature distribution data is used to characterize the local thermal anomaly information of the automotive parts surface.
3. The method according to claim 1, characterized in that, The target modal data is one modal data in the second multimodal data, and the other modal data are the remaining modal data in the second multimodal data other than the target modal data; The target modal data is used to provide the preset reference coordinate system, and the other modal data is used to map to the preset reference coordinate system to form the third multimodal data.
4. The method according to claim 1, characterized in that, The first information is information used to characterize the degree to which the type of defect to be detected depends on different modal features; The first information is obtained through at least one of the following methods: Candidate defect types are determined based on the pre-classification results corresponding to at least one modality of data; The degree of dependence of different modes on the candidate defect types is determined based on the pre-defined correspondence between defect types and modal characterization capabilities. The degree of dependence of different modalities on the candidate defect type is determined based on the contribution of different modalities output by the training model to the current defect determination; The second piece of information is used to characterize the reliability of the current data for each modality; The second information is obtained based on the operating status and data quality indicators of the acquisition devices corresponding to each mode; The operating status includes at least one of the following: device online status, frame loss status, abnormal alarm status, and operating temperature status; the data quality indicators include at least one of the following: image clarity, point cloud density, thermal image signal-to-noise ratio, and effective area coverage. The third information is information used to characterize the degree of response of the defect detection features of each modality to the background; The third information is obtained based on the response intensity of the defect detection features of each modality in the defect region; The response intensity is determined based on at least one of the following: response map, gradient distribution, activation intensity, and degree of local difference.
5. The method according to claim 1, characterized in that, The step of time-aligning the modal data in the first multimodal data based on the time stamp to obtain the second multimodal data includes: A unified master clock is used to synchronize the time of the acquisition devices corresponding to each mode; The acquisition devices corresponding to each mode are activated based on the trigger signal to acquire data. Based on the time stamps corresponding to each modal data, time matching is performed on different modal data to determine the multimodal data corresponding to the same detection time; The modal data with the highest frame rate is used as the baseline time series. For the other modal data, the corresponding baseline time nodes are matched according to the nearest time marker. When no matching data exists for the target time point, interpolation is performed on the corresponding modal data to obtain second multimodal data that corresponds one-to-one in time.
6. The method according to claim 1, characterized in that, The step of mapping other modal data in the second multimodal data to the preset reference coordinate system to obtain the third multimodal data includes: Based on the external parameter transformation relationship between each modal corresponding acquisition device and the preset reference coordinate system, the other modal data are transformed into the preset reference coordinate system; The three-dimensional contour data is represented as a point cloud, contour line, or height matrix in the coordinate system of the corresponding acquisition device, and mapped to the preset reference coordinate system according to the external parameter transformation relationship. The temperature distribution data is represented as a temperature matrix in the coordinate system of the corresponding acquisition device, and projected onto the preset reference coordinate system according to the external parameter transformation relationship; Different modal data are made to correspond to the same detection area under the preset reference coordinate system.
7. The method according to claim 1, characterized in that, The vibration detection information includes vibration displacement information and / or vibration acceleration information of the production line frame and / or sensor mounting bracket; The step of determining the preset mapping accuracy requirement based on vibration detection information used to characterize the vibration state of production line racks and / or sensor mounting brackets, and ensuring that the spatial correspondence meets the preset mapping accuracy requirement, includes: Based on the vibration detection information, the rotation and / or translation parameters in the coordinate transformation parameters used to map the other modal data to the preset reference coordinate system are corrected; When the vibration displacement reaches a preset threshold, dynamic recalibration is triggered to update the coordinate transformation parameters; Based on the updated coordinate transformation parameters, the spatial correspondence between each modality data in the second multimodal data and the same detection area is re-established.
8. The method according to claim 1, characterized in that, The process involves determining dynamic weights corresponding to the defect detection features of each modality based on at least two of the following: first information characterizing the dependence of the defect type to different modal features; second information characterizing the reliability of the current data for each modality; and third information characterizing the degree of response of the defect detection features of each modality to the background. Based on these dynamic weights, the defect detection features corresponding to the same detection region in each modality are weighted and combined to obtain fused features, including: Determine the defect type dependency coefficient corresponding to each mode based on the first information; Determine the modality confidence level corresponding to each modality based on the second information; The feature saliency corresponding to each mode is determined based on the third information; Based on at least two of the defect type dependency coefficient, modality confidence, and feature saliency, determine the dynamic weights corresponding to the defect detection features of each modality data; The defect detection features corresponding to the same detection area in each modal data are converted into a unified feature expression; The unified feature representation is combined with the corresponding dynamic weights to obtain the weighted features for each modality; The weighted features corresponding to each modality are weighted and summed and / or concatenated to obtain the fused features.
9. The method according to claim 1, characterized in that, The step of determining the defect detection result of the automotive component based on the fusion features, and determining the process step information corresponding to the defect detection result based on the correlation between the defect detection result and process parameters, includes: Based on the fusion features, at least one of the defect type, defect location, and defect level is determined as the defect detection result; Obtain the timing data of process parameters corresponding to the current automotive parts. The timing data of process parameters includes at least one of stamping pressure, mold temperature, conveying speed, process switching time point, and equipment operating status. Based on the correspondence between the defect detection results and the process parameter time series data, the target process step corresponding to the defect detection results is determined; The correspondence includes at least one of the following: the correspondence between defect location and process operation location, the correspondence between defect type and process anomaly type, and the correspondence between defect detection time and abnormal fluctuation period of process parameters.
10. A defect detection and process location system, characterized in that, include: The acquisition module is used to synchronously acquire first multimodal data of automotive parts in online transportation, and generate time stamps under a unified time reference for each modal data in the first multimodal data; A time alignment module is used to perform time alignment on each modal data in the first multimodal data based on the time marker to obtain the second multimodal data; The spatial alignment module is used to take the coordinate system of the target modal data in the second multimodal data as a preset reference coordinate system, and map other modal data in the second multimodal data to the preset reference coordinate system to obtain the third multimodal data. The third multimodal data is used to characterize the spatial correspondence of the same detection area under different modes, and the spatial correspondence satisfies the preset mapping accuracy requirements determined based on vibration detection information used to characterize the vibration state of the production line frame and / or sensor mounting bracket. The feature extraction module is used to extract defect detection features from each modality of the third multimodal data; The weight determination module is used to determine the dynamic weights corresponding to the defect detection features of each modality data based on at least two of the following: first information used to characterize the dependence of the defect type to be detected on different modal features, second information used to characterize the reliability of the current data of each modality, and third information used to characterize the degree of defect detection features of each modality relative to the background response. The fusion module is used to perform weighted combination of defect detection features corresponding to the same detection area in each modal data based on the dynamic weights, so as to obtain fused features; The determination and positioning module is used to determine the defect detection result of the automotive component based on the fusion features, and to determine the process step information corresponding to the defect detection result according to the correlation between the defect detection result and the process parameters.