A three-flow fusion visual inspection method and system for industrial sites
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-08-11
AI Technical Summary
其中,单一视觉数据源的检测系统仅依赖工业相机采集的图像或视频数据流,通过传统图像处理算法或深度学习模型完成检测任务,该类系统仅能获取视觉维度的信息,缺乏对设备物理运行状态、生产工艺节拍的上下文感知,在面对工业现场常见的光照突变、切削液飞溅、设备微颤、油污遮挡等复杂工况时,极易出现误检、漏检,鲁棒性极差
[0007] This specification's embodiments utilize a three-level collaborative triggering mechanism synchronized with a hardware-level microsecond-level clock. Using the trigger moment as a reference, it completes the timeline binding of control flow, event flow, and data flow from the data acquisition source. Furthermore, through spatiotemporal dual feature alignment, it overcomes the shortcomings of spatiotemporal fragmentation in multi-source heterogeneous data. Simultaneously, by verifying the authenticity of visual inspection results through the physical signals of the event flow, it can effectively distinguish between real defects and environmental interference, reducing both the false detection rate and the missed detection rate under complex industrial conditions.
Smart Images

Figure CN122550604A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent manufacturing technology, specifically to a three-flow fusion visual inspection method and system for industrial sites. Background Technology
[0002] With the continuous development of intelligent manufacturing, the intelligent upgrading of factories has placed high demands on real-time monitoring of the production process and full-process control of product quality, requiring high precision, high real-time performance, and strong environmental adaptability. Industrial vision inspection technology, as a core sensing means for intelligent production lines, has been widely applied in scenarios such as defect detection, dimensional measurement, and target positioning. Existing industrial vision inspection technologies are mainly single-vision data source or trigger-assisted vision inspection systems. Among them, single-vision data source inspection systems rely solely on image or video data streams collected by industrial cameras, completing inspection tasks through traditional image processing algorithms or deep learning models. These systems can only acquire information in the visual dimension, lacking contextual awareness of the physical operating status of equipment and the production process rhythm. When faced with complex working conditions commonly encountered in industrial environments, such as sudden changes in lighting, cutting fluid splashes, equipment vibrations, and oil contamination, they are prone to false positives and false negatives, exhibiting extremely poor robustness. Furthermore, trigger-assisted vision inspection systems deploy simple sensors such as photoelectric switches and limit switches on the production line. These systems use switching signals to trigger a camera to capture images after the workpiece has reached its designated position. However, these systems rely solely on the triggering function of auxiliary sensor signals and cannot analyze the core information contained within the signals, such as the physical state of the equipment and changes in operating conditions. They also cannot distinguish between workpiece arrival signals and interference signals caused by abnormal equipment vibration. Additionally, these systems do not connect to the equipment control flow data of the production line's PLC. Their vision algorithms use fixed detection thresholds and model parameters, making it impossible to automatically adapt the detection logic when the production line switches product specifications or adjusts process parameters. This necessitates manual shutdown and program reconfiguration, resulting in long production line downtime, poor flexible production capabilities, and an inability to adapt to modern production models involving multiple varieties and small batches. Moreover, the time granularity, sampling mode, and data dimensions of different types of data in industrial settings are completely mismatched. Existing technologies can only achieve coarse-grained frame-level hard trigger alignment, failing to achieve precise feature-level correspondence across the entire time sequence. This easily leads to misalignment between the physical state and the visual image at the same moment. Therefore, existing industrial vision inspection technologies generally suffer from defects such as single data source, spatiotemporal fragmentation of multi-source heterogeneous data, disconnect between perception and control, poor robustness in complex working conditions, and insufficient flexible adaptability, which cannot meet the industrial field's demand for high-precision, real-time, and highly adaptable visual perception. Summary of the Invention
[0003] This specification provides a three-stream fusion visual inspection method and system for industrial environments, the technical solution of which is as follows:
[0004] Firstly, embodiments of this specification provide a three-stream fusion visual inspection method for industrial sites, including: synchronizing the clocks of a visual sensor, an auxiliary sensor, and an industrial control terminal in the industrial site; and synchronously acquiring the data stream output by the visual sensor, the event stream output by the auxiliary sensor, and the control stream output by the industrial control terminal based on a three-level collaborative triggering mechanism; mapping the discrete data of the event stream and the control stream to their respective acquisition timestamps using the image frame timestamps of the data stream as a reference, thus achieving time alignment; establishing a coordinate system transformation relationship between the visual sensor and the industrial equipment, and mapping the physical location information in the event stream to the image of the data stream. The system calculates pixel coordinates for spatial alignment; outputs temporally and spatially aligned three-stream data; extracts features from the three-stream data to obtain visual features, physical state features, and process prior features; uses the process prior features as a query vector and dynamically calculates the fusion weights of visual features and physical state features based on an attention mechanism to generate deep fusion features; inputs the deep fusion features into a pre-trained decision model to output detection results; when the detection results meet preset conditions, it sends control commands to the industrial control terminal through the industrial communication interface; and feeds the detection results back to the attention mechanism to dynamically adjust the fusion weights or decision model parameters to complete adaptive optimization.
[0005] Secondly, this specification provides a three-stream fusion visual inspection system for industrial sites, comprising: a data acquisition module for synchronizing the clocks of the visual sensor, auxiliary sensor, and industrial control terminal in the industrial site, and synchronously acquiring the data stream output by the visual sensor, the event stream output by the auxiliary sensor, and the control stream output by the industrial control terminal based on a three-level collaborative triggering mechanism; a spatiotemporal alignment module for mapping the discrete data of the event stream and control stream to the respective acquisition times of the data stream, using the image frame timestamps of the data stream as a reference, thereby completing time alignment; establishing a coordinate system transformation relationship between the visual sensor and the industrial equipment, and mapping the physical location information in the event stream to the image frames of the data stream. The system consists of a primitive coordinate system for spatial alignment, outputting time-aligned and spatially aligned three-stream data; a multimodal fusion module for feature extraction from the three-stream data to obtain visual features, physical state features, and process prior features; and a deep fusion feature generated by dynamically calculating the fusion weights of visual features and physical state features based on an attention mechanism using the process prior features as the query vector; a loop closure detection module for inputting the deep fusion features into a pre-trained decision model to output detection results; a control command sent to the industrial control terminal via an industrial communication interface when the detection results meet preset conditions; and a feedback mechanism to dynamically adjust the fusion weights or decision model parameters to achieve adaptive optimization.
[0006] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:
[0007] This specification's embodiments utilize a three-level collaborative triggering mechanism synchronized with a hardware-level microsecond-level clock. Using the trigger moment as a reference, it completes the timeline binding of control flow, event flow, and data flow from the data acquisition source. Furthermore, through spatiotemporal dual feature alignment, it overcomes the shortcomings of spatiotemporal fragmentation in multi-source heterogeneous data. Simultaneously, by verifying the authenticity of visual inspection results through the physical signals of the event flow, it can effectively distinguish between real defects and environmental interference, reducing both the false detection rate and the missed detection rate under complex industrial conditions.
[0008] On the other hand, the embodiments of this specification can also use process prior features as query vectors, dynamically calculate the fusion weights of visual features and physical state features based on attention mechanisms, and generate deep fusion features. That is, the embodiments of this specification introduce control flow as a priori guiding signal into the visual perception system, and predict changes in production line processes in advance through the control flow pre-trigger mechanism, automatically complete the end-to-end adaptive switching of acquisition parameters, detection models, and fusion weights, without the need for manual shutdown configuration, shortening the downtime of production line product changeover, and adapting to the flexible production needs of multiple varieties and small batches.
[0009] On the other hand, the embodiments of this specification can also adopt a tiered acquisition strategy with hierarchical triggering, which only initiates full high-precision acquisition under abnormal operating conditions and only acquires necessary data under normal operating conditions, reducing invalid data transmission and computational overhead, significantly reducing the computing load of the edge computing unit, shortening the system's response time from abnormal detection to control command issuance, and meeting the needs of real-time closed-loop control in industrial sites.
[0010] On the other hand, the embodiments of this specification can also feed the detection results back to the attention mechanism to dynamically adjust the fusion weights or decision model parameters to complete adaptive optimization. That is, the embodiments of this specification construct a closed-loop link of detection, decision, control and feedback. It can not only output accurate detection results based on fusion features, but also feed the results back to the production line control system in real time to realize real-time intervention of anomalies and dynamic optimization of process parameters. This breaks the limitation of the disconnect between perception and control in the existing technology and realizes the upgrade of industrial vision from passive detection to active control.
[0011] Furthermore, the embodiments of this specification can also set a trigger latching mechanism to completely save all three-stream time-series data before and after the abnormal event is triggered, providing complete contextual data support for the root cause analysis of product quality abnormalities and equipment failures; at the same time, the collected data with spatiotemporal labels and operating condition information can be directly used as high-quality labeled samples for model iteration, greatly reducing the cost of sample labeling and realizing continuous optimization and upgrading of the system. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram illustrating the application scenario of the three-flow fusion visual inspection system for industrial sites provided in this manual.
[0014] Figure 2 This is a flowchart illustrating the three-flow fusion visual inspection method for industrial sites provided in this manual.
[0015] Figure 3 This is a flowchart illustrating the process of generating deep fusion features provided in this manual.
[0016] Figure 4 This is a schematic diagram of the structure of the three-flow fusion vision inspection system for industrial sites provided in this manual. Detailed Implementation
[0017] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings.
[0018] The terms "first," "second," etc., in the description, claims, and accompanying drawings are used to distinguish different objects and not to describe a particular order. Furthermore, the term "comprising" and any variations thereof are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0019] This specification provides a three-flow fusion visual inspection method for industrial sites through multiple embodiments. The execution subject of this three-flow fusion visual inspection method for industrial sites can be a three-flow fusion visual inspection system for industrial sites provided in the embodiments of this invention.
[0020] Before this specification elaborates on the three-flow fusion visual inspection method for industrial sites in conjunction with one or more embodiments, it first introduces the application scenarios of this three-flow fusion visual inspection method for industrial sites.
[0021] Please see Figure 1 , Figure 1This is a schematic diagram illustrating an application scenario of a three-stream fusion visual inspection method for industrial sites provided in an embodiment of the present invention. In this embodiment, the three-stream fusion visual inspection system 100 for industrial sites may include an industrial site equipment layer 110, an edge computing unit 120, etc. The industrial site equipment layer 110 is the source and control execution terminal for multi-source heterogeneous data such as data stream, event stream, and control stream. The industrial site equipment layer 110 may include visual sensors (industrial cameras, infrared thermal imagers, etc.), auxiliary sensors (vibration sensors, acoustic emission sensors, photoelectric switch sensors, temperature sensors, etc.), PLC controllers, actuators (rejection cylinders, audible and visual alarms, emergency stop modules, etc.), etc.
[0022] In this embodiment, the edge computing unit 120 can adopt the NVIDIA Jetson Xavier NX edge computing platform, equipped with the Ubuntu 20.04 operating system, running the TensorRT deep learning inference framework and real-time data processing program, communicating with all field devices through an industrial Ethernet switch, supporting the IEEE 1588 PTP precision time protocol, and realizing microsecond-level clock synchronization of all devices.
[0023] In this embodiment, the edge computing unit 120 may include a data acquisition layer, a preprocessing and feature alignment layer, a multimodal fusion layer, and a decision-making and closed-loop control layer. The data acquisition layer may include a data acquisition module, a signal conditioning module, a PLC communication interface module, a clock synchronization module, and a trigger control module; the data acquisition module may include an image acquisition module. The preprocessing and feature alignment layer may include a data preprocessing module and a spatiotemporal alignment module; the spatiotemporal alignment module may include a time alignment module and a spatial calibration module. The multimodal fusion layer may include a multimodal fusion module; the multimodal fusion module may include a feature extraction module and a cross-modal attention fusion module; the feature extraction module may include a visual feature extraction module and a multi-source feature extraction module. The decision-making and closed-loop control layer may include a closed-loop detection module and a closed-loop control feedback module; the closed-loop detection module may include a defect identification and classification module and a result output module.
[0024] In this embodiment, the vision sensor of the industrial field equipment layer 110 is used to output data streams and is connected to the image acquisition module of the data acquisition layer; the auxiliary sensor is used to output event streams and is connected to the signal conditioning module of the data acquisition layer; the PLC controller is used to output control streams and receive closed-loop control commands, and is bidirectionally connected to the PLC communication interface module of the data acquisition layer and the closed-loop control feedback module of the decision and closed-loop control layer; the actuator is used to receive control commands issued by the PLC controller and is electrically connected to the PLC controller.
[0025] In this embodiment, the data acquisition layer within the edge computing unit 120 corresponds to functions such as three-stream data (data stream, event stream, and control stream) acquisition, clock synchronization, and collaborative triggering. The image acquisition module connects to a vision sensor, receiving the raw image data stream, and also connects to the data preprocessing module of the preprocessing and feature alignment layer. The signal conditioning module connects to an auxiliary sensor, receiving the raw event stream data, and also connects to the data preprocessing module of the preprocessing and feature alignment layer. The PLC communication interface module bidirectionally connects to the PLC controller, receiving the raw control stream data, and also connects to the data preprocessing module of the preprocessing and feature alignment layer. The clock synchronization module bidirectionally connects to the image acquisition module, signal conditioning module, and PLC communication interface module, and also interfaces with the vision sensor, auxiliary sensor, and PLC controller of the industrial field equipment layer 110. The clock synchronization module outputs an IEEE 1588 PTP precision clock synchronization signal, achieving microsecond-level hardware time synchronization across all devices. The trigger control module bidirectionally connects to the image acquisition module, signal conditioning module, PLC communication interface module, and clock synchronization module, receiving three-level trigger signals and issuing hardware trigger acquisition commands to achieve three-stream collaborative triggering acquisition and data latching.
[0026] In this embodiment, the preprocessing and feature alignment layer within the edge computing unit 120 is used for data standardization and spatiotemporal feature alignment. The data preprocessing module is connected to the image acquisition module, signal conditioning module, PLC communication interface module, etc., to receive the three streams of raw data, perform corresponding preprocessing operations such as denoising, filtering, and normalization, and output standardized three streams of data; it is also connected to the time alignment module and spatial calibration module, etc. The time alignment module receives the three streams of data output by the data preprocessing module, completes unified time axis mapping alignment, and is connected to the visual feature extraction module, multi-source feature extraction module, etc. The spatial calibration module receives the three streams of data output by the data preprocessing module, completes spatial mapping alignment between pixel coordinates and device physical coordinates, and is connected to the visual feature extraction module, multi-source feature extraction module, etc.
[0027] In this embodiment, the visual feature extraction module in the multimodal fusion layer within the edge computing unit 120 is connected to the temporal alignment module and the spatial calibration module. It receives the aligned data stream, extracts visual features, and is also connected to the cross-modal attention fusion module. The multi-source feature extraction module is connected to the temporal alignment module and the spatial calibration module. It receives the aligned event stream and control stream, extracts physical state features and process prior features respectively, and is also connected to the cross-modal attention fusion module. The cross-modal attention fusion module is connected to both the visual feature extraction module and the multi-source feature extraction module. It receives three types of feature vectors: visual features, physical state features, and process prior features. Guided by the control stream, it performs dynamic weight fusion, outputs fused features, and is also connected to the defect identification and classification module.
[0028] In this embodiment, the decision-making and closed-loop control layer within the edge computing unit 120 is used for outputting detection results and coordinating closed-loop production line operations. The defect identification and classification module connects to the cross-modal attention fusion module, receives fused features, and outputs detection results (defect category, quality score, anomaly level, etc.). It also connects to the result output module and the closed-loop control feedback module. The result output module connects to the defect identification and classification module for real-time display, local storage, and historical tracing of detection results. The closed-loop control feedback module connects to the defect identification and classification module and also bidirectionally connects to the PLC controller. It sends real-time control commands to the PLC and simultaneously feeds back operating information to the front-end module, completing the entire closed loop of detection, decision-making, control, and feedback.
[0029] This embodiment provides a three-stream fusion visual inspection system for industrial sites, which can perform the following: clock synchronization of the visual sensor, auxiliary sensor, and industrial control terminal in the industrial site; synchronous acquisition of the data stream output by the visual sensor, the event stream output by the auxiliary sensor, and the control stream output by the industrial control terminal based on a three-level collaborative triggering mechanism; mapping the discrete data of the event stream and control stream to the respective acquisition times of the data stream using the image frame timestamps of the data stream as a reference, thus completing time alignment; establishing a coordinate system transformation relationship between the visual sensor and the industrial equipment, and mapping the physical location information in the event stream to the image pixel coordinates of the data stream. The system performs spatial alignment; outputs time-aligned and spatially aligned three-stream data; extracts features from the three-stream data to obtain visual features, physical state features, and process prior features; uses the process prior features as the query vector, dynamically calculates the fusion weights of visual features and physical state features based on an attention mechanism, and generates deep fusion features; inputs the deep fusion features into a pre-trained decision model to output detection results; when the detection results meet preset conditions, sends control commands to the industrial control terminal through the industrial communication interface; and feeds the detection results back to the attention mechanism to dynamically adjust the fusion weights or decision model parameters to complete adaptive optimization, etc.
[0030] It should be noted that, Figure 1 The schematic diagram of the application scenario of the three-flow fusion visual inspection system for industrial sites shown is merely an example. The three-flow fusion visual inspection system and scenario for industrial sites described in this embodiment of the invention are for the purpose of more clearly illustrating the technical solutions of this embodiment of the invention, and do not constitute a limitation on the technical solutions provided by this embodiment of the invention. As those skilled in the art will know, with the evolution of three-flow fusion visual inspection systems for industrial sites and the emergence of new scenarios, the technical solutions provided by this embodiment of the invention are also applicable to similar technical problems.
[0031] Please see Figure 2 , Figure 2This is a flowchart illustrating a three-flow fusion visual inspection method for industrial environments provided by an embodiment of the present invention. This three-flow fusion visual inspection method for industrial environments can be implemented by... Figure 1 The illustrated three-flow fusion visual inspection system 100 for industrial environments is implemented. This three-flow fusion visual inspection method for industrial environments includes:
[0032] 200. Perform clock synchronization on the vision sensors, auxiliary sensors and industrial control terminals in the industrial field, and based on a three-level collaborative triggering mechanism, synchronously collect the data stream output by the vision sensors, the event stream output by the auxiliary sensors and the control stream output by the industrial control terminal.
[0033] In this embodiment, the three-stream data refers to three heterogeneous data streams generated in the industrial field, including data stream, event stream, and control stream. The data stream consists of continuous video frames or still image data acquired by visual sensors (including industrial cameras and infrared thermal imagers); the event stream consists of discrete alarm signals, status change logs, and continuous physical state data acquired by auxiliary sensors (including vibration sensors, acoustic emission sensors, photoelectric switches, and temperature sensors); and the control stream consists of equipment operation status commands, process parameter settings, and equipment operation feedback signals issued by the PLC controller or host computer. The alignment operation in this embodiment is the process of mapping data with different modalities, different sampling frequencies, and different timestamps to a unified time axis and feature space.
[0034] In this embodiment, the vision sensor can be a global shutter CCD industrial camera, supporting hardware triggering, with a fixed-focus industrial lens, installed above the production line inspection station, and equipped with a ring light source to eliminate shadow interference. Auxiliary sensors may include piezoelectric vibration sensors, acoustic emission sensors, photoelectric switches, infrared temperature sensors, etc.; the vibration sensor is installed on the equipment spindle base; the acoustic emission sensor has a sampling frequency of 20kHz to 1MHz and is installed at the processing station; the photoelectric switch is installed at the entrance of the inspection station for workpiece arrival detection. The PLC controller can be a Siemens S7-1200 series PLC with built-in Modbus / TCP communication protocol, storing control parameters such as spindle speed, process recipe number, station feed status, and equipment operating mode in its registers; the actuator may include a rejection cylinder, an audible and visual alarm, an emergency stop control module, etc., and the actuator is electrically connected to the PLC controller to receive control commands issued by the PLC controller.
[0035] In some embodiments, a vision sensor, an auxiliary sensor, and an industrial control terminal are deployed in the industrial field. The industrial control terminal includes at least a PLC controller and an actuator. The data stream includes at least a continuous video frame sequence and static high-resolution image data acquired by the vision sensor. The event stream includes at least equipment physical status data, discrete alarm signals, and status change logs acquired by the auxiliary sensor. The control stream includes at least equipment operating status instructions, process parameter setpoints, and equipment operating feedback signals read from the registers of the industrial control terminal.
[0036] In some embodiments, based on a three-level collaborative triggering mechanism, the data stream output by the vision sensor, the event stream output by the auxiliary sensor, and the control stream output by the industrial control terminal are simultaneously acquired. The method further includes: determining the trigger time and, based on the trigger time, latching the full three-stream timing data corresponding to the data stream, event stream, and control stream within a preset time window before and after the trigger time.
[0037] In some embodiments, the three-level collaborative triggering mechanism includes control flow pre-triggering, event flow precise triggering, and data flow self-triggering. Control flow pre-triggering includes using product changeover, process adjustment, and station feed instructions issued by the PLC controller as pre-trigger signals. When the system receives the pre-trigger signal, it wakes up the vision sensor and auxiliary sensor to enter the acquisition state and preloads the acquisition parameters, decision model, and fusion weights for the corresponding working condition. Event flow precise triggering includes using the signal from the auxiliary sensor as the acquisition trigger source, supporting at least one triggering logic among threshold triggering, feature triggering, and switch quantity triggering, and simultaneously starting the vision sensor to acquire images when triggered. Data flow self-triggering includes self-triggering operations through frame difference, ROI area change detection, and suspected defect pre-identification logic. The self-triggering operations include at least automatic acquisition of workpiece in place, environmental interference parameter compensation, and continuous shooting for abnormal verification.
[0038] In this embodiment, after the system powers on, the PTP clock of all devices is first synchronized via the clock synchronization module, achieving a synchronization accuracy of ±1μs. The trigger control module then initiates a three-level collaborative triggering mechanism to synchronously acquire the three streams of data. The three-level collaborative triggering mechanism includes:
[0039] Control flow pre-triggering: The PLC communication interface module monitors the PLC register data in real time. When it detects control flow change signals such as changes in process recipe number or spindle speed adjustment, it immediately sends a pre-triggering command to the trigger control module. The system preloads the camera parameters, sensor sampling frequency, detection model and fusion weight configuration of the corresponding process recipe, and at the same time wakes up the vision sensor and auxiliary sensor to enter the waiting state.
[0040] Precise event flow triggering: When the workpiece enters the inspection station, the photoelectric switch sends a switch signal indicating that the workpiece has arrived. After receiving the signal, the trigger control module immediately sends a hardware trigger signal to the vision sensor, triggering the camera to expose and acquire a high-resolution image. At the same time, it latches all the data from the vibration sensor and acoustic emission sensor at the current moment, as well as the spindle speed, process formula, and other control flow data in the PLC register.
[0041] Data stream self-triggering: The image acquisition module performs real-time frame differential processing on the acquired images. When an anomaly such as a suspected defect or sudden change in illumination is detected in the ROI area, a self-triggering command is immediately sent to start high-speed continuous shooting of the camera and trigger high-frequency sampling of the auxiliary sensor to complete the anomaly verification data acquisition.
[0042] In some embodiments, based on a three-level collaborative triggering mechanism, the data stream output by the vision sensor, the event stream output by the auxiliary sensor, and the control stream output by the industrial control terminal are simultaneously acquired. This includes: adopting a hierarchical acquisition strategy, including the following different acquisition modes: Level 1: Regular inspection mode: triggered by the control stream station cycle signal, performing low frame rate global image acquisition and low-frequency sampling by the auxiliary sensor; Level 2: Precision detection mode: triggered by the event stream workpiece arrival signal, performing high-resolution target frame acquisition and latching all working condition data; Level 3: Anomaly verification mode: triggered by the data stream suspected defect signal or the event stream anomaly signal, performing high-speed continuous shooting and high-frequency synchronous acquisition by multiple sensors; Level 4: Safety traceability mode: triggered by the control stream fault command or the event stream emergency alarm signal, latching all three-stream time-series data within a preset time window before and after the trigger time.
[0043] In this embodiment, the system can automatically switch between four acquisition modes according to the working conditions during the data acquisition process. In the normal inspection mode, the computing power consumption is reduced by 70% to 85%, and in the abnormal verification mode, the full amount of data is acquired with high precision, taking into account both real-time performance and detection accuracy.
[0044] 210. Using the image frame timestamps of the data stream as a reference, map the discrete data of the event stream and control stream to the acquisition times of the data stream respectively to complete time alignment; establish the coordinate system transformation relationship between the vision sensor and the industrial equipment, and map the physical location information in the event stream to the image pixel coordinates of the data stream to complete spatial alignment; output the three streams of data after time alignment and spatial alignment.
[0045] In some embodiments, before time alignment is completed, the data stream, event stream, and control stream are preprocessed respectively; the preprocessing includes: denoising, distortion correction, parameter normalization, and region of interest (ROI) extraction for the data stream; filtering, baseline correction, peak detection, and feature dimension reduction for the event stream; and execution instruction parsing, parameter normalization, and state encoding for the control stream.
[0046] Before time alignment, the system in this embodiment can perform data preprocessing through a data preprocessing module. For the data stream: Gaussian filtering is used to remove image noise, distortion correction is performed using camera calibration parameters, image pixel values are normalized from 0 to 1, the Region of Interest (ROI) is extracted based on the workpiece position, and irrelevant background interference is removed. For the event stream: Butterworth filtering is used to remove low-frequency and high-frequency noise from vibration and acoustic emission signals, signal baseline correction is performed, time-domain and frequency-domain features of the signal are extracted through peak detection, and feature dimensionality reduction is performed using Principal Component Analysis (PCA). For the control stream: Instructions and parameters in the PLC registers are parsed, continuous values such as spindle speed and process parameters are normalized, and discrete values such as equipment operating status and recipe number are one-hot encoded to generate standardized control stream features.
[0047] In some embodiments, using the image frame timestamp of the data stream as a reference, the discrete data of the event stream and the control stream are mapped to each acquisition time of the data stream to complete time alignment. This includes: using the image frame timestamp of the data stream as a reference, mapping the discrete data of the event stream and the control stream to the time point of each frame image through a linear interpolation algorithm to generate image samples carrying operating condition context labels to ensure that each frame image corresponds to the synchronized physical state of the equipment and process parameters.
[0048] In this embodiment, the system performs time alignment. That is, the time alignment module uses the timestamp of each frame of the data stream as a reference to perform linear interpolation on the discrete sampled data of the event stream and control stream, mapping all data onto the same time axis, generating image samples with operating condition context labels, ensuring that each frame of the image corresponds to the physical state and process parameters of the equipment at the same time.
[0049] In this embodiment, the system performs spatial alignment, that is, the spatial calibration module completes the conversion between the camera pixel coordinate system and the device world coordinate system through hand-eye calibration, and maps the physical location of the device corresponding to the vibration and acoustic emission sensors to the corresponding pixel area of the image, thereby achieving feature alignment in the spatial dimension.
[0050] In this embodiment, the system can adopt a three-level collaborative triggering hardware-level data latching mechanism, using the trigger zero point as a unified spatiotemporal anchor point to lock the full amount of three streams of data within the trigger window, thus solving the timing misalignment problem of different sampling frequencies from the source of acquisition. Furthermore, this embodiment can also design a two-level synchronization mechanism combining PTP hardware clock synchronization with trigger zero point secondary correction. Based on achieving ±1μs-level clock synchronization across all devices, it performs secondary correction of accumulated clock deviations through trigger zero points, completely resolving the clock drift problem of distributed devices. Moreover, this embodiment proposes a segmented interpolation alignment algorithm combined with industrial prior knowledge, integrating online calibration correction algorithms. This ensures that signal characteristics are not distorted while meeting the real-time requirements of the edge side, and simultaneously achieves online adaptive correction of spatial calibration parameters, completing precise alignment in both time and space. Furthermore, this embodiment can also incorporate robust alignment logic for data verification and packet loss compensation, which can cope with abnormal operating conditions such as electromagnetic interference and data packet loss, ensuring the stability of the system's long-term continuous operation.
[0051] 220. Extract features from the three streams of data to obtain visual features, physical state features, and process prior features; and use the process prior features as the query vector to dynamically calculate the fusion weights of visual features and physical state features based on the attention mechanism to generate deep fusion features.
[0052] In some embodiments, please refer to Figure 3 , Figure 3 This is a schematic diagram of the process for generating deep fusion features according to an embodiment of the present invention. Feature extraction is performed on the three streams of data to obtain visual features, physical state features, and process prior features; using the process prior features as a query vector, the fusion weights of the visual features and physical state features are dynamically calculated based on an attention mechanism to generate deep fusion features, including:
[0053] 300. Extracting visual features from data streams based on convolutional neural networks;
[0054] 310. Extract the physical state features of the event flow and the process prior features of the control flow based on the multilayer perceptron.
[0055] 320. Based on the attention mechanism, the process prior features are used as query vectors, and visual features and physical state features are used as key-value pairs to generate fusion weights. Deep fusion features are generated by weighted summation. The fusion weights are dynamically adjusted by real-time process parameters in the control flow.
[0056] In this embodiment, the visual feature extraction module can use a lightweight ResNet18 convolutional neural network to extract features from the preprocessed ROI image and output 256-dimensional visual features; the multi-source feature extraction module can use an MLP multilayer perceptron to extract features from the preprocessed event flow and control flow data respectively, outputting 128-dimensional physical state features of the event flow and 128-dimensional process prior features of the control flow.
[0057] In this embodiment, the attention mechanism uses process prior features as the query vector and visual features and physical state features as key-value pairs to generate fusion weights. Specifically, the cross-modal attention fusion module can use the control flow feature Fc (i.e., the process prior features of the control flow) as the guiding signal to calculate the dynamic attention weights of the visual feature Fv and the event flow feature Fe (the physical state features of the event flow): First, the control flow feature is input into the fully connected layer, generating three weight coefficients α, β, and γ, whose sum is 1. Then, a weighted sum is used to generate the final 512-dimensional deep fusion feature F, calculated as: F = α·Fv + β·Fe + γ·Fc. In this embodiment, the control flow feature Fc is the query vector (prior knowledge of the industrial scenario), the visual feature Fv and the event flow feature Fe are key-value pairs, and the final output α, β, and γ are the normalized results of the attention weights. For example, when the control flow shows that the spindle speed increases, the system automatically increases the weight β of the event flow feature to verify whether the image interference under high-speed rotation is caused by equipment vibration; when the production line changes, the system automatically adjusts the weight coefficient to adapt to the new product detection logic.
[0058] 230. Input the deep fusion features into the pre-trained decision model to output the detection results; when the detection results meet the preset conditions, send control commands to the industrial control terminal through the industrial communication interface; and feed the detection results back to the attention mechanism to dynamically adjust the fusion weights or decision model parameters to complete adaptive optimization.
[0059] In some embodiments, the decision model is a classification prediction head or a regression prediction head, and the detection results include at least the defect category, quality score, equipment anomaly level, or dimensional measurement results; when the detection results exceed a preset threshold, the industrial control terminal is issued a rejection, alarm, shutdown, or process parameter adjustment command through the industrial communication interface.
[0060] In this embodiment, when the decision model is a classification prediction head, a fully connected network combined with a Softmax activation function can be used to output the defect category; when the decision model is a regression prediction head, a multilayer perceptron can be used to output the size measurement value or quality score. For example, the defect recognition and classification module in this embodiment can input deep fusion features into the pre-trained classification prediction head, which uses a Softmax activation function to output the defect category, defect confidence level, and quality score of the workpiece.
[0061] In this embodiment, the result output module can display the detection results in real time on the human-machine interface and simultaneously send the results to the closed-loop control feedback module. When the detected defect confidence exceeds a preset threshold, or when the equipment status is abnormal, the closed-loop control feedback module immediately sends a control command to the PLC via the Modbus / TCP protocol, triggering the rejection cylinder to remove the unqualified workpiece and simultaneously activating the audible and visual alarm. For serious equipment abnormalities, a shutdown command is directly issued to prevent equipment damage. In this embodiment, the system can also perform adaptive feedback, that is, the system performs correlation analysis with the detection results and control flow and event flow data, continuously optimizing the weight allocation logic of the cross-modal attention module and the detection model parameters, realizing the system's self-learning and adaptive optimization.
[0062] For example, based on the scenario of visual inspection in a precision parts machining production line, this embodiment is applied to a CNC machining production line for precision metal parts in automobiles for real-time detection of surface machining defects. It solves the problems of high false detection rates caused by cutting fluid splashing and the need for manual program configuration during product changeovers in existing technologies. The specific implementation method is as follows:
[0063] Hardware deployment: A CCD industrial camera is deployed at the material output end of the processing station, and vibration sensors and acoustic emission sensors are installed on the machine tool spindle and connected to the Siemens S7-1200 PLC of the machine tool. The edge computing unit is deployed in the production line electrical control cabinet.
[0064] Production line conditions: The production line can process two types of precision parts, corresponding to spindle speeds of 2000 r / min and 3000 r / min respectively. During the processing, there are interferences such as cutting fluid splashing and spindle vibration. The existing single vision inspection system has a false detection rate of 12.3%. Product changeover requires manual shutdown to configure the inspection program, and a single changeover takes about 30 minutes.
[0065] Detailed implementation process:
[0066] Pre-trigger adaptation: When the operator switches the product recipe in the PLC and the spindle speed changes from 2000r / min to 3000r / min, the system receives the control flow pre-trigger signal and automatically adjusts the camera shutter speed from 1 / 500s to 1 / 2000s to eliminate high-speed motion blur. The vibration sensor sampling frequency is increased from 10kHz to 50kHz. At the same time, the detection model and fusion weight configuration corresponding to 3000r / min are preloaded, and no machine stop is required throughout the process.
[0067] Precise trigger acquisition: After the part is processed and enters the inspection station, the photoelectric switch sends an in-position signal, triggering the camera to capture an image of the part surface, while simultaneously locking the current spindle speed, real-time data from the vibration sensor, and acoustic emission signals;
[0068] Interference filtering and fusion detection: Bright spots formed by cutting fluid splashes appear in the camera image, which are initially identified as suspected defects by visual inspection; the system uses a fusion algorithm to combine the high-frequency vibration signal from the vibration sensor (which matches the vibration characteristics at a speed of 3000 r / min) to determine that the bright spots are interference caused by cutting fluid splashes, rather than surface defects of the part; at the same time, it calls the detection model corresponding to the speed to perform precise detection on the surface of the part and outputs the final detection results;
[0069] Closed-loop control: If a real defect is detected in a part, the system immediately sends an instruction to the PLC to trigger the rejection cylinder to reject the defective product, and at the same time records the defect information and the corresponding working condition data.
[0070] Implementation results: In this embodiment, the system's false defect detection rate is reduced to 2.1%, which is 82.9% lower than the existing technology; the product changeover time is shortened from 30 minutes to less than 10 minutes; the system response time is ≤30ms, meeting the real-time detection requirements of the production line.
[0071] For example, based on the scenario of welding quality inspection of new energy power battery tabs, this embodiment is applied to the ultrasonic welding production line of new energy power battery tabs for quality inspection of weld seams and early warning of welding anomalies. The specific implementation method is as follows:
[0072] Hardware deployment: Industrial cameras and infrared thermal imagers are deployed at the welding station, and acoustic emission sensors and pressure sensors are installed on the welding head. They are connected to the PLC of the welding equipment to read control flow parameters such as welding current, welding time, and welding head pressure.
[0073] Specific implementation process: The system pre-triggers the control flow and, upon receiving the welding start command, activates the vision sensor and acoustic emission sensor in advance. After welding is completed, the event flow precisely triggers the camera to capture the weld image, while simultaneously locking the acoustic emission signal and welding process parameters during the welding process. Through a three-flow fusion algorithm, the welding process parameters (control flow), welding acoustic emission characteristics (event flow), and weld image characteristics (data flow) are combined to jointly determine the welding quality. This can effectively identify defects such as incomplete welds, over-welds, and incomplete welds, while also providing early warnings of equipment abnormalities such as weld head wear.
[0074] Implementation results: In this embodiment, the accuracy rate of welding defect detection reaches 99.2%, which is 8.7% higher than the existing technology. It can provide early warning of abnormal wear of welding heads 2 hours in advance, and significantly reduce product defect rate and equipment downtime.
[0075] This specification's embodiments utilize a three-level collaborative triggering mechanism synchronized with a hardware-level microsecond-level clock. Using the trigger moment as a reference, it completes the timeline binding of control flow, event flow, and data flow from the data acquisition source. Furthermore, through spatiotemporal dual feature alignment, it overcomes the shortcomings of spatiotemporal fragmentation in multi-source heterogeneous data. Simultaneously, by verifying the authenticity of visual inspection results through the physical signals of the event flow, it can effectively distinguish between real defects and environmental interference, reducing both the false detection rate and the missed detection rate under complex industrial conditions.
[0076] Furthermore, the embodiments of this specification can also use process prior features as query vectors and dynamically calculate the fusion weights of visual features and physical state features based on an attention mechanism to generate deep fusion features. That is, the embodiments of this specification introduce control flow as a priori guiding signal into the visual perception system, and predict changes in production line processes in advance through a control flow pre-triggering mechanism. This automatically completes the end-to-end adaptive switching of acquisition parameters, detection models, and fusion weights without manual downtime configuration, shortening the downtime of production line product changeovers and adapting to the flexible production needs of multiple varieties and small batches.
[0077] Furthermore, the embodiments in this specification can also adopt a tiered acquisition strategy with hierarchical triggering, which initiates full high-precision acquisition only under abnormal operating conditions and acquires only necessary data under normal operating conditions. This reduces invalid data transmission and computational overhead, significantly reduces the computing load on the edge computing unit, and shortens the system's response time from abnormal detection to the issuance of control commands, thus meeting the needs of real-time closed-loop control in industrial sites.
[0078] Furthermore, the embodiments of this specification can also feed the detection results back to the attention mechanism to dynamically adjust the fusion weights or decision model parameters to achieve adaptive optimization. In other words, the embodiments of this specification construct a closed-loop link of detection, decision-making, control, and feedback. It can not only output accurate detection results based on fusion features, but also feed the results back to the production line control system in real time to realize real-time intervention in anomalies and dynamic optimization of process parameters. This breaks through the limitations of the existing technology where perception and control are disconnected, and realizes the upgrade of industrial vision from passive detection to active control.
[0079] Furthermore, the embodiments of this specification can also set a trigger latching mechanism to completely save all three-stream time-series data before and after the abnormal event is triggered, providing complete contextual data support for the root cause analysis of product quality abnormalities and equipment failures; at the same time, the collected data with spatiotemporal labels and operating condition information can be directly used as high-quality labeled samples for model iteration, greatly reducing the cost of sample labeling and realizing continuous optimization and upgrading of the system.
[0080] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0081] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a three-flow fusion visual inspection system for industrial sites, provided as an embodiment of this specification.
[0082] like Figure 4 As shown, this three-stream fusion visual inspection system for industrial sites can include at least a data acquisition module 400, a spatiotemporal alignment module 410, a multimodal fusion module 420, and a closed-loop detection module 430, among which:
[0083] The data acquisition module 400 is used to synchronize the clocks of vision sensors, auxiliary sensors and industrial control terminals in the industrial field, and to synchronously acquire the data stream output by the vision sensors, the event stream output by the auxiliary sensors and the control stream output by the industrial control terminal based on a three-level collaborative triggering mechanism.
[0084] The spatiotemporal alignment module 410 is used to map the discrete data of the event stream and control stream to the acquisition times of the data stream based on the image frame timestamps of the data stream, thus completing time alignment; establish the coordinate system transformation relationship between the vision sensor and the industrial equipment, and map the physical location information in the event stream to the image pixel coordinates of the data stream, thus completing spatial alignment; and output the three streams of data after time alignment and spatial alignment.
[0085] The multimodal fusion module 420 is used to extract features from the three streams of data to obtain visual features, physical state features, and process prior features; and using the process prior features as the query vector, it dynamically calculates the fusion weights of visual features and physical state features based on an attention mechanism to generate deep fusion features.
[0086] The closed-loop detection module 430 is used to input deep fusion features into a pre-trained decision model to output detection results. When the detection results meet preset conditions, it sends control commands to the industrial control terminal through the industrial communication interface. It also feeds back the detection results to the attention mechanism to dynamically adjust the fusion weights or decision model parameters to complete adaptive optimization.
[0087] In some embodiments, a vision sensor, an auxiliary sensor, and an industrial control terminal are deployed in the industrial field. The industrial control terminal includes at least a PLC controller and an actuator. The data stream includes at least a continuous video frame sequence and static high-resolution image data acquired by the vision sensor. The event stream includes at least equipment physical status data, discrete alarm signals, and status change logs acquired by the auxiliary sensor. The control stream includes at least equipment operating status instructions, process parameter setpoints, and equipment operating feedback signals read from the registers of the industrial control terminal.
[0088] In some embodiments, the three-level collaborative triggering mechanism includes control flow pre-triggering, event flow precise triggering, and data flow self-triggering. Control flow pre-triggering includes using product changeover, process adjustment, and station feed instructions issued by the PLC controller as pre-trigger signals. When the system receives the pre-trigger signal, it wakes up the vision sensor and auxiliary sensor to enter the acquisition state and preloads the acquisition parameters, decision model, and fusion weights for the corresponding working condition. Event flow precise triggering includes using the signal from the auxiliary sensor as the acquisition trigger source, supporting at least one triggering logic among threshold triggering, feature triggering, and switch quantity triggering, and simultaneously starting the vision sensor to acquire images when triggered. Data flow self-triggering includes self-triggering operations through frame difference, ROI area change detection, and suspected defect pre-identification logic. The self-triggering operations include at least automatic acquisition of workpiece in place, environmental interference parameter compensation, and continuous shooting for abnormal verification.
[0089] In some embodiments, the data acquisition module 400 further includes a data latching module, which is used to: determine the trigger time, and latch the full three-stream timing data corresponding to the data stream, event stream and control stream within a preset time window before and after the trigger time, based on the trigger time.
[0090] In some embodiments, the data acquisition module 400 includes a data acquisition submodule, which is used to adopt a hierarchical acquisition strategy, including the following different acquisition modes: Level 1 Regular Inspection Mode: triggered by the control flow station cycle signal, performing low frame rate global image acquisition and low-frequency sampling by auxiliary sensors; Level 2 Precision Detection Mode: triggered by the event flow workpiece arrival signal, performing high-resolution target frame acquisition and latching full working condition data; Level 3 Anomaly Verification Mode: triggered by the data flow suspected defect signal or the event flow anomaly signal, performing high-speed continuous shooting and high-frequency synchronous acquisition by multiple sensors; Level 4 Safety Traceability Mode: triggered by the control flow fault command or the event flow emergency alarm signal, latching full three-stream time-series data within a preset time window before and after the trigger time.
[0091] In some embodiments, the spatiotemporal alignment module 410 includes a time alignment module, which is used to: map the discrete data of the event stream and the control stream to the time point of each frame image respectively using a linear interpolation algorithm based on the image frame timestamp of the data stream, and generate image samples carrying operating condition context labels to ensure that each frame image corresponds to the synchronized equipment physical state and process parameters.
[0092] In some embodiments, the multimodal fusion module 420 includes a fusion submodule, which is used to: extract visual features of the data stream based on a convolutional neural network; extract physical state features of the event stream and process prior features of the control stream based on a multilayer perceptron; generate fusion weights using the process prior features as query vectors and the visual features and physical state features as key-value pairs using an attention mechanism; and generate deep fusion features by weighted summation. The fusion weights are dynamically adjusted by real-time process parameters in the control stream.
[0093] In some embodiments, the decision model is a classification prediction head or a regression prediction head, and the detection results include at least the defect category, quality score, equipment anomaly level, or dimensional measurement results; when the detection results exceed a preset threshold, the industrial control terminal is issued a rejection, alarm, shutdown, or process parameter adjustment command through the industrial communication interface.
[0094] In some embodiments, the three-stream fusion visual inspection system for industrial sites further includes a data preprocessing module, which is used to: denoise, correct distortion, normalize parameters, and extract regions of interest (ROIs) from the data stream; filter, correct baselines, detect peaks, and reduce feature dimensions from the event stream; and parse execution instructions, normalize parameters, and encode states from the control stream.
[0095] Based on the descriptions of the three-stream fusion visual inspection system for industrial sites in several embodiments of this specification, it can be seen that the embodiments of this specification can synchronize with a hardware-level microsecond-level clock through a three-level collaborative triggering mechanism. Using the triggering time as a reference, the time axis binding of control flow, event flow, and data flow is completed from the source of data acquisition. Furthermore, by aligning with dual spatiotemporal features, the spatiotemporal fragmentation of multi-source heterogeneous data is resolved. Simultaneously, the authenticity of the visual inspection results is verified through the physical signals of the event flow, effectively distinguishing between real defects and environmental interference. This reduces both the false detection rate and the missed detection rate under complex industrial conditions.
[0096] Furthermore, the embodiments of this specification can also use process prior features as query vectors and dynamically calculate the fusion weights of visual features and physical state features based on an attention mechanism to generate deep fusion features. That is, the embodiments of this specification introduce control flow as a priori guiding signal into the visual perception system, and predict changes in production line processes in advance through a control flow pre-triggering mechanism. This automatically completes the end-to-end adaptive switching of acquisition parameters, detection models, and fusion weights without manual downtime configuration, shortening the downtime of production line product changeovers and adapting to the flexible production needs of multiple varieties and small batches.
[0097] Furthermore, the embodiments in this specification can also adopt a tiered acquisition strategy with hierarchical triggering, which initiates full high-precision acquisition only under abnormal operating conditions and acquires only necessary data under normal operating conditions. This reduces invalid data transmission and computational overhead, significantly reduces the computing load on the edge computing unit, and shortens the system's response time from abnormal detection to the issuance of control commands, thus meeting the needs of real-time closed-loop control in industrial sites.
[0098] Furthermore, the embodiments of this specification can also feed the detection results back to the attention mechanism to dynamically adjust the fusion weights or decision model parameters to achieve adaptive optimization. In other words, the embodiments of this specification construct a closed-loop link of detection, decision-making, control, and feedback. It can not only output accurate detection results based on fusion features, but also feed the results back to the production line control system in real time to realize real-time intervention in anomalies and dynamic optimization of process parameters. This breaks through the limitations of the existing technology where perception and control are disconnected, and realizes the upgrade of industrial vision from passive detection to active control.
[0099] Furthermore, the embodiments of this specification can also set a trigger latching mechanism to completely save all three-stream time-series data before and after the abnormal event is triggered, providing complete contextual data support for the root cause analysis of product quality abnormalities and equipment failures; at the same time, the collected data with spatiotemporal labels and operating condition information can be directly used as high-quality labeled samples for model iteration, greatly reducing the cost of sample labeling and realizing continuous optimization and upgrading of the system.
[0100] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the embodiment of the three-flow fusion visual inspection system for industrial sites is relatively simple in description because it is fundamentally similar to the embodiment of the three-flow fusion visual inspection method for industrial sites; relevant parts can be referred to the description of the method embodiment.
[0101] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.
[0102] The above embodiments are merely preferred embodiments described in this specification and are not intended to limit the scope of this specification. Any modifications and improvements made by those skilled in the art to the technical solutions of this specification without departing from the spirit of this specification should fall within the protection scope defined by the claims of this specification.
Claims
1. A three-stream fusion visual inspection method for industrial sites, characterized in that, include: The system synchronizes the clocks of the vision sensors, auxiliary sensors, and industrial control terminals in the industrial field, and synchronously collects the data streams output by the vision sensors, the event streams output by the auxiliary sensors, and the control streams output by the industrial control terminals based on a three-level collaborative triggering mechanism. Based on the image frame timestamps of the data stream, the discrete data of the event stream and the control stream are mapped to each acquisition time of the data stream to complete time alignment. Establish a coordinate system transformation relationship between the vision sensor and the industrial equipment, and map the physical location information in the event stream to the image pixel coordinates of the data stream to complete spatial alignment; Output the time-aligned and space-aligned three-stream data; Feature extraction is performed on the three streams of data to obtain visual features, physical state features, and process prior features; Using the prior features of the process as the query vector, the fusion weights of the visual features and the physical state features are dynamically calculated based on the attention mechanism to generate deep fusion features; The deep fusion features are input into a pre-trained decision model to output detection results. When the detection results meet preset conditions, control commands are sent to the industrial control terminal through the industrial communication interface. The detection results are then fed back to the attention mechanism to dynamically adjust the fusion weights or decision model parameters to complete adaptive optimization.
2. The method according to claim 1, characterized in that, The industrial site is equipped with vision sensors, auxiliary sensors, and an industrial control terminal. The industrial control terminal includes at least a PLC controller and an actuator. The data stream includes at least a continuous video frame sequence and static high-resolution image data collected by the vision sensors. The event stream includes at least the equipment physical status data, discrete alarm signals, and status change logs collected by the auxiliary sensors. The control stream includes at least the equipment operating status instructions, process parameter setpoints, and equipment operating feedback signals read from the registers of the industrial control terminal.
3. The method according to claim 1, characterized in that, The three-level collaborative triggering mechanism includes control flow pre-triggering, event flow precise triggering, and data flow self-triggering; The control flow pre-triggering includes using product change, process adjustment, and station feed instructions issued by the PLC controller as pre-trigger signals. When the system receives the pre-trigger signal, it wakes up the vision sensor and the auxiliary sensor to enter the waiting state and preloads the acquisition parameters, decision model, and fusion weights of the corresponding working condition. The precise triggering of the event stream includes using the signal from the auxiliary sensor as the acquisition trigger source, supporting at least one of the triggering logics of threshold triggering, feature triggering, and switch quantity triggering, and simultaneously starting the visual sensor to acquire images when triggered. The data stream self-triggering includes self-triggering operations through frame differential, ROI region change detection, and suspected defect pre-identification logic. The self-triggering operations include at least automatic acquisition of workpiece arrival, environmental interference parameter compensation, and continuous shooting for abnormal verification.
4. The method according to claim 1, characterized in that, The three-level collaborative triggering mechanism, which synchronously collects the data stream output by the vision sensor, the event stream output by the auxiliary sensor, and the control stream output by the industrial control terminal, further includes: determining the trigger time, and using the trigger time as a reference, latching the full three-stream timing data corresponding to the data stream, the event stream, and the control stream within a preset time window before and after the trigger time.
5. The method according to claim 1, characterized in that, The three-level collaborative triggering mechanism synchronously collects the data stream output by the visual sensor, the event stream output by the auxiliary sensor, and the control stream output by the industrial control terminal, including: employing a hierarchical, tiered acquisition strategy, comprising the following different acquisition modes: Level 1 Regular Inspection Mode: Triggered by the control flow station cycle signal, it performs low frame rate global image acquisition and low frequency sampling of auxiliary sensors; Level 2 Precision Detection Mode: Triggered by the workpiece arrival signal in the event flow, high-resolution target frame acquisition is performed, and all working condition data is latched; Level 3 anomaly verification mode: triggered by suspected defect signals in the data stream or anomaly signals in the event stream, it performs high-speed continuous shooting and high-frequency synchronous acquisition by multiple sensors; Level 4 safety traceability mode: triggered by control flow fault command or event flow emergency alarm signal, latching all three-stream time sequence data within a preset time window before and after the trigger time.
6. The method according to claim 1, characterized in that, The step of mapping the discrete data of the event stream and the control stream to each acquisition time of the data stream, based on the image frame timestamps of the data stream, to achieve time alignment, includes: Based on the image frame timestamps of the data stream, the discrete data of the event stream and the control stream are mapped to the time points of each frame image through a linear interpolation algorithm, generating image samples carrying operating context labels to ensure that each frame image corresponds to the synchronized physical state of the equipment and process parameters.
7. The method according to claim 1, characterized in that, The three streams of data are subjected to feature extraction to obtain visual features, physical state features, and process prior features; and using the process prior features as a query vector, the fusion weights of the visual features and the physical state features are dynamically calculated based on an attention mechanism to generate deep fusion features, including: Visual features of the data stream are extracted based on a convolutional neural network; The physical state features of the event flow and the process prior features of the control flow are extracted based on the multilayer perceptron. Based on the attention mechanism, the process prior features are used as query vectors, and the visual features and physical state features are used as key-value pairs to generate fusion weights. Deep fusion features are generated by weighted summation, and the fusion weights are dynamically adjusted by real-time process parameters in the control flow.
8. The method according to claim 1, characterized in that, The decision model is a classification prediction head or a regression prediction head, and the detection results include at least the defect category, quality score, equipment anomaly level, or dimensional measurement results; when the detection results exceed a preset threshold, the industrial control terminal is issued a rejection, alarm, shutdown, or process parameter adjustment command through the industrial communication interface.
9. The method according to claim 1, characterized in that, Before the completion of time alignment, preprocessing is performed on the data stream, event stream, and control stream respectively; the preprocessing includes: The data stream is subjected to denoising, distortion correction, parameter normalization, and region of interest (ROI) extraction. The event stream is then filtered, baseline corrected, peak detected, and feature dimension reduced. The control flow is subjected to instruction parsing, parameter normalization, and state encoding.
10. A three-flow fusion visual inspection system for industrial sites, characterized in that, include: The data acquisition module is used to synchronize the clocks of the vision sensors, auxiliary sensors, and industrial control terminals in the industrial field, and based on a three-level collaborative triggering mechanism, synchronously acquire the data streams output by the vision sensors, the event streams output by the auxiliary sensors, and the control streams output by the industrial control terminals. The spatiotemporal alignment module is used to map the discrete data of the event stream and the control stream to each acquisition time of the data stream, based on the image frame timestamps of the data stream, to complete the time alignment. Establish a coordinate system transformation relationship between the vision sensor and the industrial equipment, and map the physical location information in the event stream to the image pixel coordinates of the data stream to complete spatial alignment; Output the time-aligned and space-aligned three-stream data; The multimodal fusion module is used to extract features from the three streams of data to obtain visual features, physical state features, and process prior features. Using the prior features of the process as the query vector, the fusion weights of the visual features and the physical state features are dynamically calculated based on the attention mechanism to generate deep fusion features; The closed-loop detection module is used to input the deep fusion features into a pre-trained decision model to output detection results; when the detection results meet preset conditions, it sends control commands to the industrial control terminal through the industrial communication interface; and feeds back the detection results to the attention mechanism to dynamically adjust the fusion weights or decision model parameters to complete adaptive optimization.