Lightweight industrial hazard detection method based on time sequence feature and causal decoupling
Patent Information
- Application Number
- CN202610906957.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-01
AI Technical Summary
[0004]然而,上述现有方法在面向冶金、化工等典型高危工业场景进行实时隐患检测与预警时,仍存在以下显著技术瓶颈:
1、通过构建嵌入多尺度卷积与通道注意力的局部特征增强模块,显著提升了弱目标隐患的感知精度;同时引入物理引导的因果解耦机制与可微反事实干预,将图像特征分解为真实隐患相关特征、环境干扰相关特征与背景残差特征,并强制真实隐患相关特征与环境干扰相关特征正交,从而大幅增强抗光照突变、粉尘遮挡等环境干扰的能力;在此基础上,利用光流法对齐连续帧特征并借助门控循环单元进行时序关联建模,通过帧间相似度比较有效区分持续性真实隐患与瞬时性环境干扰,实现时序因果判别;模型还能直接输出工业隐患类别及干扰判别理由,赋予工业隐患检测结果完整的可解释性;通过精简编码器层数、通道剪枝、稀疏注意力与模型量化等轻量化重构措施,使模型参数量与计算量显著降低,满足端侧实时部署要求。
Smart Images

Figure CN122676253A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial safety testing technology, and in particular to a lightweight ViT industrial hazard detection method that decouples temporal characteristics from causal factors. Background Technology
[0002] Intelligent detection of industrial hazards is a key technology for ensuring safe production in high-risk environments (such as metallurgy, chemical industry, high-temperature molten metal, and flammable and explosive equipment). With the development of computer vision, deep learning-based methods, especially the Vision Transformer (ViT), have shown great potential in the field of industrial defect detection due to their superior global feature modeling capabilities. ViT and its variants, through a self-attention mechanism, can effectively capture long-range dependencies in images, achieving performance superior to traditional convolutional neural networks (CNNs) in tasks such as molten metal leak detection and equipment surface defect identification.
[0003] Existing research has explored various aspects of the application of Transformer in industrial inspection. For example, to alleviate the dependence of Transformer models on large-scale labeled data and the high computational resource requirements, lightweight models such as TRSBi-YOLO have achieved a balance between parameter quantity and accuracy in printed circuit board (PCB) defect detection by introducing C3TR modules and the SimAM parameterless attention mechanism. For multi-scale feature extraction, some studies have employed parallel multi-size convolutional kernels or switchable dilated convolutions to expand the receptive field, thereby enhancing the ability to capture features of small targets. Regarding the utilization of temporal information, recurrent neural networks such as gated recurrent units (GRUs) and long short-term memory networks (LSTMs) have been used to model temporal dependencies between consecutive frames. Some works have also attempted to eliminate motion effects by using optical flow methods for inter-frame feature alignment.
[0004] However, the existing methods described above still have the following significant technical bottlenecks when used for real-time hazard detection and early warning in typical high-risk industrial scenarios such as metallurgy and chemical industry: First, the ability to suppress environmental interference is limited. Existing methods mostly employ fixed-weight feature fusion strategies or static filtering templates, making it difficult to dynamically adapt to real industrial environments with drastic changes such as sudden changes in lighting, dust obstruction, and uneven reflection. When the interference intensity exceeds a preset threshold, the model's detection performance drops sharply, leading to a large number of false positives or false negatives.
[0005] Second, the model lacks sufficient perception of weak targets and small-scale hazards. While the traditional ViT global self-attention mechanism captures the global context of the image, its ability to represent local details, especially the features of weak targets such as micro-cracks and early leaks, is relatively weak. This makes it difficult for the model to accurately locate small-sized hazard areas, easily leading to missed detections of key risk points.
[0006] Third, insufficient mining of temporal information makes it difficult to distinguish between real hidden dangers and transient interference. Most existing methods still rely on single-frame images for independent detection, failing to effectively utilize the rich continuous frame temporal information in industrial scenarios. Therefore, the model struggles to accurately distinguish between the continuous development process of hidden dangers (such as leakage expansion) and transient, non-causally related environmental interferences such as flying insects and flickering light and shadow, resulting in a high false alarm rate and affecting the reliability of the system.
[0007] Fourth, the model lacks interpretability and cannot identify the source of interference. Current mainstream methods are essentially "black box" models, only outputting detection results (category, location, confidence level). When false detections occur, engineers cannot determine whether the error is caused by dust, changes in lighting, or other specific interference factors. This lack of interpretability makes it difficult for the system to gain sufficient trust in high-risk industrial scenarios and also increases the difficulty of subsequent fault localization and model optimization.
[0008] Fifth, the high model complexity makes it difficult to meet the real-time deployment requirements of edge devices. Existing Transformer models have a large number of parameters and high computational complexity, placing stringent demands on hardware resources. This sharply contradicts the limited computing power and power sensitivity of edge devices such as robots and smart cameras, limiting their large-scale real-time deployment in real-world industrial inspection scenarios.
[0009] Therefore, how to provide a lightweight ViT industrial hazard detection method that decouples temporal features from causality, thereby improving the anti-environmental interference capability, weak target perception accuracy, temporal causal discrimination capability between real hazards and instantaneous interference, model interpretability, and real-time deployment efficiency on the edge side, has become an urgent technical problem to be solved. Summary of the Invention
[0010] The technical problem to be solved by this invention is to provide a lightweight ViT industrial hazard detection method that decouples temporal features from causality, thereby improving the anti-environmental interference capability, weak target perception accuracy, temporal causal discrimination capability between real hazards and instantaneous interference, model interpretability, and real-time deployment efficiency at the edge.
[0011] This invention is implemented as follows: a lightweight ViT industrial hazard detection method decoupled from temporal characteristics and causality, comprising the following steps: Step S1: Obtain the inspection image sequence of the industrial scene, and preprocess the inspection image sequence to obtain standardized images; Step S2: Construct a Vision Transformer network and pre-train the Vision Transformer network, introducing causal consistency training loss during the pre-training process; Step S3: Improve the pre-trained Vision Transformer network by embedding a local feature enhancement module into the encoder of the Vision Transformer network and perform lightweight reconstruction on the improved Vision Transformer network. Step S4: Using the lightweight reconstructed Vision Transformer network, extract features from the standardized image and output an enhanced local feature map; Step S5: Introduce a physically guided hazard-interference causal decoupling mechanism to perform causal decoupling on the local feature map. The causal decoupling includes: decomposing the local feature map into real hazard-related features, environmental interference-related features, and background residual features, and then performing differentiable counterfactual intervention: simulating the removal of environmental interference based on the real hazard-related features to generate counterfactual features, comparing the similarity between the real hazard-related features and the counterfactual features, determining whether it is a real hazard or environmental interference based on the comparison result, and outputting the reason for interference discrimination. Step S6: Introduce a cross-frame temporal feature fusion mechanism to perform temporal alignment and correlation modeling on the local feature maps of multiple consecutive frames, so as to distinguish between the development of real hidden dangers and instantaneous environmental interference, and output temporal fusion features; Step S7: Input the standardized image and time-series fusion features into the pre-trained Vision Transformer network that has undergone lightweight reconstruction, and output the industrial hazard detection result in conjunction with the interference discrimination reason.
[0012] Furthermore, in step S1, the preprocessing specifically includes: The inspection image sequence is noise-removed using Gaussian filtering. The Gaussian filtering function is: ; in, Indicates the Gaussian filter kernel in coordinates The weight value at the location; σ represents the position coordinates of a pixel within the Gaussian filter kernel relative to the kernel center; σ represents the standard deviation of the Gaussian distribution, used to control the width and smoothness of the Gaussian filter kernel. This represents the normalization coefficient, used to ensure that the sum of all weights of the Gaussian filter kernel is 1; The Gaussian-filtered inspection image sequence is normalized using a linear stretching algorithm. The normalization formula is as follows: ; in, This represents the normalized image in pixel coordinates. The grayscale value at that location is mapped to the range of 0~255; This indicates the pixel coordinates of the inspection image sequence after Gaussian filtering. The grayscale value at that location; This represents the minimum gray value of all pixels in the Gaussian-filtered inspection image sequence; 255 represents the maximum gray value of all pixels in the Gaussian filtered inspection image sequence; 255 represents the target upper limit for gray-level normalization. The grayscale-normalized inspection image sequence is scaled proportionally to a preset size to obtain a standardized image. The scaling process uses bilinear interpolation, and the interpolation formula is as follows: ; in, This indicates the normalized image obtained by scaling in coordinates. Pixel value at; This represents the pixel coordinates in the normalized image; m and n both represent summation indices, where m=0,1 represents two adjacent integer coordinates in the horizontal direction, and n=0,1 represents two adjacent integer coordinates in the vertical direction. This indicates that in the grayscale normalized inspection image sequence, the image is in line with... Integer coordinates of four adjacent pixels; This indicates that in the grayscale-normalized inspection image sequence, in the coordinates... The pixel value at that location.
[0013] Furthermore, in step S2, the pre-trained Vision Transformer network is obtained through the following process: Construct a dedicated dataset of potential hazards in high-risk industrial scenarios, covering various types of industrial hazards and different on-site working conditions; The dataset specifically designed for potential hazards in high-risk industrial scenarios is divided into a training set, a validation set, and a test set, and data augmentation is performed on the training set. The VisionTransformer network is iteratively trained using the weighted sum of cross-entropy loss and position regression loss as the total loss function. Training of the Vision Transformer network stops when the average detection accuracy on the validation set stabilizes, and the performance of the trained Vision Transformer network is validated using the test set.
[0014] Furthermore, in step S2, the causal consistency training loss is used to constrain the VisionTransformer network to learn real-world hazard-related features that are independent of the learning environment. The causal consistency training loss is a joint loss function that includes task loss, feature decoupling loss, and counterfactual consistency loss. The feature decoupling loss is achieved through mutual information minimization constraints, which force the real hidden danger-related features and environmental interference-related features to be orthogonal in the feature space, so that the information carried by the real hidden danger-related features and environmental interference-related features does not overlap.
[0015] Furthermore, in step S2, the criteria for determining whether training is successful during the pre-training process are as follows: Average detection accuracy ≥96%, small target hazard detection accuracy ≥90%, edge-side inference speed ≥30 frames / second.
[0016] Furthermore, in step S3, the workflow of the local feature enhancement module is as follows: Multi-scale local receptive field units are constructed, and convolutional kernels of three sizes (3×3, 5×5, and 7×7) are connected in parallel. Parallel convolution operations are performed on the input standardized image to capture local features at the three scales respectively. The convolution operation formula is as follows: ; in, This indicates that when the kernel size is k, the output local features are located at... The pixel value at that location; k represents the size of the convolution kernel, which can be 3, 5, or 7. Represents the relative coordinates within the convolution kernel, with values ranging from 0 to k−1, used to traverse each element of the convolution kernel; This represents the relative positions of a k×k convolution kernel. Weight parameters at the location; Indicates the normalized image at location Pixel value at that location, It is the top-left anchor point of the standardized image corresponding to the output position; The local features at the three scales are normalized and then channel-level stitched together to obtain a fused feature map. The fused feature map is input into a feature enhancement convolutional layer with an embedded channel attention mechanism. The importance weights of each feature channel are calculated through the channel attention mechanism, and the importance weights are multiplied by the fused feature map to output the enhanced local feature map.
[0017] Furthermore, in step S3, the lightweight reconstruction specifically includes: The number of encoder layers in the improved Vision Transformer network is reduced to 6-8 layers; A channel pruning algorithm is used to compress the channel dimensions of each layer in the improved Vision Transformer network, and effective channels are selected based on channel importance scores; the formula for calculating the channel importance score is as follows: ; in, This represents the importance score of the c-th channel; c represents the channel index. Indicates the row index within the convolution kernel; This represents the column index within the convolution kernel; N represents the height of the convolution kernel; M represents the width of the convolution kernel; This indicates that the convolution kernel of the c-th channel is located at position... Weight parameters at the location; Sparse attention computation is adopted to replace the multi-head self-attention mechanism, and attention weights are calculated only for candidate regions of potential hazards. The sparse attention computation formula is as follows: ; in, This represents the result of sparse attention calculation; Represents the query matrix; Represents the key matrix; Represents a value matrix; Represents the transpose of the key matrix; Indicates the dimension of the key matrix; Indicates the scaling factor; This represents the activation function, used to convert the attention score into a probability distribution; The spatial mask matrix represents a value of 1 for candidate regions of potential hazards and a value of 0 for background regions. Model quantization converts 32-bit floating-point parameters to 16-bit floating-point parameters.
[0018] Furthermore, in step S5, the differentiable counterfactual intervention further includes: performing at least one of dust intervention, light intervention, and motion intervention on the candidate hazard area; The dust intervention involves forcibly assuming that the dust concentration is zero, and then re-extracting the real hazard-related features in the decoupled feature space to obtain the first counterfactual hazard features. The lighting intervention involves forcibly assuming that the lighting intensity is a standard reference value, correcting the image lighting to ideal conditions, and then re-propagating it forward to extract real hazard-related features and obtain the second counterfactual hazard features. The motion intervention involves forcibly assuming zero motion interference, eliminating the effects of timing jitter and instantaneous motion, and then re-extracting the relevant features of the real hidden danger to obtain the third counterfactual hidden danger features. The similarity of the real hidden danger-related features with the first counterfactual hidden danger features, the second counterfactual hidden danger features, and the third counterfactual hidden danger features is compared, and the real hidden danger or environmental interference is determined based on the comparison results.
[0019] Furthermore, in step S6, the cross-frame temporal feature fusion mechanism specifically includes: Feature extraction is performed on the local feature maps of 3 to 8 consecutive frames to obtain the high-level local features corresponding to each frame; Optical flow is used to perform temporal alignment of the high-level local features in consecutive frames to eliminate feature position offsets caused by device movement; A gated loop unit is constructed to perform correlation modeling on the time-aligned high-level local features and extract the time-fusion features of the potential hazard targets; The similarity of the temporal fusion features of consecutive frames is calculated. When the similarity is higher than a preset threshold, it is determined to be a real hidden danger development. When the similarity is lower than the preset threshold and only a single frame shows feature abnormality, it is determined to be instantaneous environmental interference.
[0020] Furthermore, in step S7, the industrial hazard detection result carries the industrial hazard category, industrial hazard location, hazard confidence level, instantaneous interference filtering mark, and interference discrimination reason.
[0021] The advantages of this invention are: 1. By constructing a local feature enhancement module embedded with multi-scale convolution and channel attention, the perception accuracy of weak target hazards is significantly improved. At the same time, a physically guided causal decoupling mechanism and differentiable counterfactual intervention are introduced to decompose image features into features related to real hazards, features related to environmental interference, and background residual features. The features related to real hazards and features related to environmental interference are forced to be orthogonal, thereby greatly enhancing the ability to resist environmental interference such as sudden changes in illumination and dust occlusion. On this basis, the optical flow method is used to align continuous frame features and a gated recurrent unit is used to perform temporal correlation modeling. By comparing the similarity between frames, the model can effectively distinguish between persistent real hazards and transient environmental interference, and realize temporal causal discrimination. The model can also directly output the industrial hazard category and the reason for interference discrimination, giving the industrial hazard detection results complete interpretability. Through lightweight reconstruction measures such as reducing the number of encoder layers, channel pruning, sparse attention, and model quantization, the number of model parameters and computational cost are significantly reduced, meeting the requirements of real-time deployment on the edge.
[0022] 2. Integrating local enhancement and global modeling to improve feature representation accuracy: A local feature enhancement module is embedded in the Vision Transformer (ViT) encoder. Local features are extracted by parallel 3×3, 5×5, and 7×7 multi-scale convolutional kernels, and feature enhancement is performed by combining channel attention mechanism. Compared with the traditional ViT which only relies on global self-attention, it can capture subtle local textures and wide-area contextual information in industrial scenarios at the same time. It effectively solves the problem of variable target scale and easy loss of details in hidden dangers, and significantly improves the detection capability of small hidden dangers (such as cracks and leaks).
[0023] 3. Introducing physically guided causal decoupling and counterfactual intervention to enhance anti-interference and interpretability: By decomposing image features into three parts—features related to real hazards, features related to environmental interference, and background residual features—and performing differentiable counterfactual interventions (such as dust, lighting, and motion interventions), it can actively simulate the ideal features after "removing interference." By comparing the similarity with the original features, the authenticity of the hazard can be determined. This mechanism not only significantly reduces false alarms caused by dust, lighting changes, and equipment vibration in complex industrial environments, but also outputs reasons for interference judgment (such as "current high dust levels lead to false detection"), making the detection results interpretable and facilitating on-site personnel to trace and verify them.
[0024] 4. Lightweight Reconstruction Design to Meet Real-Time Deployment Needs on the Edge: To address the limited computing power of edge computing devices in industrial inspections, the ViT network has undergone system-level lightweighting: the number of encoder layers has been reduced to 6-8, a channel pruning algorithm has been used to compress the channel dimension, global multi-head self-attention has been replaced with sparse attention (only calculating candidate regions for potential hazards), and the parameters have been quantized from 32-bit floating-point to 16-bit floating-point. These measures significantly reduce the number of model parameters and computational overhead while maintaining high accuracy, enabling inference speeds of over 30 frames per second, making it suitable for real-time detection scenarios on mobile devices such as drones and robots.
[0025] 5. Cross-frame temporal feature fusion effectively distinguishes between transient interference and persistent hazards: By temporal alignment (optical flow method) and gated recurrent unit (GRU) modeling of 3 to 8 consecutive frames of images, the dynamic laws of hazard evolution over time can be explored; when an abnormal feature appears in a frame but the temporal similarity drops sharply, it is judged as transient environmental interference (such as flying insects, reflections); if the features of multiple frames change continuously and the similarity is stable, it is judged as the development of a real hazard; this cleverly utilizes the characteristics that hazards in industrial scenarios are usually gradual while interference is mostly transient, further filtering out false detections in a single frame and improving detection reliability.
[0026] 6. Comprehensive and quantifiable pre-training strategies ensure model generalization and robustness: The pre-training process constructs a dedicated dataset covering various hazard types and on-site working conditions, and employs joint optimization using cross-entropy loss, location regression loss, and causal consistency training loss (including feature decoupling loss and counterfactual consistency loss); among them, minimizing mutual information forces the orthogonality of real hazard-related features and environmental interference-related features, prompting the model to learn hazard representations that are independent of the environment; the training qualification standards are clearly set as an average detection accuracy of ≥96%, small target hazard accuracy of ≥90%, and edge frame rate of ≥30 frames / second, providing clear and verifiable performance thresholds for product deployment.
[0027] 7. End-to-end output of multi-dimensional detection results to assist operation and maintenance decisions: The final output of industrial hazard detection results not only includes the industrial hazard category, industrial hazard location, and hazard confidence level, but also carries instantaneous interference filtering markers and specific interference discrimination reasons (such as "motion intervention leads to inconsistent features"). This information-rich output format facilitates the backend system to automatically filter false alarms, while providing a basis for human-machine collaborative review, significantly improving the efficiency and reliability of industrial site hazard handling. Attached Figure Description
[0028] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0029] Figure 1 This is a flowchart of a lightweight ViT industrial hazard detection method based on the decoupling of temporal characteristics and causality, according to the present invention. Detailed Implementation
[0030] The technical solution in this application embodiment follows the following general approach: By embedding a local feature enhancement module consisting of multi-scale convolution and channel attention into the encoder of the Vision Transformer network, the perception accuracy of minor hazards is improved; a physically guided causal decoupling mechanism is introduced to decompose image features into three parts: hazard, interference, and background, and differentiable counterfactual intervention is performed to simulate the ideal features after removing interference such as dust and illumination. The authenticity of the hazard is determined by similarity comparison, and the reason for interference discrimination is output, thereby enhancing the anti-interference capability and interpretability; at the same time, a lightweight strategy of reducing the number of encoder layers, channel pruning, sparse attention, and model quantization is adopted to significantly reduce computational overhead to meet the real-time inference requirements of the edge side; combined with a cross-frame temporal feature fusion mechanism, the optical flow method is used to align continuous frame features and the temporal correlation is modeled with gated recurrent units to effectively distinguish between persistent real hazards and transient environmental interference; finally, through a joint training strategy including causal consistency loss, the model is forced to learn environment-independent hazard representations, and finally, an interpretable detection result carrying the industrial hazard category, industrial hazard location, hazard confidence, transient interference filtering label, and interference discrimination reason is output.
[0031] Please refer to Figure 1 As shown, a preferred embodiment of the lightweight ViT industrial hazard detection method of the present invention, which decouples temporal characteristics and causality, includes the following steps: Step S1: Obtain the inspection image sequence of the industrial scene, and preprocess the inspection image sequence to obtain standardized images; Industrial scenarios include metallurgical and chemical industrial parks, high-temperature molten metal operation areas, flammable and explosive equipment inspection areas, and high-pressure pipeline inspection areas; the inspection image sequence is acquired by a high-definition vision sensor carried by the inspection robot, with a frame rate of no less than 30 frames / second to meet the needs of real-time detection; Step S2: Construct an improved Vision Transformer network (improved ViT model), embed a local feature enhancement module in the encoder of the improved Vision Transformer network, and perform lightweight reconstruction of the improved Vision Transformer network. Step S3: Using the pre-trained and improved Vision Transformer network, extract features from the standardized image and output an enhanced local feature map. Step S4: Introduce a physically guided hazard-interference causal decoupling mechanism to perform causal decoupling on the local feature map. The causal decoupling includes: decomposing the local feature map into real hazard-related features, environmental interference-related features, and background residual features, and then performing differentiable counterfactual intervention: simulating the removal of environmental interference based on the real hazard-related features to generate counterfactual features, comparing the similarity between the real hazard-related features and the counterfactual features, determining whether it is a real hazard or environmental interference based on the comparison result, and outputting the reason for interference discrimination. Step S5: Introduce a cross-frame temporal feature fusion mechanism to perform temporal alignment and correlation modeling on the local feature maps of multiple consecutive frames in order to distinguish between the development of real hidden dangers and instantaneous environmental interference, and output temporal fusion features; Step S6: Input the standardized image and time-series fusion features into the pre-trained Vision Transformer network that has undergone lightweight reconstruction, and output the industrial hazard detection result in conjunction with the interference discrimination reason.
[0032] In step S1, the preprocessing specifically includes: The inspection image sequence is noise-removed using Gaussian filtering. The Gaussian filtering function is: ; in, Indicates the Gaussian filter kernel in coordinates The weight value at the location; σ represents the position coordinates of the pixels within the Gaussian filter kernel relative to the kernel center; σ represents the standard deviation of the Gaussian distribution, used to control the width and smoothness of the Gaussian filter kernel. Its value range is dynamically adjusted according to the noise intensity of the industrial site, and is usually set to 0.8 to 1.2. When the dust in the site is heavy and the image noise increases, the σ value can be appropriately increased to enhance the smoothing effect; conversely, in a relatively clean environment, a smaller σ value is taken to retain more hidden details. This represents the normalization coefficient, used to ensure that the sum of all weights of the Gaussian filter kernel is 1, that is, to ensure that the overall brightness of the image remains unchanged before and after filtering; through this filtering, while removing interference, key features such as the contours and texture abrupt changes of potential problems are preserved, thus avoiding false filtering of features; The Gaussian-filtered inspection image sequence is normalized using a linear stretching algorithm to eliminate brightness differences caused by abrupt changes in illumination and uneven reflection. The normalization formula is as follows: ; in, This represents the normalized image in pixel coordinates. The grayscale value at that location is mapped to the range of 0~255; This indicates the pixel coordinates of the inspection image sequence after Gaussian filtering. The grayscale value at that location; This represents the minimum gray value of all pixels in the Gaussian-filtered inspection image sequence; 255 represents the maximum gray value of all pixels in the Gaussian filtered inspection image sequence; 255 represents the target upper limit for gray-level normalization. The grayscale-normalized inspection image sequence is scaled proportionally to a preset size (e.g., 640×480, set according to the inspection robot's sensor parameters and model inference requirements) to obtain a standardized image. The scaling process uses bilinear interpolation, and the interpolation formula is as follows: ; in, This indicates the normalized image obtained by scaling in coordinates. Pixel value at; This represents the pixel coordinates in the normalized image; m and n both represent summation indices, where m=0,1 represents two adjacent integer coordinates in the horizontal direction, and n=0,1 represents two adjacent integer coordinates in the vertical direction. This indicates that in the grayscale normalized inspection image sequence, the image is in line with... The integer coordinates of four adjacent pixels are used to ensure the feature integrity of potential targets such as molten metal leakage and tank corrosion, without stretching distortion; This indicates that in the grayscale-normalized inspection image sequence, in the coordinates... The pixel value at that location.
[0033] Step S1 aims to eliminate image quality inconsistencies caused by factors such as sensor noise and drastic changes in lighting in industrial settings, providing stable and consistent input for subsequent models. For any frame of the inspection image sequence, perform the following operations: First, a two-dimensional Gaussian filter is used for noise removal. Considering that dust and sensor vibration, which are common in industrial environments, can cause high-frequency noise, while real hidden dangers such as cracks and leak edges are usually manifested as key contours and texture information, Gaussian filtering can smooth out noise while preserving these structural features to the maximum extent.
[0034] Subsequently, a linear stretching algorithm was used to normalize the grayscale of the denoised image to eliminate the differences in image brightness caused by sudden changes in illumination and uneven reflection. This operation forcibly linearly maps pixel values from the current distribution range to the standard range [0, 255], so that similar potential hazards collected at different times and under different illumination conditions have consistent grayscale statistical characteristics, thereby improving the stability of feature extraction.
[0035] Finally, the normalized image is scaled proportionally to a preset size, such as 640×480 pixels, to match the input requirements of the improved ViT network. This preset size needs to balance the model's inference speed with the resolution requirements of small target hazards. The scaling operation uses bilinear interpolation, which determines the new pixel value by calculating the weighted average of the four neighboring pixels around the target pixel. This effectively avoids the jagged edges and mosaic effects caused by nearest-neighbor interpolation, ensuring the smoothness of the contours and the integrity of features of potential hazards such as tank corrosion patches and tiny leaks of molten metal, and avoiding stretching distortion.
[0036] In step S2, the pre-trained Vision Transformer network is obtained through the following process: Construct a dedicated dataset of potential hazards in high-risk industrial scenarios, covering various types of industrial hazards and different on-site working conditions; The dataset specifically designed for potential hazards in high-risk industrial scenarios is divided into a training set, a validation set, and a test set, and data augmentation is performed on the training set. The VisionTransformer network is iteratively trained using the weighted sum of cross-entropy loss and position regression loss as the total loss function. Training of the Vision Transformer network stops when the average detection accuracy on the validation set stabilizes, and the performance of the trained Vision Transformer network is validated using the test set.
[0037] In step S2, the causal consistency training loss is used to constrain the Vision Transformer network to learn real hazard-related features that are independent of the environment. The causal consistency training loss is a joint loss function that includes task loss, feature decoupling loss, and counterfactual consistency loss. The feature decoupling loss is achieved through mutual information minimization constraints, which force the real hidden danger-related features and environmental interference-related features to be orthogonal in the feature space, so that the information carried by the real hidden danger-related features and environmental interference-related features does not overlap.
[0038] In step S2, during the pre-training process, the criteria for determining whether the training is successful are as follows: Average detection accuracy ≥96%, small target hazard detection accuracy ≥90%, edge-side inference speed ≥30 frames / second.
[0039] To ensure the improved ViT model meets the stringent performance requirements of high-risk industrial scenarios, a comprehensive and quantifiable pre-training strategy was designed. First, a dedicated, high-coverage dataset of potential hazards in high-risk industrial scenarios was constructed. This dataset must include various hazard types such as molten metal leaks, tank corrosion, pipe cracks, equipment deformation, and valve loosening, and cover diverse and complex real-world conditions including varying lighting, dust obstruction, significant differences in target scale, and complex backgrounds, ensuring the diversity and representativeness of the training data. The dataset was then divided into training, validation, and test sets in an 8:1:1 ratio. Online data augmentation techniques, such as random flipping, rotation, brightness and contrast adjustment, and noise addition, were applied to the training set to simulate more realistic conditions and improve the model's generalization ability.
[0040] Training is supervised using a joint loss function, which includes not only task-oriented losses for detection accuracy (such as a weighted sum of cross-entropy classification loss and location regression loss), but also creatively introduces a causal consistency training loss. This loss consists of feature decoupling loss and counterfactual consistency loss. Feature decoupling loss forces the two to be orthogonal in the feature space by minimizing the mutual information between features related to real hazards and features related to environmental interference, requiring the network to learn robust representations of the hazards themselves that are independent of environmental factors (light, dust, etc.). Counterfactual consistency loss penalizes situations where the predicted results after counterfactual intervention are inconsistent with the expected results (the correct results that should be obtained without interference), further strengthening the model's robust reasoning ability from a causal perspective. The training process uses stochastic gradient descent (SGD) or its variant optimizer, with learning rate warm-up and cosine annealing decay strategies. Training is stopped when the average detection accuracy (mAP) of the improved ViT model on the validation set tends to stabilize and shows no significant improvement over several consecutive epochs to avoid overfitting. Ultimately, a successfully trained improved ViT model must simultaneously meet the following quantifiable performance thresholds on the test set: average detection accuracy ≥ 96%, small target hazard detection accuracy ≥ 90%, and edge inference speed ≥ 30 frames / second. This rigorous process and clear acceptance criteria provide a solid guarantee for the industrial-grade productization of the technical solution.
[0041] In step S3, the workflow of the local feature enhancement module is as follows: Multi-scale local receptive field units are constructed, and convolutional kernels of three sizes (3×3, 5×5, and 7×7) are connected in parallel. Parallel convolution operations are performed on the input standardized image to capture local features at the three scales respectively. The convolution operation formula is as follows: ; in, This indicates that when the kernel size is k, the output local features are located at... The pixel value at that location; k represents the size of the convolution kernel, which can be 3, 5, or 7. Represents the relative coordinates within the convolution kernel, with values ranging from 0 to k−1, used to traverse each element of the convolution kernel; This represents the relative positions of a k×k convolution kernel. Weight parameters at the location; Indicates the normalized image at location Pixel value at that location, It is the top-left anchor point of the standardized image corresponding to the output position; The local features at the three scales are normalized to unify the numerical scale, avoid weight imbalance, and then channel-level concatenation is performed to obtain a fused feature map, fully preserving multi-dimensional local details; the normalization formula is: ; in, This represents the normalized feature map obtained after layer normalization operations; Let these represent the mean and variance of the local features, respectively. , indicating the prevention of minute values where the denominator is 0; γ and β both represent learnable parameters; The fused feature map is input into a feature enhancement convolutional layer with an embedded channel attention mechanism. The importance weights of each feature channel are calculated through the channel attention mechanism, and the importance weights are multiplied by the fused feature map to output the enhanced local feature map.
[0042] Step S3 is the fundamental architectural innovation of this invention, aiming to solve the problem of insufficient local detail capture capability and high missed detection rate of minor issues caused by the reliance on global self-attention in traditional ViT. The core is to embed a local feature enhancement module after the PatchEmbedding layer and before the multi-head attention layer of the ViT encoder. The local feature enhancement module consists of three parts: First, multi-scale local receptive field units are constructed. Unlike conventional single-scale convolutions, this invention uses three different sizes of depthwise separable convolutional kernels (3×3, 5×5, and 7×7) in parallel to perform parallel computation on the input feature map F_in (normalized image). The small 3×3 convolutional kernel has a small receptive field, focusing on capturing fine-grained local features such as the edges of molten metal leaks and minute cracks on the tank surface; the medium-sized 5×5 convolutional kernel is responsible for extracting medium-contour information of the leak area; and the large 7×7 convolutional kernel provides coarser-grained overall regional features. Through this multi-scale parallel extraction mechanism, comprehensive, cross-scale hazard feature capture is achieved, ranging from minute defects to large-scale anomalies.
[0043] Secondly, multi-scale feature normalization and concatenation fusion are performed. Since the numerical distribution ranges of feature maps output by convolutional kernels of different sizes differ, direct fusion can easily lead to weight imbalance. This step performs LayerNorm normalization on the output feature maps of the three scales sequentially, unifying them to similar numerical scales. Then, the normalized feature maps are concatenated along the channel dimension to form a highly information-rich fused feature map containing multi-dimensional local details. .
[0044] Finally, a feature enhancement convolutional layer is applied for weighted amplification. The fused feature maps are then... Input a feature-enhanced convolutional layer with a channel-embedded attention mechanism (CBAM). The core of the feature-enhanced convolutional layer is adaptively learning the importance weights of each feature channel: ; in, σ represents the channel attention weight, i.e., the importance weight; σ represents the Sigmoid activation function; MLP represents a fully connected network. Indicates global average pooling; This indicates global max pooling.
[0045] The channel attention module learns to automatically identify key channels corresponding to small targets and weak features (such as microcracks) and assigns them high weights; simultaneously, it assigns low weights to channels representing background interference or meaningless textures. Then, based on these weights, it... The various channels are weighted, amplified, and suppressed to ultimately output an enhanced local feature map. This process significantly strengthens the expression of the core features of potential hazards and effectively suppresses background noise, fundamentally solving the technical pain point of traditional ViT's weak ability to perceive small-sized hazards.
[0046] In step S3, the lightweight reconstruction specifically includes: The number of encoder layers in the improved Vision Transformer network is reduced to 6-8 layers; A channel pruning algorithm is used to compress the channel dimensions of each layer in the improved Vision Transformer network, and effective channels are selected based on channel importance scores; the formula for calculating the channel importance score is as follows: ; in, This represents the importance score of the c-th channel; c represents the channel index. Indicates the row index within the convolution kernel; This represents the column index within the convolution kernel; N represents the height of the convolution kernel; M represents the width of the convolution kernel; This indicates that the convolution kernel of the c-th channel is located at position... Weight parameters at the location; Sparse attention computation is adopted to replace the multi-head self-attention mechanism, and attention weights are calculated only for candidate regions of potential hazards. The sparse attention computation formula is as follows: ; in, This represents the result of sparse attention calculation; Represents the query matrix; Represents the key matrix; Represents a value matrix; Represents the transpose of the key matrix; Indicates the dimension of the key matrix; Indicates the scaling factor; This represents the activation function, used to convert the attention score into a probability distribution; The spatial mask matrix represents a value of 1 for candidate regions of potential hazards and a value of 0 for background regions. Model quantization converts 32-bit floating-point parameters to 16-bit floating-point parameters.
[0047] Lightweighting aims to address the challenges of ViT models' massive computational and storage requirements, making them difficult to deploy on edge devices such as inspection robots with limited computing power. It involves multiple lightweight restructurings at the system level without significantly sacrificing detection accuracy: First, the number of encoder layers is reduced. The standard ViT model's 12-layer encoder is reduced to 6 to 8 layers, eliminating redundant coding layers that are responsible for high-level semantic abstraction but do not provide much benefit for specific industrial hazard detection tasks, thus directly reducing the computational load caused by network depth.
[0048] Secondly, a channel pruning algorithm is used to compress the channel dimensions of each layer. The contribution of each channel to the final detection task is evaluated by calculating its importance score, which is the sum of the absolute values of the convolutional kernel weights for that channel. During pruning, a compression ratio of 30% to 50% is set, retaining only the top 50% to 70% of highly important channels and removing a large number of channels carrying redundant information or with extremely weak responses, thus significantly reducing the number of parameters without damaging the network structure.
[0049] Secondly, the attention computation mechanism is optimized. The resource-intensive global multi-head self-attention is replaced with sparse attention. The core idea is that in industrial inspection scenarios, most background areas (such as open factory grounds and the sky) do not require complex attention interaction calculations. Therefore, a lightweight candidate region generation network first identifies hazard candidate regions with a confidence level higher than a threshold T. Then, a spatial mask is used to force attention computation to occur only within these hazard candidate regions, ignoring background areas. This reduces the computation complexity, which was originally quadratic with the image size, to a value proportional only to the number of hazard candidate regions, significantly improving inference speed.
[0050] Finally, model quantization is performed. The parameters of the trained 32-bit floating-point (FP32) model are converted to 16-bit floating-point (FP16). This operation immediately compresses the model storage volume by about 50%, and on most edge computing hardware that supports half-precision computing, inference speed can be improved by more than 30%. The accuracy loss of this magnitude (controlled within 2%) is within the acceptable range for industrial testing, making it an effective means of achieving high real-time performance at a low cost.
[0051] In step S5, the differentiable counterfactual intervention further includes: performing at least one of dust intervention, light intervention, and motion intervention on the candidate area of potential hazards; The dust intervention involves forcibly assuming that the dust concentration is zero, and then re-extracting the real hazard-related features in the decoupled feature space to obtain the first counterfactual hazard features. The lighting intervention involves forcibly assuming that the lighting intensity is a standard reference value, correcting the image lighting to ideal conditions, and then re-propagating it forward to extract real hazard-related features and obtain the second counterfactual hazard features. The motion intervention involves forcibly assuming zero motion interference, eliminating the effects of timing jitter and instantaneous motion, and then re-extracting the relevant features of the real hidden danger to obtain the third counterfactual hidden danger features. The similarity of the real hidden danger-related features with the first counterfactual hidden danger features, the second counterfactual hidden danger features, and the third counterfactual hidden danger features is compared, and the real hidden danger or environmental interference is determined based on the comparison results.
[0052] This step is the core innovation of this invention, aiming to break through the limitations of existing "black box" models that passively fit data and cannot distinguish between hidden dangers and interference, and endow the model with anti-interference ability and interpretability from the perspective of causal mechanism. Specifically, it includes: First, a lightweight causal graph for industrial scenarios is constructed. A three-layer causal structure model is systematically established to address the three most common types of disturbances in high-risk scenarios such as metallurgy and chemical engineering—dust occlusion, abrupt changes in illumination, and motion interference. This model includes: a latent variable layer with dust concentration, illumination intensity, and motion interference as nodes; an observation layer with local texture, edge gradient, and color distribution as nodes; and a target layer with "whether it is a real hazard" as a node. The causal strength relationships between these nodes are obtained through joint learning using prior knowledge from physical optics and imaging as constraints, combined with a small amount of finely labeled causal data, forming an interpretable and transferable lightweight causal graph.
[0053] Secondly, causal feature decoupling encoding is performed. Based on the structure of the lightweight causal graph in the industrial scenario, the local feature map enhanced in step S2 is forcibly decomposed into three mutually orthogonal components in the feature space with non-overlapping information: features related to real hazards, features related to environmental interference, and background residual features. The core constraint of this decoupling process is the minimization of mutual information. That is, in network training, by minimizing the mutual information between hazard features and features related to environmental interference, the information carried by these two types of features is forced to be completely unrelated, fundamentally avoiding false detections caused by feature confusion (such as mislearning "light spots" as "leakage reflection features").
[0054] The specific causal decoupling is achieved through a three-way parallel feature encoder network architecture, which receives enhanced local feature maps and forces them to be mapped to three mutually orthogonal feature subspaces: Main encoder: Responsible for extracting general high-level semantic features.
[0055] Hazard Feature Encoder: Contains multiple convolutional layers and residual blocks, with a projection head connected to the top to project the features of the main encoder into a high-dimensional space representing "hazard correlation", outputting true hazard-related features. .
[0056] Interference Feature Encoder: A mirror image of the structure and hazard feature encoder, but with non-shared weights. It projects the features of the main encoder onto a space representing "interference correlation," outputting environmental interference-related features. .
[0057] Background Feature Encoder: Responsible for capturing the remaining background information not covered by the aforementioned two encoders, and outputting background residual features. .
[0058] To enforce the orthogonality of the feature space, a variational approximation loss based on a lower bound of mutual information is introduced during training as the feature decoupling loss. Specifically, the CLUB mutual information estimator is used to approximate and minimize... and Mutual information between : ; in, It is a variational distribution parameterized by a small neural network to approximate the true conditional distribution, where N is the batch size. By jointly optimizing the task loss and the mutual information minimization loss, the network is forced to learn to separate the hazard features from the environmental disturbance-related features and place them in mutually orthogonal subspaces, thereby avoiding feature confusion from a mechanistic perspective.
[0059] Finally, differentiable counterfactual interventions are implemented. This is crucial for achieving interpretable judgments. During model inference, for a candidate hazard area, the system proactively raises a "counterfactual question": "If there were no dust at this time, would it still be a hazard?" Specifically, three types of causal interventions are performed: dust intervention (artificially setting the dust concentration node to 0), illumination intervention (artificially correcting the illumination intensity node to the standard reference value), and motion intervention (artificially eliminating the influence of temporal jitter). The pre-trained network is then re-propagated forward to generate a set of "counterfactual features" under an "ideal interference-free" environment. The original hazard features are then compared with these counterfactual features for similarity calculation. If, after removing all interference, the hazard features remain highly similar and stable, it is judged as a real hazard; if the features change drastically or essentially disappear, it indicates that the target is a false hazard caused entirely by environmental interference. At this point, the model can not only provide accurate detection results but also output interpretable judgment reasons to engineers based on the type of counterfactual intervention that has changed significantly, such as "false detection due to dust obstruction" or "false detection due to sudden changes in illumination." To ensure that the counterfactual intervention process is fully differentiable and supports end-to-end training, the above intervention behavior is achieved by injecting learnable masks and correction parameters into specific feature layers: Dust intervention: Instead of directly modifying the input image, it uses a learnable "dust mask" channel in the decoupled feature space. All activation values are forcibly set to zero, and the perturbed features are then forward-propagated to simulate the ideal feature state without dust.
[0060] Illumination Intervention: A lightweight adaptive instance normalization module is introduced at the front end of the original ViT encoder, whose affine transformation parameters are determined by a small network. The predictions are based on the current feature statistics. When performing the intervention, these parameters are forced to be overridden to the pre-trained standard reference values, thereby normalizing the features to standard lighting conditions before subsequent embedding and extraction.
[0061] Motion intervention: After temporal alignment, instead of simply removing jitter, when weighting and summing the features of adjacent frames after feature alignment, the temporal weight of the previous frame is differentially reduced to 10% of its original value to simulate the state where motion blur information is suppressed, and then the fused features are extracted.
[0062] By allowing differentiable adjustment of the control variables in the feature space, the entire counterfactual reasoning process can be tracked by a computational graph, enabling the gradient of the counterfactual consistency loss to propagate back and be used to optimize the parameters related to the feature extractor and intervention strategy.
[0063] In step S6, the cross-frame temporal feature fusion mechanism specifically includes: The improved Vision Transformer network is used to extract features from the local feature maps of 3 to 8 consecutive frames to obtain the high-level local features corresponding to each frame. Optical flow is used to temporally align the high-level local features of consecutive frames to eliminate feature position offsets caused by device movement; the optical flow field calculation formula is as follows: ; in, It represents the optical flow vector field, that is, the displacement vector of the pixel located at pixel coordinates (x,y) in the t-th frame of the image when it moves to the (t+1)-th frame; This represents the brightness value of the image at pixel coordinates (x, y) in the t-th frame; Represents the optical flow vectors in the horizontal and vertical directions; Indicates adjacent frame images; This represents the brightness value corresponding to the position that the (x, y) pixel in frame t moves to in frame t+1. A gated recurrent unit is constructed to perform correlation modeling on the temporally aligned high-level local features, extracting the temporal fusion features of the potential hazard targets; the update formula for the gated recurrent unit is: ; ; ; ; in, These represent updating the door and resetting the door, respectively. Indicates the hidden state of the previous frame; Indicates the candidate hidden state; Represents the temporal fusion features of the current frame; These represent the weights of the update gate, reset gate, and candidate state, respectively. This represents the input features at the current moment, i.e., high-level local features; The similarity of the temporal fusion features of consecutive frames is calculated. When the similarity is higher than a preset threshold, it is determined to be a real potential hazard. When the similarity is lower than the preset threshold and the feature anomaly only occurs in a single frame, it is determined to be transient environmental interference. The similarity calculation formula is as follows: .
[0064] Step S6 leverages the characteristics of video inspection to further enhance detection robustness from a temporal perspective, resolving issues that single-frame detection cannot distinguish between "flying insects" and "cracks," or "light and shadow" and "leakage." Specific implementation includes: First, for a sequence of 3 to 8 consecutive preprocessed image frames, high-level local features of each frame are extracted by improving the ViT model.
[0065] Secondly, to address pixel displacement between adjacent frames caused by robot movement or camera shake, optical flow is used for temporal alignment. By calculating the optical flow vector of each pixel between two consecutive frames—that is, the movement velocity in the horizontal and vertical directions—the correspondence between pixels between frames can be accurately established. Based on this vector, coordinate correction is performed on the feature maps of subsequent frames, aligning features belonging to the same real-world location in different frames and eliminating feature misalignment caused by motion.
[0066] Next, a temporal feature fusion network is constructed. The aligned consecutive frame feature sequences are fed into a gated recurrent unit (GRU) for temporal correlation modeling. The GRU uses its internal reset gate... and the update gate It can memorize and integrate the long-term patterns of how potential hazards evolve over time, and output the current frame temporal fusion feature ht, which contains historical information. For example, it can learn the typical feature pattern of "the leak area slowly expanding over time".
[0067] Finally, a real-world hazard assessment is performed. The preset threshold (temporal feature similarity threshold) α ranges from 0.6 to 0.8. The temporal fusion features of the GRU outputs for consecutive frames are calculated. and The cosine similarity is used. If the similarity of multiple consecutive frames is consistently higher than the preset threshold α, it indicates that the target feature is changing continuously and gradually, consistent with the development pattern of real hazards (such as crack expansion or liquid seepage), and is therefore identified as a real hazard. Conversely, if a feature suddenly appears in a frame, but its similarity with the preceding and following frames drops sharply below the preset threshold α, it indicates that this is an isolated, transient abnormal signal, consistent with the characteristics of interference such as flying insects or flickering light and shadow, and is therefore identified as transient interference and filtered out. The preset threshold α can be dynamically adjusted based on the robot's movement speed and prior knowledge of hazard development.
[0068] In step S7, the industrial hazard detection result carries the industrial hazard category, industrial hazard location, hazard confidence level, instantaneous interference filtering mark, and interference discrimination reason.
[0069] Step S7 is the complete inference and output closed loop. It normalizes the inputs of all the aforementioned modules, performs forward computation through the trained improved ViT model, and post-processes the results to generate deliverable industrial hazard detection results. First, it loads the trained improved ViT model weights and switches to inference mode, unifying the standardized images and temporal fusion features processed in steps S1-S5 into the tensor format required by the improved ViT model. Then, it inputs the tensor into the improved ViT model, sequentially passing through modules such as local feature enhancement, causal decoupling, lightweight encoder inference, and temporal fusion, to obtain the original inference results containing industrial hazard category, industrial hazard location (detection box), hazard confidence, and transient interference filtering labels, as well as the interference discrimination reasons generated by the counterfactual intervention in step S3. Next, it post-processes the original inference results: sets a confidence threshold (e.g., 0.5) to filter out low-confidence industrial hazard locations; applies the non-maximum suppression (NMS) algorithm to remove multiple overlapping detection boxes for the same hazard target, retaining only the unique box with the highest hazard confidence. Then, the normalized bounding box coordinates output by the model are back-mapped back to the actual pixel positions of the inspected image sequence, ensuring that the output positions correspond to the physical world. Finally, the final industrial hazard detection results are output in accordance with the industrial inspection standard format. For example, an industrial hazard detection result might be output as: "Industrial hazard category: molten metal leakage, industrial hazard location: (x1, y1, x2, y2), hazard confidence: 0.97, judgment: real hazard, interference analysis: dust and light interference are eliminated through causal decoupling"; or for a false alarm signal, the output might be: "Industrial hazard category: no leakage, transient target, judgment: instantaneous interference (flying insect), reason: temporal feature similarity is below the threshold, and counterfactual features disappear under motion intervention." This multi-dimensional information output greatly facilitates the backend system in automatically blocking false alarms and provides clear and reliable decision-making basis for human-machine collaborative judgment by on-site engineers, solving the model interpretability problem that has plagued the industry.
[0070] In summary, the advantages of this invention are as follows: 1. By constructing a local feature enhancement module embedded with multi-scale convolution and channel attention, the perception accuracy of weak target hazards is significantly improved. At the same time, a physically guided causal decoupling mechanism and differentiable counterfactual intervention are introduced to decompose image features into features related to real hazards, features related to environmental interference, and background residual features. The features related to real hazards and features related to environmental interference are forced to be orthogonal, thereby greatly enhancing the ability to resist environmental interference such as sudden changes in illumination and dust occlusion. On this basis, the optical flow method is used to align continuous frame features and a gated recurrent unit is used to perform temporal correlation modeling. By comparing the similarity between frames, the model can effectively distinguish between persistent real hazards and transient environmental interference, and realize temporal causal discrimination. The model can also directly output the industrial hazard category and the reason for interference discrimination, giving the industrial hazard detection results complete interpretability. Through lightweight reconstruction measures such as reducing the number of encoder layers, channel pruning, sparse attention, and model quantization, the number of model parameters and computational cost are significantly reduced, meeting the requirements of real-time deployment on the edge.
[0071] 2. Integrating local enhancement and global modeling to improve feature representation accuracy: A local feature enhancement module is embedded in the Vision Transformer (ViT) encoder. Local features are extracted by parallel 3×3, 5×5, and 7×7 multi-scale convolutional kernels, and feature enhancement is performed by combining channel attention mechanism. Compared with the traditional ViT which only relies on global self-attention, it can capture subtle local textures and wide-area contextual information in industrial scenarios at the same time. It effectively solves the problem of variable target scale and easy loss of details in hidden dangers, and significantly improves the detection capability of small hidden dangers (such as cracks and leaks).
[0072] 3. Introducing physically guided causal decoupling and counterfactual intervention to enhance anti-interference and interpretability: By decomposing image features into three parts—features related to real hazards, features related to environmental interference, and background residual features—and performing differentiable counterfactual interventions (such as dust, lighting, and motion interventions), it can actively simulate the ideal features after "removing interference." By comparing the similarity with the original features, the authenticity of the hazard can be determined. This mechanism not only significantly reduces false alarms caused by dust, lighting changes, and equipment vibration in complex industrial environments, but also outputs reasons for interference judgment (such as "current high dust levels lead to false detection"), making the detection results interpretable and facilitating on-site personnel to trace and verify them.
[0073] 4. Lightweight Reconstruction Design to Meet Real-Time Deployment Needs on the Edge: To address the limited computing power of edge computing devices in industrial inspections, the ViT network has undergone system-level lightweighting: the number of encoder layers has been reduced to 6-8, a channel pruning algorithm has been used to compress the channel dimension, global multi-head self-attention has been replaced with sparse attention (only calculating candidate regions for potential hazards), and the parameters have been quantized from 32-bit floating-point to 16-bit floating-point. These measures significantly reduce the number of model parameters and computational overhead while maintaining high accuracy, enabling inference speeds of over 30 frames per second, making it suitable for real-time detection scenarios on mobile devices such as drones and robots.
[0074] 5. Cross-frame temporal feature fusion effectively distinguishes between transient interference and persistent hazards: By temporal alignment (optical flow method) and gated recurrent unit (GRU) modeling of 3 to 8 consecutive frames of images, the dynamic laws of hazard evolution over time can be explored; when an abnormal feature appears in a frame but the temporal similarity drops sharply, it is judged as transient environmental interference (such as flying insects, reflections); if the features of multiple frames change continuously and the similarity is stable, it is judged as the development of a real hazard; this cleverly utilizes the characteristics that hazards in industrial scenarios are usually gradual while interference is mostly transient, further filtering out false detections in a single frame and improving detection reliability.
[0075] 6. Comprehensive and quantifiable pre-training strategies ensure model generalization and robustness: The pre-training process constructs a dedicated dataset covering various hazard types and on-site working conditions, and employs joint optimization using cross-entropy loss, location regression loss, and causal consistency training loss (including feature decoupling loss and counterfactual consistency loss); among them, minimizing mutual information forces the orthogonality of real hazard-related features and environmental interference-related features, prompting the model to learn hazard representations that are independent of the environment; the training qualification standards are clearly set as an average detection accuracy of ≥96%, small target hazard accuracy of ≥90%, and edge frame rate of ≥30 frames / second, providing clear and verifiable performance thresholds for product deployment.
[0076] 7. End-to-end output of multi-dimensional detection results to assist operation and maintenance decisions: The final output of industrial hazard detection results not only includes the industrial hazard category, industrial hazard location, and hazard confidence level, but also carries instantaneous interference filtering markers and specific interference discrimination reasons (such as "motion intervention leads to inconsistent features"). This information-rich output format facilitates the backend system to automatically filter false alarms, while providing a basis for human-machine collaborative review, significantly improving the efficiency and reliability of industrial site hazard handling.
[0077] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A lightweight ViT industrial hazard detection method decoupled from temporal characteristics and causal factors, characterized in that: Includes the following steps: Step S1: Obtain the inspection image sequence of the industrial scene, and preprocess the inspection image sequence to obtain standardized images; Step S2: Construct a Vision Transformer network and pre-train the Vision Transformer network, introducing causal consistency training loss during the pre-training process; Step S3: Improve the pre-trained Vision Transformer network by embedding a local feature enhancement module into the encoder of the Vision Transformer network and perform lightweight reconstruction on the improved Vision Transformer network. Step S4: Using the lightweight reconstructed Vision Transformer network, extract features from the standardized image and output an enhanced local feature map; Step S5: Introduce a physically guided hazard-interference causal decoupling mechanism to perform causal decoupling on the local feature map. The causal decoupling includes: decomposing the local feature map into real hazard-related features, environmental interference-related features, and background residual features, and then performing differentiable counterfactual intervention: simulating the removal of environmental interference based on the real hazard-related features to generate counterfactual features, comparing the similarity between the real hazard-related features and the counterfactual features, determining whether it is a real hazard or environmental interference based on the comparison result, and outputting the reason for interference discrimination. Step S6: Introduce a cross-frame temporal feature fusion mechanism to perform temporal alignment and correlation modeling on the local feature maps of multiple consecutive frames, so as to distinguish between the development of real hidden dangers and instantaneous environmental interference, and output temporal fusion features; Step S7: Input the standardized image and time-series fusion features into the pre-trained Vision Transformer network that has undergone lightweight reconstruction, and output the industrial hazard detection result in conjunction with the interference discrimination reason.
2. The lightweight ViT industrial hazard detection method based on temporal characteristics and causal decoupling as described in claim 1, characterized in that: In step S1, the preprocessing specifically includes: The inspection image sequence is noise-removed using Gaussian filtering. The Gaussian filtering function is: ; in, Indicates the Gaussian filter kernel in coordinates The weight value at the location; σ represents the position coordinates of a pixel within the Gaussian filter kernel relative to the kernel center; σ represents the standard deviation of the Gaussian distribution, used to control the width and smoothness of the Gaussian filter kernel. This represents the normalization coefficient, used to ensure that the sum of all weights of the Gaussian filter kernel is 1; The Gaussian-filtered inspection image sequence is normalized using a linear stretching algorithm. The normalization formula is as follows: ; in, This represents the normalized image in pixel coordinates. The grayscale value at that location is mapped to the range of 0~255; This indicates the pixel coordinates of the inspection image sequence after Gaussian filtering. The grayscale value at that location; This represents the minimum gray value of all pixels in the Gaussian-filtered inspection image sequence; 255 represents the maximum gray value of all pixels in the Gaussian filtered inspection image sequence; 255 represents the target upper limit for gray-level normalization. The grayscale-normalized inspection image sequence is scaled proportionally to a preset size to obtain a standardized image. The scaling process uses bilinear interpolation, and the interpolation formula is as follows: ; in, This indicates the normalized image obtained by scaling in coordinates. Pixel value at; This represents the pixel coordinates in the normalized image; m and n both represent summation indices, where m=0,1 represents two adjacent integer coordinates in the horizontal direction, and n=0,1 represents two adjacent integer coordinates in the vertical direction. This indicates that in the grayscale normalized inspection image sequence, the image is in line with... Integer coordinates of four adjacent pixels; This indicates that in the grayscale-normalized inspection image sequence, in the coordinates... The pixel value at that location.
3. The lightweight ViT industrial hazard detection method based on temporal characteristics and causal decoupling as described in claim 1, characterized in that: In step S2, the pre-trained Vision Transformer network is obtained through the following process: Construct a dedicated dataset of potential hazards in high-risk industrial scenarios, covering various types of industrial hazards and different on-site working conditions; The dataset specifically designed for potential hazards in high-risk industrial scenarios is divided into a training set, a validation set, and a test set, and data augmentation is performed on the training set. The Vision Transformer network is iteratively trained using the weighted sum of cross-entropy loss and position regression loss as the total loss function. Training of the Vision Transformer network stops when the average detection accuracy on the validation set stabilizes, and the performance of the trained Vision Transformer network is validated using the test set.
4. The lightweight ViT industrial hazard detection method based on temporal characteristics and causal decoupling as described in claim 1, characterized in that: In step S2, the causal consistency training loss is used to constrain the Vision Transformer network to learn real hazard-related features that are independent of the environment. The causal consistency training loss is a joint loss function that includes task loss, feature decoupling loss, and counterfactual consistency loss. The feature decoupling loss is achieved through mutual information minimization constraints, which force the real hidden danger-related features and environmental interference-related features to be orthogonal in the feature space, so that the information carried by the real hidden danger-related features and environmental interference-related features does not overlap.
5. The lightweight ViT industrial hazard detection method based on temporal characteristics and causal decoupling as described in claim 1, characterized in that: In step S2, during the pre-training process, the criteria for determining whether the training is successful are as follows: Average detection accuracy ≥96%, small target hazard detection accuracy ≥90%, edge-side inference speed ≥30 frames / second.
6. The lightweight ViT industrial hazard detection method based on temporal characteristics and causal decoupling as described in claim 1, characterized in that: In step S3, the workflow of the local feature enhancement module is as follows: Multi-scale local receptive field units are constructed, and convolutional kernels of three sizes (3×3, 5×5, and 7×7) are connected in parallel. Parallel convolution operations are performed on the input standardized image to capture local features at the three scales respectively. The convolution operation formula is as follows: ; in, This indicates that when the kernel size is k, the output local features are located at... The pixel value at that location; k represents the size of the convolution kernel, which can be 3, 5, or 7. Represents the relative coordinates within the convolution kernel, with values ranging from 0 to k−1, used to traverse each element of the convolution kernel; This represents the relative positions of a k×k convolution kernel. Weight parameters at the location; Indicates the normalized image at location Pixel value at that location, It is the top-left anchor point of the standardized image corresponding to the output position; The local features at the three scales are normalized and then channel-level stitched together to obtain a fused feature map. The fused feature map is input into a feature enhancement convolutional layer with an embedded channel attention mechanism. The importance weights of each feature channel are calculated through the channel attention mechanism, and the importance weights are multiplied by the fused feature map to output the enhanced local feature map.
7. The lightweight ViT industrial hazard detection method based on temporal characteristics and causal decoupling as described in claim 1, characterized in that: In step S3, the lightweight reconstruction specifically includes: The number of encoder layers in the improved Vision Transformer network is reduced to 6-8 layers; A channel pruning algorithm is used to compress the channel dimensions of each layer in the improved Vision Transformer network, and effective channels are selected based on channel importance scores; the formula for calculating the channel importance score is as follows: ; in, This represents the importance score of the c-th channel; c represents the channel index. Indicates the row index within the convolution kernel; This represents the column index within the convolution kernel; N represents the height of the convolution kernel; M represents the width of the convolution kernel; This indicates that the convolution kernel of the c-th channel is located at position... Weight parameters at the location; Sparse attention computation is adopted to replace the multi-head self-attention mechanism, and attention weights are calculated only for candidate regions of potential hazards. The sparse attention computation formula is as follows: ; in, This represents the result of sparse attention calculation; Represents the query matrix; Represents the key matrix; Represents a value matrix; Represents the transpose of the key matrix; Indicates the dimension of the key matrix; Indicates the scaling factor; This represents the activation function, used to convert the attention score into a probability distribution; The spatial mask matrix represents a value of 1 for candidate regions of potential hazards and a value of 0 for background regions. Model quantization converts 32-bit floating-point parameters to 16-bit floating-point parameters.
8. The lightweight ViT industrial hazard detection method based on temporal characteristics and causal decoupling as described in claim 1, characterized in that: In step S5, the differentiable counterfactual intervention further includes: performing at least one of dust intervention, light intervention, and motion intervention on the candidate area of potential hazards; The dust intervention involves forcibly assuming that the dust concentration is zero, and then re-extracting the real hazard-related features in the decoupled feature space to obtain the first counterfactual hazard features. The lighting intervention involves forcibly assuming that the lighting intensity is a standard reference value, correcting the image lighting to ideal conditions, and then re-propagating it forward to extract real hazard-related features and obtain the second counterfactual hazard features. The motion intervention involves forcibly assuming zero motion interference, eliminating the effects of timing jitter and instantaneous motion, and then re-extracting the relevant features of the real hidden danger to obtain the third counterfactual hidden danger features. The similarity of the real hidden danger-related features with the first counterfactual hidden danger features, the second counterfactual hidden danger features, and the third counterfactual hidden danger features is compared, and the real hidden danger or environmental interference is determined based on the comparison results.
9. The lightweight ViT industrial hazard detection method based on temporal characteristics and causal decoupling as described in claim 1, characterized in that: In step S6, the cross-frame temporal feature fusion mechanism specifically includes: Feature extraction is performed on the local feature maps of 3 to 8 consecutive frames to obtain the high-level local features corresponding to each frame; Optical flow is used to perform temporal alignment of the high-level local features in consecutive frames to eliminate feature position offsets caused by device movement; A gated loop unit is constructed to perform correlation modeling on the time-aligned high-level local features and extract the time-fusion features of the potential hazard targets; The similarity of the temporal fusion features of consecutive frames is calculated. When the similarity is higher than a preset threshold, it is determined to be a real hidden danger development. When the similarity is lower than the preset threshold and only a single frame shows feature abnormality, it is determined to be instantaneous environmental interference.
10. The lightweight ViT industrial hazard detection method based on temporal characteristics and causal decoupling as described in claim 1, characterized in that: In step S7, the industrial hazard detection result carries the industrial hazard category, industrial hazard location, hazard confidence level, instantaneous interference filtering mark, and interference discrimination reason.