Chemical industry park equipment fault detection method and system based on image recognition

By constructing a hierarchical fault detection mechanism and analyzing weak motion characteristics in the spatiotemporal domain, the problem of insufficient sensitivity in detecting minor leaks in equipment in chemical industrial parks has been solved, achieving highly sensitive identification and early warning of minor faults.

CN122289203APending Publication Date: 2026-06-26ZHI YUAN SHU ZI KE JI (SHAN DONG) YOU XIAN GONG SI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610402362.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-30
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing image recognition technology lacks sensitivity in detecting minute leaks in equipment in chemical industrial parks, making it difficult to effectively identify early minor fault characteristics. Especially when the leak hole diameter is extremely small and the leak flow rate is slow, existing detection models are unable to effectively accumulate and amplify weak changes in the spatiotemporal domain.

Method used

A hierarchical fault detection mechanism is constructed, which combines conventional image recognition models and time-accumulated amplification analysis algorithms for weak motion features in the spatiotemporal domain. By performing multimodal feature extraction and adaptive time-scale fusion on continuous video sequences, a highly sensitive identification of minor faults is achieved.

Benefits of technology

It improves the early detection capability and safety warning reliability of equipment failure detection in chemical industrial parks, and can promptly identify minor leaks and weak faults in equipment, thereby reducing safety hazards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289203A_ABST
    Figure CN122289203A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for equipment fault detection in chemical industrial parks based on image recognition, belonging to the field of image processing technology. The method includes: real-time acquisition of monitoring video streams of equipment in the chemical industrial park; detection of significant fault features in the current frame using a conventional image recognition model, outputting a first alarm if present; if no fault features are detected, performing temporal cumulative amplification of motion features on the continuous video sequence to extract subtle spatiotemporal change features to determine minor faults, outputting a second alarm if present. This solves the technical problem in existing image recognition-based equipment fault detection methods where minor faults, due to their extremely small pixel scale and inter-frame variations below noise levels, are difficult to effectively identify due to their weak visual features and slow changes. It achieves the technical effect of identifying early-stage minor faults by accumulating and amplifying subtle spatiotemporal changes on the basis of conventional visual detection, thereby improving fault detection sensitivity and early warning capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically to a method and system for detecting equipment faults in chemical industrial parks based on image recognition. Background Technology

[0002] Chemical industrial parks contain numerous critical equipment such as reaction units, pipelines, valves, and storage tanks. During long-term operation, these devices are prone to malfunctions such as leaks, abnormal vibrations, and structural damage due to corrosion, seal aging, structural fatigue, or abnormal operation. Failure to detect these malfunctions in a timely manner can lead to leaks of flammable and explosive gases, the spread of toxic substances, or a chain reaction of production failures, ultimately causing safety accidents. Therefore, how to continuously monitor the operating status of chemical equipment and promptly detect early malfunctions has always been an important research direction in the field of chemical safety production.

[0003] With the development of video surveillance technology and deep learning image recognition technology, deploying cameras in chemical industrial parks and using image recognition algorithms to automatically detect equipment operating status has become an important safety monitoring method. Existing technologies typically utilize convolutional neural networks (CNNs) or object detection networks to identify abnormal phenomena in video images, such as obvious visual features like liquid leaks, equipment damage, smoke, flames, or abnormal personnel behavior, thereby enabling automatic alarms for equipment malfunctions or safety incidents.

[0004] However, in the early-stage detection of minute leaks in chemical equipment, existing image recognition-based methods still have significant limitations due to the obvious physical characteristics that limit the initial stage of leakage. In monitoring minute leaks in pipelines or valves in chemical industrial parks, the leak orifice is extremely small and the leakage velocity is slow. The spatial scale of the leaking gas in the image is typically only 1-2 pixels, and the inter-frame variation is often lower than the image noise level. This results in a significant deficiency in the sensitivity of existing image recognition methods for identifying minute faults. The root cause lies in the fact that existing detection models mainly rely on significant visual features or short-term motion features in a single frame image for identification, making it difficult to effectively accumulate and amplify subtle changes in the spatiotemporal domain. Therefore, how to effectively identify early-stage minute leaks or weak fault characteristics in chemical equipment under existing video surveillance conditions, and improve the detection sensitivity and reliability of minute faults, has become an urgent technical problem to be solved in the field of intelligent safety monitoring in chemical industrial parks. Summary of the Invention

[0005] This application provides a method and system for equipment fault detection in chemical industrial parks based on image recognition. The key is to address the technical obstacle that conventional image recognition models cannot effectively detect early minor faults in chemical industrial park equipment monitoring scenarios, which are caused by the extremely small spatial scale of minor leaks in images, the smaller inter-frame variation amplitude than image noise, and the lack of obvious texture or color features. By constructing a hierarchical fault detection mechanism and introducing a time-accumulation amplification analysis algorithm for weak motion features in the spatiotemporal domain, and combined with a data processing flow of multimodal feature extraction and time-scale adaptive fusion for continuous video sequences, a highly sensitive identification of minor equipment faults with weak visual features and slow changes is achieved, thereby improving the early detection capability and safety warning reliability of equipment faults in chemical industrial parks.

[0006] The first aspect of this application provides a method for equipment fault detection in chemical industrial parks based on image recognition, the method comprising:

[0007] The system acquires real-time video streams of equipment monitoring captured by cameras within the chemical industrial park. For the current image frame in the video stream, a first fault detection is performed. This first fault detection uses a preset conventional image recognition model to identify whether there are visually significant fault features in the current image frame. If such features are found, a first alarm message is output. If the first fault detection does not identify any visually significant fault features, a second-level minor fault detection is initiated for the continuous video sequence corresponding to the current image frame in the video stream. This involves performing time-accumulated amplification analysis of the motion features of the continuous video sequence, extracting subtle spatiotemporal change features, and determining whether a minor fault exists based on these subtle change features. If such a fault exists, a second alarm message is output.

[0008] A second aspect of this application provides an image recognition-based equipment fault detection system for chemical industrial parks, the system comprising:

[0009] Video acquisition module: acquires real-time video streams of equipment monitoring captured by cameras within the chemical industrial park; First fault detection module: performs first fault detection on the current image frame in the equipment monitoring video stream. The first fault detection uses a preset conventional image recognition model to identify whether there are visually significant fault features in the current image frame. If so, it outputs a first alarm message; Second fault detection module: if the first fault detection does not identify visually significant fault features, it initiates a second-level minor fault detection for the continuous video sequence corresponding to the current image frame in the video stream. It performs time-accumulated amplification analysis of motion features on the continuous video sequence, extracts weak spatiotemporal change features, and determines whether there is a minor fault based on the weak change features. If so, it outputs a second alarm message.

[0010] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0011] First, real-time video streams from cameras in the chemical industrial park are acquired, and preliminary fault detection is performed on the current image frame. A pre-set conventional image recognition model is used to determine whether there are obvious fault features in the image. When a significant anomaly is detected, a first alarm message is generated directly. If no obvious fault is detected, a second-level detection is performed on the continuous video sequence containing the current image frame. By accumulating and amplifying the motion features in the video sequence over time, subtle changes in the spatiotemporal domain are extracted, and the presence of minor faults is determined accordingly. When a minor fault is identified, a second alarm message is generated, thus achieving graded detection and alarming of both significant equipment faults and early minor faults. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A schematic diagram of the process for a chemical industrial park equipment fault detection method based on image recognition provided in this application embodiment.

[0014] Figure 2 A schematic diagram of the structure of an image recognition-based equipment fault detection system for chemical industrial parks provided in this application embodiment.

[0015] Explanation of reference numerals in the attached diagram: Video acquisition module 11, first fault detection module 12, second fault detection module 13. Detailed Implementation

[0016] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0017] Example 1, as Figure 1 As shown, this application provides a method for equipment fault detection in chemical industrial parks based on image recognition, wherein the method includes:

[0018] Real-time acquisition of equipment monitoring video streams collected by cameras within the chemical industrial park.

[0019] In this embodiment, industrial-grade network cameras are pre-deployed in key equipment areas within the chemical industrial park, such as pipe connections, valve areas, tank interfaces, near pumps, and other key monitoring locations where leaks or anomalies may occur. These cameras are connected to the park's monitoring terminal via wired or wireless networks. The monitoring terminal uses a video access protocol to uniformly access and manage each camera, continuously receiving real-time video data streams from each camera. These video data streams are continuously acquired and transmitted to the system server at a preset frame rate. On the server side, the received video streams are decoded, parsing the continuous video data into a sequence of image frames arranged in chronological order. Each frame is appended with a corresponding timestamp and camera device identification information. The parsed image frame sequence is then cached in a video processing buffer to ensure that subsequent image recognition algorithms can stably read the video data in chronological order, thereby providing continuous, stable, and real-time monitoring video data input for subsequent equipment fault detection and anomaly identification.

[0020] A first fault detection is performed on the current image frame in the monitoring video stream of the device. The first fault detection uses a preset conventional image recognition model to identify whether there are visually significant fault features in the current image frame. If there are, a first alarm message is output.

[0021] In one embodiment, the system sequentially extracts the current image frame from the device monitoring video stream according to a preset sampling frequency. The system then preprocesses the current image frame, including at least one of image scaling, brightness correction, noise reduction, contrast enhancement, and region of interest cropping, to improve the stability and computational efficiency of subsequent recognition. Subsequently, the preprocessed current image frame is input into a preset conventional image recognition model. This model can be a convolutional neural network-based object detection model, image classification model, semantic segmentation model, or multi-task deep learning detection network, used to identify surface damage, liquid leakage traces, smoke, flames, abnormal liquid accumulation, abnormal displacement, missing parts, unauthorized personnel approaching, or other visually significant fault features that can be directly presented in a single image frame. The model outputs category labels, location regions, and confidence scores corresponding to various fault features. The system then compares the confidence scores with pre-set recognition thresholds. When the confidence score of any fault feature reaches or exceeds the corresponding threshold, it is determined that the current image frame has a visually significant fault feature, and a first alarm message is generated. This first alarm message includes at least the alarm type, alarm time, camera number, fault feature category, fault area coordinates, and corresponding captured image. The first alarm message is sent to the monitoring platform for display and storage, and is available for subsequent confirmation and handling by operators, thereby achieving rapid identification and real-time early warning of significant equipment faults in the chemical industrial park.

[0022] Furthermore, the conventional image recognition model is a multi-task deep learning detection network, which includes multiple salient feature recognition branches; when the detection result of any branch exceeds the corresponding threshold, it is determined that there is a visually significant fault feature and the first alarm information is output.

[0023] Preferably, the current image frame is extracted from the device monitoring video stream, and image preprocessing is performed on the current image frame to obtain a standardized input image. Subsequently, the standardized input image is input into the shared feature extraction backbone network of a conventional image recognition model. This conventional image recognition model is a multi-task deep learning detection network. This shared feature extraction backbone network can adopt a ResNet, CSPDarknet, or MobileNet structure. Through multi-layer convolution operations and downsampling operations, it extracts multi-scale features of the image to generate feature maps containing spatial and semantic information. Subsequently, the feature maps are input into multiple salient feature recognition branches, such as open flame detection, dense smoke detection, large-scale leak detection, and equipment deformation detection. Each recognition branch detects different types of visually significant faults. The open flame detection branch, based on the characteristic high brightness, warm color, and irregular edges of flame areas, processes the feature map through convolutional and classification layers, outputting the bounding boxes, category labels, and corresponding confidence scores of candidate flame areas. The dense smoke detection branch, considering the diffusion pattern, low contrast, and blurred texture of smoke, generates and classifies candidate regions from the feature map to identify whether smoke diffusion exists around the equipment, outputting the location and confidence score of the smoke area. The large-scale leak detection branch identifies large areas of liquid spraying, flowing, or accumulating around the equipment, detecting liquid texture features, flow direction, and area, outputting the leak area boundary and leak confidence score. The equipment deformation detection branch learns from the equipment's contour structural features to detect structural deformations such as pipe bending, flange misalignment, shell bulging, or abnormal support structures, outputting the corresponding abnormal areas and deformation confidence scores. Each significant feature recognition branch outputs the detection category, target location coordinates, and corresponding confidence score. The system compares the confidence score output by each branch with a preset threshold. When the confidence score of any recognition branch is higher than the corresponding threshold, it is determined that there is a visually significant fault feature in the current image frame, and a first alarm message is generated. This first alarm message includes at least the alarm type, fault category, alarm time, camera number, fault area coordinates, and corresponding captured image. The system sends this first alarm message to the monitoring platform for real-time display and storage, so that operators can view it in a timely manner and take subsequent actions.

[0024] For example, taking the open flame detection branch as an example, a shared feature extraction backbone network is set up in the multi-task deep learning detection network. This backbone network can adopt a ResNet-50 structure to extract features from the input image at multiple scales. Specifically, the input image is uniformly adjusted to 640×640 pixels and pixel normalization is performed so that the pixel values ​​are mapped to the [0,1] interval. Subsequently, the image is input into the backbone network, and features are extracted layer by layer through convolutional layers, batch normalization layers, and ReLU activation functions, outputting three feature maps of different scales, such as 80×80×256, 40×40×512, and 20×20×1024 feature tensors, which are used for subsequent detection branch processing. In the open flame detection branch, the above multi-scale feature maps are input into the feature fusion module. This feature fusion module can be an FPN structure, which upsamples and laterally connects and fuses features of different scales to enhance the expressive ability of small target flame regions. The fused feature maps are compressed by 3×3 convolutional layers and then generated as candidate detection feature maps by 1×1 convolutional layers. Subsequently, three anchor boxes are set at each detection scale. The anchor box sizes can be set to 10×13, 16×30, and 33×23 for small-scale flames; 30×61, 62×45, and 59×119 for medium-scale flames; and 116×90, 156×198, and 373×326 for large-scale flames. For each anchor box, the network output includes the target center point coordinate offset, predicted target width and height, target presence probability, and flame category probability. During model training, a dataset containing both flame and non-flame image samples is constructed. Flame regions are labeled manually to generate bounding boxes. The loss function can be a weighted combined loss function, including localization loss, confidence loss, and classification loss. Model training can use the Adam optimizer with an initial learning rate of 1×10⁻⁶. -4 The batch size was set to 16, and training was conducted for 100-200 epochs, with a learning rate decay strategy implemented using cosine annealing. During actual detection, after the open flame detection branch outputs all candidate flame targets, a non-maximum suppression algorithm is used to remove duplicate detection boxes. When the confidence level of any flame target in the detection result is greater than or equal to 0.5 and the target area exceeds a preset area threshold, it is determined that an open flame fault feature exists in the current image frame, and the first alarm message is generated.

[0025] If the first fault detection fails to identify visually significant fault features, then for the continuous video sequence in the video stream corresponding to the current image frame, a second-level minor fault detection is initiated. The continuous video sequence is subjected to time-accumulated amplification analysis of motion features to extract weak spatiotemporal change features, and the presence of a minor fault is determined based on the weak change features. If a minor fault exists, a second alarm message is output.

[0026] In one embodiment, if the first fault detection fails to identify visually significant fault features such as open flames, dense smoke, large-scale leaks, or equipment deformation in the current image frame, the system does not immediately determine that the equipment is in a normal state. Instead, it further initiates a second-level minor fault detection for the continuous video sequence corresponding to the current image frame. Specifically, the system uses the current image frame as the center and extracts several consecutive frames from the equipment monitoring video stream according to a preset time window to form a video segment to be analyzed. This time window can be set according to the camera frame rate and detection requirements. For example, it can select a predetermined number of image frames before and after the current image frame, or select multiple consecutively acquired images before the current image frame to form a continuous video sequence. Subsequently, the continuous video sequence is preprocessed. This preprocessing includes at least image registration, noise reduction filtering, brightness compensation, and background stabilization to eliminate the interference of slight camera shake, ambient light fluctuations, and random noise on the detection of minor faults. After preprocessing, the system performs a cumulative temporal analysis of motion-related information such as pixel changes, local texture disturbances, edge drift, and background refraction distortion in continuous video sequences. It superimposes and amplifies subtle changes that are difficult to distinguish from noise in a single frame at multiple time points, gradually revealing minute anomalies that were initially difficult to identify in a single frame due to their extremely small pixel scale and amplitude below the image noise level. This yields spatiotemporal features that reflect the subtle motion trends around the equipment. For example, for slow airflow disturbances, refraction changes, or local grayscale fluctuations caused by minor leaks, the system calculates grayscale differences between consecutive frames, the intensity of motion changes, and the consistency of changes in local areas along the time axis. This separates persistent but low-amplitude changes from random noise and forms corresponding subtle change feature representations. These spatiotemporal subtle change features are then input into a pre-defined minor fault classifier for matching analysis to determine whether a minor fault exists within the current equipment monitoring area. When subtle changes in characteristics meet preset fault judgment conditions, the system determines the existence of a minor fault and generates a second alarm message. This second alarm message includes at least the alarm time, camera number, fault location, minor fault type, characteristic intensity value, and corresponding video clip or keyframe image. Finally, the second alarm message is sent to the monitoring platform for display and storage, thereby improving the detection sensitivity of early minor faults in equipment with weak visual characteristics and slow changes, and realizing early identification and safety warning of equipment faults in chemical industrial parks.

[0027] Furthermore, a second-level minor fault detection is initiated, which involves performing temporal cumulative amplification analysis of the motion features of the continuous video sequence, extracting subtle spatiotemporal variation features, and determining the presence of minor faults based on these subtle variation features, including:

[0028] The continuous video sequence is input into a long-term motion memory unit (LTM). The gradient of grayscale value over time is calculated for each pixel location, and the absolute values ​​of the gradients in the time dimension are accumulated to form a spatiotemporal gradient accumulation feature map. Each frame of the continuous video sequence is phase-magnified to reconstruct a phase-magnified video sequence. In the phase-magnified video sequence, static background regions are identified, and the optical flow field of the background region in each frame is calculated. The second derivative of the optical flow field is then calculated to generate a schlieren feature map sequence reflecting texture distortion caused by gas refraction. The spatiotemporal gradient accumulation feature map, the phase-magnified video sequence, and the schlieren feature map sequence are input into a temporal pooling layer. This temporal pooling layer uses an attention mechanism to adaptively select a time scale matching the current leakage rate and performs multimodal feature weighted fusion to obtain a comprehensive micro-fault feature tensor. The comprehensive micro-fault feature tensor is then input into a trained micro-fault classifier to complete the micro-fault identification.

[0029] Preferably, the continuous video sequence is first input into a long-term motion memory (LTM) unit. This LTM unit is constructed as a circular buffer in system memory or a dedicated cache to continuously store historical video frame sequences. Its storage duration is no less than 5 minutes of surveillance video frames from the past period. For example, when the camera frame rate is 25 frames per second, the LTM unit stores at least approximately 7500 historical images. The system reads the frame sequence in chronological order, calculates the grayscale difference between adjacent frames for each pixel position (x, y), and obtains the temporal gradient. ,in, This represents the grayscale value at pixel position (x, y) in frame t. Subsequently, the absolute values ​​of the calculated temporal gradients are accumulated and summed over time, gradually amplifying pixel changes that have a long duration but extremely small amplitude. The accumulated result is then normalized and contrast-enhanced to generate a spatiotemporal gradient accumulation feature map, making subtle changes appear more prominent in the image. Simultaneously with obtaining the spatiotemporal gradient accumulation feature map, a phase amplification network is used to phase amplify each frame in the continuous video sequence. The image is then reconstructed through inverse transform, resulting in a phase-amplified video sequence. This enhances minute movements or air refraction disturbances that are difficult to observe with the naked eye. Afterward, stable structural regions are extracted from the phase-amplified video sequence as static background regions, and a background region mask is generated. This static background region includes structural regions that maintain a fixed position during normal operation, such as pipe surface textures, equipment nameplates, and equipment support structures. Next, the dense optical flow field between adjacent frames is calculated within the static background region to obtain the pixel displacement vector field F(x,y)=(u,v). Then, the second derivative of this optical flow field is calculated, including divergence and curl features. Divergence describes the diffusion or contraction changes in a local region, while curl describes the rotation changes. Since minute gas leaks alter the air refractive index, causing slight distortions in the background texture, a schlieren feature map sequence reflecting the texture distortion caused by gas refraction can be generated by calculating the divergence and curl and performing low-pass filtering in the temporal dimension. Then, the spatiotemporal gradient accumulation feature map, the phase-magnified video sequence, and the schlieren feature map sequence are input into a temporal pooling layer for joint analysis.

[0030] Specifically, the three types of input features are first processed for time alignment and size unification. The spatiotemporal gradient accumulation feature map is used as the first modality feature, the frame-level feature map obtained by convolutional feature extraction from each frame in the phase magnification video sequence is used as the second modality feature, and the divergence feature map and curl feature map at each time step in the schlieren feature map sequence are combined as the third modality feature. The three types of modality features are then divided into multiple time windows of different lengths according to a unified time axis. This time window includes at least a short time window, a medium time window, and a long time window. The short time window is used to characterize the instantaneous changes caused by rapid leakage, the medium time window is used to characterize the continuous changes caused by slow leakage, and the long time window is used to characterize the weak anomalous changes formed by the accumulation of extremely slow leakage over a long period of time. For example, they can be set to 5 seconds, 30 seconds, and 300 seconds, respectively. For each time window, the system calculates the statistical response values ​​of three modal features within that window. These statistical response values ​​include at least one of the following: feature mean, feature variance, local peak intensity, duration proportion, and spatial continuity index. Specifically, for spatiotemporal gradient accumulation feature maps, the system calculates the area proportion of high-response regions and the average gradient intensity; for phase-magnified video sequences, it calculates the average amplitude, peak amplitude, and temporal continuity of phase change responses in each frame; and for schlieren feature map sequences, it calculates the anomalous response intensity, response range, and duration (in frames) of divergence and curl within the background region. Next, the statistical response values ​​corresponding to each time window are input into a temporal attention calculation unit to generate temporal attention weights for each time window. This temporal attention calculation unit can be implemented using a fully connected network, a softmax normalization function, or a gating unit, allowing time windows with stronger, more persistent responses and better conforming to the evolution of small leaks to receive higher weights. After obtaining the attention weights for each time window, the system performs weighted temporal pooling on the three modal features. Specifically, for each modality, the feature vectors extracted under different time windows are multiplied by the corresponding temporal attention weights, and then weighted summation or weighted averaging is performed along the time dimension to obtain the temporal augmentation feature representation for that modality. Next, cross-modal fusion is performed on the temporal augmentation feature representations of the three modalities. This cross-modal fusion can be achieved through feature concatenation, element-wise weighted summation, or attention-based cross-fusion. For example, the spatiotemporal gradient accumulation feature, phase amplification feature, and schlieren feature can be mapped to a unified feature dimension, and then modal weight coefficients can be generated based on the current response intensity of each modality. This allows the modality that is more sensitive in the current scene to contribute more. For instance, in scenes with poor lighting but clear background texture, the weight of the schlieren feature can be increased; in scenes with simple backgrounds and stable temporal changes, the weight of the spatiotemporal gradient accumulation feature can be increased.After completing cross-modal weighted fusion, a unified comprehensive micro-fault feature tensor is output. This comprehensive micro-fault feature tensor retains the ability of spatiotemporal gradient accumulation features to represent long-term weak changes, the ability of phase amplification features to enhance small displacements and vibrations, and the sensitivity of schlieren features to background distortion caused by gas leaks. Therefore, it can more comprehensively reflect weak abnormal changes in continuous video sequences and provide micro-fault classifiers with input features that have higher discriminative power.

[0031] Finally, the comprehensive micro-fault feature tensor is input into a pre-trained micro-fault classifier for judgment. This micro-fault classifier can employ a convolutional neural network or a lightweight temporal classification network, and its output is the probability value of whether a micro-leak exists in the currently monitored area. When the probability value exceeds a preset threshold, such as 0.7, the system determines that a minor fault exists and generates a second alarm message. This second alarm message includes at least the alarm time, camera number, suspected leak area location, detection probability value, and corresponding video clips or keyframe images. The alarm message is then sent to the monitoring platform for display and recording, so as to achieve automatic identification and early warning of early minor leaks in chemical equipment.

[0032] This minor fault classifier is a lightweight spatiotemporal fusion binary classification network. During construction, it first collects equipment monitoring videos under various operating conditions in a chemical industrial park as training samples. These training samples include at least normal operation video samples and video samples with minor leaks. The minor leak video samples can cover leakage scenarios with different leak apertures, leakage rates, shooting distances, lighting conditions, and background complexities. The collected video samples are preprocessed, and according to the aforementioned second-level minor fault detection process, corresponding spatiotemporal gradient cumulative feature maps, phase-magnified video sequences, and schlieren feature map sequences are extracted. After time pooling and multimodal weighted fusion processing, a comprehensive minor fault feature tensor corresponding to each video sample is generated. Simultaneously, a classification label is assigned to the comprehensive minor fault feature tensor based on whether a minor leak actually exists in each video segment; samples with minor leaks are labeled as positive, and those without are labeled as negative. The network structure can be designed as a three-level structure including a feature compression layer, a temporal modeling layer, and a classification output layer. Specifically, the comprehensive micro-fault feature tensor is input into the feature compression layer. Several convolutional layers, batch normalization layers, and activation function layers are used to perform channel compression and local pattern extraction on the input features to reduce feature dimensionality and highlight discriminative information related to micro-leaks. The compressed feature sequence is then input into the temporal modeling layer. This temporal modeling layer can employ any of the following: Gated Recurrent Unit (GRU), Long Short-Term Memory (LSTM), or Lightweight Temporal Convolutional Network (TCN). This layer learns the continuous evolution of micro-leaks over time, enabling the classifier to identify not only instantaneous anomalies but also slowly accumulating weak anomaly changes. Next, the temporal features output from the temporal modeling layer are input into the classification output layer. This classification output layer includes at least one fully connected layer and a sigmoid activation function layer, outputting the probability value of a micro-leak existing in the currently monitored area. During training, the labeled comprehensive micro-fault feature tensor is divided into training, validation, and test sets. A binary cross-entropy loss function is used for supervised training of the micro-fault classifier, and the network parameters are updated using backpropagation. To improve the classifier's adaptability to complex working conditions, data augmentation can be performed on the training samples during the training phase. This data augmentation includes at least one of the following: brightness perturbation, contrast variation, random noise superposition, local occlusion simulation, and background texture variation. This simulates the actual imaging conditions of a chemical industrial park under daytime, nighttime, rainy / foggy weather, and complex industrial backgrounds. Furthermore, the class imbalance problem caused by the relatively small number of minor leak samples can be alleviated by adjusting the ratio of positive to negative samples or introducing class weights, thereby improving the classifier's sensitivity to minor leak samples. After training, the network parameters that perform best on the validation set are saved as the target minor fault classifier model and deployed.In actual operation, the minor fault classifier receives the comprehensive minor fault feature tensor output by the aforementioned time pooling layer and calculates the probability value of a minor leak in the current video sequence. When the probability value exceeds a preset threshold, it is determined that a minor fault exists in the current monitored area, and a second alarm message is generated; when the probability value does not exceed the preset threshold, it is determined that no minor fault is detected in the current video sequence.

[0033] Furthermore, the continuous video sequence is input into a long-term motion memory unit (LTM), the gradient of grayscale value over time is calculated for each pixel location, and the absolute values ​​of the gradients in the time dimension are accumulated to form a spatiotemporal gradient accumulation feature map, including:

[0034] For each pixel position of each frame in the continuous video sequence, the grayscale difference between adjacent frames is calculated; the long-term motion memory unit is set as a circular queue of length N, the queue is updated for each frame processed, and the grayscale difference of all frames in the current queue is accumulated to obtain a cumulative gradient map; the cumulative gradient map is normalized and contrast enhanced, and the spatiotemporal gradient cumulative feature map is output.

[0035] Optionally, for each pixel position in each frame of a continuous video sequence, the grayscale difference between adjacent frames is first calculated in chronological order. By performing this operation on each adjacent frame in the continuous video sequence, a grayscale difference map corresponding to each frame is obtained, which is used to characterize the intensity of local changes in the image over time. Subsequently, the long-term motion memory unit is set as a circular queue of length N, where N corresponds to the total number of frames within a preset long-term time window. For example, when the camera captures frames at a frame rate of 25 frames / second and the preset time window is 5 minutes, N can be set to 7500; when the camera captures frames at a frame rate of 20 frames / second, N can be set to 6000. Each time the system receives a new grayscale difference map, it pushes the grayscale difference map to the tail of the circular queue and removes the earliest grayscale difference map exceeding length N from the head of the queue, thus always maintaining the grayscale difference map sequence corresponding to the most recent 5 minutes stored in the circular queue. For each frame processed, the grayscale difference images in the current circular queue are accumulated pixel by pixel to obtain a cumulative gradient map. The value of each pixel in this cumulative gradient map represents the cumulative result of the grayscale change at that location within the most recent preset time window, thus gradually accumulating and amplifying persistent, weak changes with small amplitudes that are difficult to separate from noise in a single frame. Next, the cumulative gradient map is normalized, mapping each pixel value to a preset grayscale range, for example, 0 to 255, to eliminate the influence of differences in numerical ranges across different time windows or scenes. Based on normalization, the cumulative gradient map undergoes contrast enhancement processing. This contrast enhancement can employ histogram equalization, adaptive histogram equalization, or linear stretching to improve the distinguishability between areas of weak change and the background. After normalization and contrast enhancement, a spatiotemporal gradient cumulative feature map is output for further extraction of early minor leaks or weak abnormal change features in subsequent minor fault identification processes.

[0036] Furthermore, phase magnification is performed on each frame of the continuous video sequence to reconstruct a phase-magnified video sequence, including:

[0037] A complex wavelet transform is performed on each frame of the continuous video sequence to obtain a phase domain representation; the phase domain representation is input into a phase amplification network to amplify the phase change according to a preset small motion frequency range, and then the phase amplified video sequence is reconstructed by inverse transform.

[0038] Optionally, images are first read frame by frame from a continuous video sequence in chronological order. A complex wavelet transform is then performed on each frame to decompose the image into sub-bands of multiple scales and directions, such as horizontal, vertical, and diagonal sub-bands. For each sub-band, corresponding complex wavelet coefficients are obtained. Each complex wavelet coefficient can be represented by two parts: amplitude and phase. The amplitude reflects the local texture intensity of the image, while the phase information reflects subtle changes in the spatial location of the local structure. By extracting the phase information from each scale sub-band, a phase domain representation corresponding to that frame can be obtained. After obtaining the phase domain representation, the phase information corresponding to the continuous video frames is input into a phase amplification network for processing. This phase amplification network can employ a lightweight convolutional neural network structure. Its input is a sequence of phase changes from continuous frames. The network first performs frequency analysis on the phase sequence in the time dimension, for example, by extracting the spectral features of the phase changes through one-dimensional convolution or short-time Fourier transform, thereby identifying phase change components within a preset range of minute motion frequencies, such as low-frequency or mid-to-low-frequency change signals generated by minute gas leaks or subtle equipment vibrations. Subsequently, the network applies gain amplification processing to the corresponding phase changes according to a preset frequency range. That is, the phase change within the selected frequency range is multiplied by a preset amplification factor, thereby enhancing the originally small and easily observable weak phase changes, while keeping phase changes in other frequency ranges unchanged or suppressing them to avoid amplifying environmental noise or random disturbances. After phase amplification, the amplified phase information is recombined with the original amplitude information to form new complex wavelet coefficients. An inverse complex wavelet transform is then performed on these coefficients to restore the enhanced information in the frequency domain back to the image spatial domain, reconstructing phase-amplified image frames. All processed image frames are then combined in chronological order to form a phase-amplified video sequence. Through this process, minute movements, refraction changes, or local structural disturbances that are difficult to detect in the original video can be significantly enhanced, providing more obvious visual change information for subsequent extraction and identification of minor fault features.

[0039] Furthermore, a lightweight convolutional neural network is constructed as the phase amplification network. The phase amplification network learns sensitivity to different motion frequencies through training, so that phase changes that correspond to the frequency of the micro-leakage feature obtain the corresponding amplification coefficient.

[0040] Optionally, the phase amplification network employs a lightweight convolutional neural network architecture, including an input layer, a feature extraction layer, a frequency response learning layer, a gain prediction layer, and an output layer. The input layer receives phase change data obtained from a continuous video sequence after complex wavelet transform. The input data can be a phase difference sequence between adjacent frames or a phase temporal tensor stacked within a fixed time window. The feature extraction layer consists of multiple convolutional layers and nonlinear activation layers, used to extract local features of phase changes in both spatial and temporal dimensions. To reduce computational complexity, the convolutional layers can employ depthwise separable convolution, pointwise convolution, or small convolutional kernel structures. The frequency response learning layer learns the response intensities corresponding to different frequency components from the extracted temporal features to establish a correspondence between phase change frequencies and leakage features. The gain prediction layer outputs amplification coefficients corresponding to different frequency ranges based on the frequency response learning results. The output layer applies these amplification coefficients to the input phase change data to obtain the amplified phase change result. Subsequently, training samples are constructed to train the phase amplification network. These training samples can originate from actual collected monitoring videos of chemical equipment, simulated leakage experiment videos, and synthetic samples generated through data augmentation. For each training sample, a complex wavelet transform is first performed on the original video frames to extract phase information, and the phase change sequence between consecutive frames is calculated. Simultaneously, based on the leakage experiment annotation results or manual annotation results, the presence of micro-leakage and the corresponding frequency characteristic range of the leak are determined. During training, the phase change sequence is input into the phase amplification network. The network first extracts the local spatiotemporal features of the phase change through convolutional layers, and then analyzes the change patterns at different time scales through a frequency response learning layer. Specifically, the frequency response learning layer can encode components corresponding to different frequencies in the phase change sequence using one-dimensional temporal convolution, dilated convolution, or temporal pooling, enabling the network to distinguish between low-frequency slow changes, target micro-leakage frequency changes, and high-frequency random noise changes. Then, the gain prediction layer outputs amplification coefficients corresponding to each frequency component based on the encoded frequency characteristics. For phase changes that match the characteristic frequency of micro-leakage, a larger amplification coefficient is output; for changes deviating from the characteristic frequency range of micro-leakage, a smaller amplification coefficient or suppression coefficient is output, thus avoiding excessive amplification of irrelevant disturbances and noise.

[0041] Next, the training objective of the phase amplification network is set, for example, to maximize the enhancement of subtle motion or refraction perturbations associated with minute leaks after inverse transformation reconstruction of the amplified phase changes, while suppressing background noise and non-leaking motion. To this end, a joint loss function can be constructed to optimize the network. This joint loss function includes at least an amplification-reconstruction loss and a leak-sensitivity loss. The amplification-reconstruction loss constrains the spatial structure of the amplified phase changes after image reconstruction to maintain stability and avoid significant artifacts; the leak-sensitivity loss constrains the network to respond more strongly to phase changes within the target leak frequency range and less strongly to non-target frequency changes. During training, the network parameters are updated using backpropagation until the amplification coefficients output by the network stably reflect the matching relationship between different motion frequencies and minute leaks. After the network training is complete, the trained phase amplification network is deployed to the system. In actual operation, for the input phase domain representation, the network automatically analyzes the phase change components at different frequencies and outputs the corresponding frequency-sensitive amplification coefficients. This allows the phase changes that correspond to the characteristic frequencies of minor leaks to obtain higher gains. The amplified phase information is then combined with the original amplitude information, and the phase-amplified video sequence is reconstructed through inverse complex wavelet transform. Through this process, the slow airflow disturbances, background refraction distortions, or subtle vibrations caused by minor leaks in the original video can be made more apparent, thus providing stronger feature support for subsequent minor fault identification.

[0042] Furthermore, in the phase-magnified video sequence, static background regions are identified, the optical flow field of the background region in each frame is calculated, and then the second derivative of the optical flow field is calculated to generate a schlieren feature map sequence reflecting the texture distortion caused by gas refraction, including:

[0043] A static background region mask is extracted from the first frame of the continuous video sequence through semantic segmentation; for each frame in the phase-magnified video sequence, the divergence field and curl field of the dense optical flow field are calculated within the static background region mask; the spatial gradients of the divergence field and the curl field are calculated respectively to obtain the second derivative features; the second derivative features are low-pass filtered in the time dimension to generate the schlieren feature map sequence.

[0044] Optionally, the first frame of the continuous video sequence is first input into a pre-trained semantic segmentation model to perform pixel-level classification of different regions in the image. Based on the changes in the mask in subsequent frames, the model distinguishes the equipment body, pipe surface, equipment nameplate, support structure, ground background, and potentially moving non-background target regions. Then, regions with relatively fixed positions, stable textures, and the ability to reflect gas refraction disturbances during monitoring are selected from the segmentation results as static background regions, and corresponding static background region masks are generated. The static background region mask can be represented as a binary image with the same size as the original image. Pixels with a mask value of 1 represent static background pixels participating in subsequent optical flow calculations, while pixels with a mask value of 0 represent non-target regions not participating in the calculation. These masks typically remain unchanged in subsequent frames. Then, dense optical flow estimation is performed on the phase-magnified images of two adjacent frames within the mask-defined region to obtain the displacement vector field at each pixel position. This displacement vector field can be represented as... ,in, Let represent the horizontal displacement component of the pixel at time t. This represents the vertical displacement component of a pixel at time t. Dense optical flow calculation methods can employ the Farnebäck optical flow method, the TV-L1 optical flow method, or a deep learning-based optical flow estimation network. Since the phase-magnified video sequence has already enhanced the weak perturbations, the slight distortion of the background texture originally caused by a small gas leak will appear as a continuous, slow, and locally concentrated displacement change in the optical flow field.

[0045] After obtaining the dense optical flow field, its divergence field and curl field are further calculated. Specifically, for the optical flow field of each frame... Calculate the divergence: ; and curl: In this model, divergence reflects the degree of expansion or contraction of optical flow in a local region, characterizing the outward diffusion or inward convergence of background texture caused by gas leakage. Curl reflects the degree of rotation of optical flow in a local region, characterizing local torsional texture changes caused by airflow disturbance or inhomogeneous refraction. Since colorless gas leakage usually does not directly form a visible appearance, but changes the local air refractive index, causing slight bending, stretching, or rotation of static background texture, the divergence and curl fields can be used to highlight this type of schlieren phenomenon. Next, the spatial gradient is calculated by differentiating the divergence field in the horizontal and vertical directions, and then by differentiating the curl field in the same horizontal and vertical directions. This gradient is used as the second derivative feature, and can be discretized and approximated using the Sobel operator, Scharr operator, or differential convolution kernel. By further calculating the spatial gradients of the divergence and curl fields, the response at the local refractive distortion boundary can be enhanced, making the subtle texture distortion caused by small leaks appear as a more prominent high-response region in the feature map. Finally, the divergence field gradient and curl field gradient are combined according to preset weights to form the second-order derivative feature map corresponding to the current frame. These preset weights can be set based on experimental results or learned automatically through training. Next, the second-order derivative feature maps from multiple consecutive time points are arranged in chronological order, and the time-series signal at the same pixel location is low-pass filtered to suppress high-frequency random noise, instantaneous jitter, and non-continuous disturbances, retaining only the schlieren variation components that evolve slowly and continuously over time. This low-pass filtering can be implemented using moving average filtering, Gaussian time filtering, or first- or second-order digital low-pass filters. After low-pass filtering, the resulting schlieren feature map sequence can more stably reflect the continuous distortion features of the background texture caused by gas leakage and reduce spurious responses caused by camera noise, instantaneous illumination fluctuations, and occasional moving targets, thereby improving the detection sensitivity and recognition accuracy of refraction disturbances caused by minor leaks.

[0046] Furthermore, after outputting the first alarm message or the second alarm message, it also includes:

[0047] The system receives manual confirmation from operators regarding the first or second alarm information. If the confirmation result indicates a genuine fault, it automatically sends a linkage control command to the distributed control components of the chemical industrial park based on a preset linkage rule base, triggering the action of at least one emergency device, including emergency shut-off valves, sprinkler systems, and audible and visual alarms.

[0048] Preferably, after the system generates the first or second alarm message, the alarm message is first sent to the park's monitoring and management platform and displayed visually on the monitoring terminal interface. This alarm message may include the alarm type, alarm time, camera number, suspected fault area location, detection confidence level, and corresponding captured images or video clips. Operators view the alarm message and corresponding video footage through the monitoring terminal and manually confirm it based on the on-site situation. That is, operators mark the alarm event on the monitoring platform for confirmation. Manual confirmation results include two categories: genuine fault marking and false alarm marking. Specifically, when operators confirm that the equipment does indeed have a leak, open flame, smoke, or other abnormal state, the alarm event is marked as a genuine fault; when operators determine that the alarm is due to environmental interference, misidentification, or a non-fault condition, the alarm event is marked as a false alarm. After receiving the manual confirmation result, the system records the confirmation result and associates the alarm information with the manual confirmation result for storage. When the alarm result is confirmed as a false alarm, the system terminates the current alarm linkage process and only records the alarm event in the historical alarm database for subsequent statistical analysis and model optimization. When the alarm result is confirmed as a genuine fault, the system generates a corresponding linkage control strategy based on a pre-configured linkage rule base. This linkage rule base stores the correspondence between different fault types and emergency response actions. For example, when a leak is detected, an emergency shut-off valve is triggered to close the relevant pipeline; when a flame or smoke is detected, a sprinkler system is activated and an audible and visual alarm is triggered; when a serious anomaly is detected, multiple emergency response measures are executed simultaneously. After determining the corresponding linkage control strategy, the system sends linkage control commands to the distributed control components of the chemical industrial park through the industrial communication network. These distributed control components may include devices such as programmable logic controllers (PLCs), remote terminal units (RTUs), or distributed control systems (DCS). Upon receiving the linkage control command, the control component drives the corresponding actuator to act according to the command content, thereby triggering the activation of at least one emergency device. For example, closing the emergency shut-off valve to stop the continued delivery of the leaking medium, activating the sprinkler system to cool the equipment area or dilute the concentration of combustible gases, and activating the audible and visual alarm to warn on-site personnel. Through the above process, after the operator confirms the actual fault, the on-site safety equipment can be quickly activated, enabling timely handling and risk control of sudden faults in the chemical industrial park.

[0049] Furthermore, after outputting the first alarm message or the second alarm message, it also includes:

[0050] The alarm information, corresponding video clips, and manual confirmation results are associated and stored in the historical alarm database. Data is extracted from the historical alarm database according to a preset period and multidimensional statistical analysis is performed, including calculating the false alarm rate and false negative rate of each detection model under different operating conditions, identifying the scene patterns of frequent false alarms, and obtaining statistical analysis results. Based on the statistical analysis results, an incremental training dataset is constructed to incrementally learn the image recognition models used in the first fault detection and the second-level minor fault detection.

[0051] Optionally, after the system outputs the first or second alarm message, it first extracts and saves the structured and unstructured information related to the alarm event. The structured information includes the alarm number, alarm type, alarm time, camera number, device number, fault category, alarm confidence level, alarm area coordinates, environmental parameters at the time of alarm triggering, and the manual confirmation result provided by the operator. The unstructured information includes video clips, keyframe images, and feature maps or detection result maps corresponding to the alarm area within a preset time range before and after the alarm. Subsequently, the system establishes a unique event identifier for the same alarm event and associates the alarm information, corresponding video clips, and manual confirmation results according to the unique event identifier, storing them in the historical alarm database to form a traceable alarm event sample. After accumulating historical data, data is extracted from the historical alarm database for multidimensional statistical analysis according to a preset period, which can be set daily, weekly, or monthly. When extracting data, the system categorizes and summarizes historical alarm events according to dimensions such as model type, camera location, equipment category, operating conditions, and time interval. Operating conditions can include daytime or nighttime, sunny or rainy / foggy weather, normal operation, high-temperature operation, high-pressure operation, and maintenance / repair operation. For the image recognition models used in the first-level fault detection and the second-level minor fault detection, the system counts the total number of alarms, the number of times manually confirmed as real faults, and the number of times manually confirmed as false alarms under various operating conditions, and calculates the false alarm rate accordingly. Simultaneously, the system can also combine manually supplemented actual fault logs, maintenance records, or on-site inspection records to identify the number of events where the model did not trigger an alarm but actual faults existed, thereby calculating the missed alarm rate. Through the above statistical analysis, the performance of each detection model under different cameras, different equipment types, and different operating conditions can be obtained.

[0052] Subsequently, the system extracts relevant features from the false alarm samples, including image brightness, noise level, weather conditions, equipment background texture, motion interference type, reflective area distribution, shadow changes, and camera shake at the time of the alarm. It then uses clustering analysis, association rule analysis, or statistical ranking to identify high-frequency false alarm patterns. For example, it can identify scenarios where equipment nameplate reflections under low light conditions at night are easily misinterpreted as open flames, water stains on pipe surfaces after rain are easily misinterpreted as leaks, or swaying tree shadows in the background are easily misinterpreted as abnormal movement. The system outputs the frequently false alarm scenario patterns along with the corresponding sample data and statistical results to form statistical analysis results. After obtaining the statistical analysis results, it selects alarm samples that are manually confirmed as real faults from the historical alarm database as positive samples, and selects negative samples from the representative high-frequency false alarm samples that are manually confirmed as false alarms. It then combines data from different operating conditions, different equipment types, and different camera perspectives for balanced sampling to construct an incremental training dataset covering real fault scenarios and typical false alarm scenarios. For the image recognition model used in the first-level fault detection, the incremental training dataset can include significant fault samples such as open flames, dense smoke, large-scale leaks, and equipment deformation, as well as corresponding false alarm samples. For the image recognition model used in the second-level minor fault detection, the incremental training dataset can include video clips of minor leaks, corresponding spatiotemporal gradient cumulative feature maps, phase magnification video sequences, schlieren feature map sequences, and corresponding feature data in false alarm scenarios. Finally, while retaining the original model's basic parameters, the model is further trained using new samples, enhancing its adaptability to newly emerging fault modes and high-frequency false alarm scenarios without losing its original recognition capabilities. After training, the updated model is redeployed to the system for subsequent online detection. Through this process, the system can continuously correct model parameters and recognition boundaries based on historical alarm data and manual confirmation results, reducing false alarm and false negative rates, and improving its continuous recognition capability and scenario adaptability for equipment faults, especially minor faults, in chemical industrial parks.

[0053] In summary, the embodiments of this application have at least the following technical effects:

[0054] First, real-time video streams of equipment monitoring captured by cameras within the chemical industrial park are acquired. Next, a first fault detection is performed on the current image frame in the video stream. This first fault detection uses a preset conventional image recognition model to identify whether there are visually significant fault features in the current image frame. If so, a first alarm message is output. Finally, if the first fault detection does not identify visually significant fault features, a second-level minor fault detection is initiated for the continuous video sequence corresponding to the current image frame in the video stream. This involves performing temporal cumulative amplification analysis of the motion features of the continuous video sequence to extract subtle spatiotemporal change features. Based on these subtle change features, it is determined whether a minor fault exists. If so, a second alarm message is output. This solves the technical problem in existing image recognition-based equipment fault detection methods where minor faults, due to their extremely small pixel scale and frame-to-frame variations below noise levels, are difficult to effectively identify due to their weak visual features and slow changes. This achieves the technical effect of identifying early-stage minor faults in equipment by accumulating and amplifying subtle spatiotemporal changes on the basis of conventional visual detection, thereby improving fault detection sensitivity and early warning capabilities.

[0055] Example 2 is based on the same inventive concept as the image recognition-based equipment fault detection method in the aforementioned examples, such as... Figure 2 As shown, this application provides an image recognition-based equipment fault detection system for chemical industrial parks, wherein the system includes:

[0056] Video acquisition module 11: acquires real-time video streams of equipment monitoring captured by cameras within the chemical industrial park; First fault detection module 12: performs first fault detection on the current image frame in the equipment monitoring video stream. The first fault detection uses a preset conventional image recognition model to identify whether there are visually significant fault features in the current image frame. If so, it outputs a first alarm message; Second fault detection module 13: if the first fault detection does not identify visually significant fault features, it initiates a second-level minor fault detection for the continuous video sequence corresponding to the current image frame in the video stream. It performs time-accumulated amplification analysis of motion features on the continuous video sequence, extracts weak spatiotemporal change features, and determines whether there is a minor fault based on the weak change features. If so, it outputs a second alarm message.

[0057] Furthermore, the first fault detection module 12 is used to perform the following method:

[0058] The conventional image recognition model is a multi-task deep learning detection network, which includes multiple salient feature recognition branches. When the detection result of any branch exceeds the corresponding threshold, it is determined that there is a visually significant fault feature and the first alarm information is output.

[0059] Furthermore, the second fault detection module 13 is used to perform the following method:

[0060] The continuous video sequence is input into a long-term motion memory unit (LTM). The gradient of grayscale value over time is calculated for each pixel location, and the absolute values ​​of the gradients in the time dimension are accumulated to form a spatiotemporal gradient accumulation feature map. Each frame of the continuous video sequence is phase-magnified to reconstruct a phase-magnified video sequence. In the phase-magnified video sequence, static background regions are identified, and the optical flow field of the background region in each frame is calculated. The second derivative of the optical flow field is then calculated to generate a schlieren feature map sequence reflecting texture distortion caused by gas refraction. The spatiotemporal gradient accumulation feature map, the phase-magnified video sequence, and the schlieren feature map sequence are input into a temporal pooling layer. This temporal pooling layer uses an attention mechanism to adaptively select a time scale matching the current leakage rate and performs multimodal feature weighted fusion to obtain a comprehensive micro-fault feature tensor. The comprehensive micro-fault feature tensor is then input into a trained micro-fault classifier to complete the micro-fault identification.

[0061] Furthermore, the second fault detection module 13 is used to perform the following method:

[0062] For each pixel position of each frame in the continuous video sequence, the grayscale difference between adjacent frames is calculated; the long-term motion memory unit is set as a circular queue of length N, the queue is updated for each frame processed, and the grayscale difference of all frames in the current queue is accumulated to obtain a cumulative gradient map; the cumulative gradient map is normalized and contrast enhanced, and the spatiotemporal gradient cumulative feature map is output.

[0063] Furthermore, the second fault detection module 13 is used to perform the following method:

[0064] A complex wavelet transform is performed on each frame of the continuous video sequence to obtain a phase domain representation; the phase domain representation is input into a phase amplification network to amplify the phase change according to a preset small motion frequency range, and then the phase amplified video sequence is reconstructed by inverse transform.

[0065] Furthermore, the second fault detection module 13 is used to perform the following method:

[0066] A lightweight convolutional neural network is constructed as the phase amplification network. The phase amplification network learns sensitivity to different motion frequencies through training, so that the phase change that corresponds to the frequency of the small leakage feature obtains the corresponding amplification coefficient.

[0067] Furthermore, the second fault detection module 13 is used to perform the following method:

[0068] A static background region mask is extracted from the first frame of the continuous video sequence through semantic segmentation; for each frame in the phase-magnified video sequence, the divergence field and curl field of the dense optical flow field are calculated within the static background region mask; the spatial gradients of the divergence field and the curl field are calculated respectively to obtain the second derivative features; the second derivative features are low-pass filtered in the time dimension to generate the schlieren feature map sequence.

[0069] Furthermore, the second fault detection module 13 is used to perform the following method:

[0070] The system receives manual confirmation from operators regarding the first or second alarm information. If the confirmation result indicates a genuine fault, it automatically sends a linkage control command to the distributed control components of the chemical industrial park based on a preset linkage rule base, triggering the action of at least one emergency device, including emergency shut-off valves, sprinkler systems, and audible and visual alarms.

[0071] Furthermore, the second fault detection module 13 is used to perform the following method:

[0072] The alarm information, corresponding video clips, and manual confirmation results are associated and stored in the historical alarm database. Data is extracted from the historical alarm database according to a preset period and multidimensional statistical analysis is performed, including calculating the false alarm rate and false negative rate of each detection model under different operating conditions, identifying the scene patterns of frequent false alarms, and obtaining statistical analysis results. Based on the statistical analysis results, an incremental training dataset is constructed to incrementally learn the image recognition models used in the first fault detection and the second-level minor fault detection.

[0073] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for equipment fault detection in chemical industrial parks based on image recognition, characterized in that, include: Real-time acquisition of equipment monitoring video streams captured by cameras within the chemical industrial park; Perform a first fault detection on the current image frame in the monitoring video stream of the device. The first fault detection uses a preset conventional image recognition model to identify whether there are visually significant fault features in the current image frame. If there are, output a first alarm message. If the first fault detection fails to identify visually significant fault features, then for the continuous video sequence in the video stream corresponding to the current image frame, a second-level minor fault detection is initiated. The continuous video sequence is subjected to time-accumulated magnification analysis of motion features to extract weak spatiotemporal change features, and the presence of a minor fault is determined based on the weak change features. If a minor fault exists, a second alarm message is output.

2. The image recognition-based equipment fault detection method for chemical industrial parks as described in claim 1, characterized in that, The second level of minor fault detection is initiated by performing temporal cumulative amplification analysis of motion features on the continuous video sequence, extracting weak spatiotemporal variation features, and determining whether a minor fault exists based on the weak variation features, including: The continuous video sequence is input into the long-term motion memory unit, the gradient of gray value change with time is calculated for each pixel position, and the absolute value of the gradient in the time dimension is accumulated to form a spatiotemporal gradient accumulation feature map. Phase amplification is performed on each frame of the continuous video sequence to reconstruct the phase-amplified video sequence; In the phase-amplified video sequence, static background regions are identified, the optical flow field of the background region in each frame is calculated, and then the second derivative of the optical flow field is calculated to generate a schlieren feature map sequence that reflects the texture distortion caused by gas refraction. The spatiotemporal gradient cumulative feature map, the phase magnification video sequence, and the schlieren feature map sequence are input into the temporal pooling layer. The temporal pooling layer uses an attention mechanism to adaptively select a time scale that matches the current leakage rate and performs multimodal feature weighted fusion to obtain a comprehensive micro-fault feature tensor. The comprehensive micro-fault feature tensor is input into the trained micro-fault classifier to complete the micro-fault judgment.

3. The image recognition-based equipment fault detection method for chemical industrial parks as described in claim 2, characterized in that, The continuous video sequence is input into a long-term motion memory unit (LTM). The gradient of grayscale value over time is calculated for each pixel location, and the absolute values ​​of the gradients in the time dimension are accumulated to form a spatiotemporal gradient accumulation feature map, including: For each pixel position in each frame of the continuous video sequence, calculate the grayscale difference between adjacent frames; The long-term motion memory unit is set as a circular queue of length N. The queue is updated for each frame processed, and the grayscale difference of all frames in the current queue is accumulated to obtain the cumulative gradient map. The accumulated gradient map is normalized and its contrast is enhanced to output the spatiotemporal gradient accumulated feature map.

4. The image recognition-based equipment fault detection method for chemical industrial parks as described in claim 2, characterized in that, Phase magnification is performed on each frame of the continuous video sequence to reconstruct a phase-magnified video sequence, including: Perform a complex wavelet transform on each frame of the continuous video sequence to obtain a phase domain representation; The phase domain representation is input into the phase amplification network to amplify the phase change according to a preset range of minute motion frequencies, and then the phase amplification video sequence is reconstructed by inverse transformation.

5. The image recognition-based equipment fault detection method for chemical industrial parks as described in claim 4, characterized in that, A lightweight convolutional neural network is constructed as the phase amplification network. The phase amplification network learns sensitivity to different motion frequencies through training, so that the phase change that corresponds to the frequency of the small leakage feature obtains the corresponding amplification coefficient.

6. The image recognition-based equipment fault detection method for chemical industrial parks as described in claim 2, characterized in that, In the phase-magnified video sequence, static background regions are identified, the optical flow field of the background region in each frame is calculated, and then the second derivative of the optical flow field is calculated to generate a schlieren feature map sequence reflecting the texture distortion caused by gas refraction, including: Extract a static background region mask from the first frame of the continuous video sequence using semantic segmentation; For each frame in the phase-amplified video sequence, the divergence field and curl field of the dense optical flow field are calculated within the static background region mask; The spatial gradients of the divergence field and the curl field are calculated respectively to obtain the second derivative characteristics; The second derivative features are low-pass filtered in the time dimension to generate the schlieren feature map sequence.

7. The image recognition-based equipment fault detection method for chemical industrial parks as described in claim 1, characterized in that, The conventional image recognition model is a multi-task deep learning detection network, which includes multiple salient feature recognition branches; When the detection result of any branch exceeds the corresponding threshold, it is determined that there is a visually significant fault feature and the first alarm information is output.

8. The method for equipment fault detection in chemical industrial parks based on image recognition as described in claim 1, characterized in that, After outputting the first alarm message or the second alarm message, the following is also included: Receive manual confirmation from the operator regarding the first or second alarm information; If the result is confirmed to be a genuine fault, then according to the preset linkage rule library, the linkage control command is automatically sent to the distributed control component of the chemical industrial park to trigger the action of at least one emergency device, including emergency shut-off valve, sprinkler system, and audible and visual alarm.

9. The image recognition-based equipment fault detection method for chemical industrial parks as described in claim 8, characterized in that, include: The alarm information, the corresponding video clip, and the manual confirmation result are associated and stored in the historical alarm database; According to a preset cycle, data is extracted from the historical alarm database and multidimensional statistical analysis is performed, including calculating the false alarm rate and false alarm rate of each detection model under different working conditions, identifying the scene patterns of frequent false alarms, and obtaining statistical analysis results. Based on the statistical analysis results, an incremental training dataset is constructed to incrementally learn the image recognition model used in the first fault detection and the second-level minor fault detection.

10. A chemical industrial park equipment fault detection system based on image recognition, characterized in that, The method for detecting equipment faults in chemical industrial parks based on image recognition, as described in any one of claims 1-9, comprises: Video acquisition module: Real-time acquisition of equipment monitoring video streams captured by cameras within the chemical industrial park; First fault detection module: Performs first fault detection on the current image frame in the monitoring video stream of the device. The first fault detection uses a preset conventional image recognition model to identify whether there are visually significant fault features in the current image frame. If there are, it outputs first alarm information. Second fault detection module: If the first fault detection fails to identify visually significant fault features, then for the continuous video sequence in the video stream corresponding to the current image frame, a second-level minor fault detection is initiated. The continuous video sequence is subjected to time-accumulated magnification analysis of motion features to extract weak spatiotemporal change features, and the presence of minor faults is determined based on the weak change features. If a minor fault exists, a second alarm message is output.