Intelligent filling process operation monitoring method based on video analysis
By integrating multimodal video data stream acquisition and cross-modal attention mechanisms, and combining time synchronization and spatial registration processing, the problem of low recognition accuracy in complex environments of existing video surveillance systems is solved, realizing intelligent monitoring of filling operations and ensuring safe production and process compliance.
Patent Information
- Application Number
- CN202511091970.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-28
AI Technical Summary
Existing video surveillance systems widely used at filling operation sites lack the ability to intelligently identify operational behavior, equipment status, and process integrity in real time. In particular, the identification accuracy is low in complex environments, making it difficult to meet the requirements of safe production and process compliance.
Multimodal video data stream acquisition is adopted, and multimodal feature maps are interactively fused through cross-modal attention mechanism. Combined with time synchronization and spatial registration processing, local visual features are extracted using convolutional neural network to generate multi-scale feature maps, identify personnel posture, equipment status and operation actions in the work area, and construct work process state transition map for verification.
It enables comprehensive and multi-dimensional visual perception of filling operations in complex environments, improving recognition accuracy and robustness. It can automatically identify the compliance of operators and the status of equipment, ensuring process integrity, reducing labor costs and improving supervision efficiency.
Smart Images

Figure CN121033752A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of filling process operation monitoring, in particular to a smart filling process operation monitoring method based on video analysis. BACKGROUND
[0002] With the continuous improvement of industrial automation level, the filling operation of high-risk substances such as dangerous goods, liquefied gas and chemical raw materials is gradually developing towards intelligence and unmanned. In this process, how to ensure the safety of the operators and the standardized execution of the operation process has become one of the key technical problems in the field of industrial safety supervision.
[0003] At present, the video monitoring system widely used in filling operation sites is mainly composed of traditional single-mode camera equipment (such as visible light camera), whose functions are mainly concentrated in image acquisition and post-playback of the operation area, lacking real-time intelligent recognition ability for operation behavior, equipment state and process integrity. Such systems usually rely on manual duty or simple motion detection algorithm for abnormal judgment, which has high false alarm rate, low recognition accuracy and lagging response, etc., and is difficult to meet the growing needs of modern industry in safety production, process compliance and process traceability.
[0004] In addition, most of the existing video monitoring systems only focus on single visual information, which cannot meet the identification challenges in complex environments. For example, in poor lighting conditions, smoke interference, night operation and other scenes, the quality of visible light images decreases, leading to target recognition failure; at the same time, the lack of perception ability for key parameters such as temperature change and object distance also makes the system have blind spots in judging gas leakage, personnel approaching dangerous areas and other safety hazards. SUMMARY
[0005] The present application aims to solve the above problems, and proposes a smart filling process operation monitoring method based on video analysis, which integrates multi-modal video data stream acquisition, interactive fusion of multi-modal feature maps through cross-modal attention mechanism, and multi-modal video data stream acquisition and spatial registration processing.
[0006] The present application specifically adopts the following technical solutions:
[0007] A smart filling process operation monitoring method based on video analysis, specifically comprising the following steps:
[0008] S1: acquiring multi-modal video data stream of the filling operation area, the multi-modal video data stream comprising a visible light image sequence, an infrared thermal imaging image sequence and a depth image sequence;
[0009] S2: performing time synchronization and spatial registration processing on each modal image sequence to generate a synchronous multi-modal image frame group;
[0010] S3: Extract local visual features of each modality image frame based on a convolutional neural network, and obtain multi-scale visual feature maps through a feature pyramid structure;
[0011] S4: Interactively fuse multimodal feature maps using a cross-modal attention mechanism to generate a unified fused visual feature representation;
[0012] S5: Input the fused visual features into the target detection module to identify personnel posture, equipment status and operating actions within the work area;
[0013] S6: Construct a job flow state transition diagram based on the identified action sequence and equipment status information;
[0014] S7: Match and verify the current job process status according to the preset standard process path;
[0015] S8: If the current process status is not executed in the standard order or an abnormal status occurs, an early warning signal will be triggered and output to the control terminal.
[0016] Preferably, in S2, the synchronization processing during time synchronization and spatial registration utilizes a hardware clock synchronization mechanism to ensure that the video streams acquired by multiple cameras are aligned at the frame level, and uses timestamps to mark each frame. During spatial registration, a calibration board is used to determine the spatial coordinate mapping relationship between different modal cameras. Then, based on the calibration results, an affine transformation is used to geometrically correct the misaligned images, so that the images of each modality are presented in the same spatial coordinate system. The affine transformation is represented by the following matrix form:
[0017]
[0018] Where a, b, c, and d control scaling, rotation, and shearing operations, respectively, and t x ,t y These are the translation parameters.
[0019] Preferably, the acquisition of multi-scale visual feature maps in S3 includes: setting multiple convolutional layers with different receptive fields in the convolutional neural network, using a feature pyramid network (FPN) to fuse shallow detail features with deep semantic features, and outputting multi-scale feature maps with high resolution and strong semantic expressive power for subsequent detection tasks. The specific process is as follows:
[0020] Design a multi-layer convolutional neural network, each layer including a convolutional layer, an activation function, and a pooling layer;
[0021] For each input modality image sequence, local visual features are extracted through forward propagation, and the convolution operation is represented as:
[0022] z (l) =f(W (l) *x(l-1) +b (l) )
[0023] Among them, W (l) ,b (l) These are the weights and biases of the l-th layer, x (l-1) It is the output of the previous layer, and f is the activation function;
[0024] The Feature Pyramid Network generates multi-scale feature maps by combining deep semantic information and shallow detail information. The specific process is as follows:
[0025] (1) Set up convolutional layers with different receptive fields at different levels of the network;
[0026] (2) Upsample the high-level semantic feature map and fuse it with the corresponding low-level feature map to enhance the detection capability of small targets;
[0027] (3) Use weighted summation or splicing to combine these features to form a feature pyramid containing rich details and strong semantic information.
[0028] Preferably, the cross-modal attention mechanism in S4 includes: constructing an intermodal correlation matrix, calculating the attention weights between different modal features, fusing the features of each modality using a weighted summation method, and introducing a channel attention mechanism to enhance and suppress the channels of the fused feature map. The specific process is as follows:
[0029] (1) First, construct an intermodal correlation matrix to calculate the similarity or correlation between different modal features;
[0030] (2) Adjust the importance of each modality feature based on the calculated relevance score as the attention weight;
[0031] (3) The modal features are fused using a weighted summation method, i.e.:
[0032]
[0033] Among them, w i F is the attention weight of the i-th mode. i This is the feature map of this mode.
[0034] The fused feature map is further optimized through an attention mechanism, which enhances useful information and suppresses noise by learning the importance of each channel.
[0035] Preferably, the object detection module in S5 includes: using a YOLOv5 or Faster R-CNN model to perform object detection on the fused feature map, and outputting object category labels, bounding box coordinates, and confidence scores.
[0036] Preferably, the construction of the job process state transition diagram in S6 includes: defining several nodes in the standard job process, each node representing a job stage, each node containing the corresponding target detection result and action recognition result, and the nodes being connected by directed edges to represent a legal job process transition path.
[0037] Preferably, the process state matching verification in S7 includes: using breadth-first search (BFS) to traverse the state transition graph to find whether the current process exists in the standard path set. If the current path does not exist in the standard path set and deviates from the set threshold number of steps, it is determined that the process is abnormal; otherwise, the current job path is continued to be tracked until the process is completed.
[0038] The present invention has the following beneficial effects:
[0039] Through the design of multimodal video data stream acquisition, visible light images, infrared thermal imaging images and depth images can be acquired and fused for analysis simultaneously, realizing a comprehensive and multi-dimensional visual perception capability of the work site, and achieving the technical effect of being able to stably identify targets and abnormal states even in complex lighting, smoke, obstruction and other environments.
[0040] By designing time synchronization and spatial registration processing for image sequences of various modalities, video frames acquired by different camera devices can be kept consistent in both time and space, achieving accurate alignment and collaborative analysis of multi-source information, and thus improving the technical effect of subsequent image fusion accuracy and recognition accuracy.
[0041] By extracting local visual features based on convolutional neural networks and combining them with feature pyramid structures to generate multi-scale feature maps, the system can simultaneously take into account image details and high-level semantic understanding, achieving high-precision recognition of small targets (such as pressure gauge pointers and glove wear), thus improving the overall robustness and applicability of recognition.
[0042] By using a cross-modal attention mechanism to interactively fuse multimodal feature maps, the system can adaptively adjust the weights of each modality feature according to the current environment, achieving a better feature fusion strategy and noise suppression capability, and achieving the technical effect of maintaining good recognition performance even when a single modality fails.
[0043] By integrating visual features into the target detection module to identify personnel posture, equipment status, and operational actions, the system can automatically identify whether operators are wearing protective equipment in compliance with regulations and whether the equipment is in normal working condition. This achieves intelligent identification of key objects and behaviors in filling operations, thereby improving regulatory efficiency and reducing labor costs.
[0044] By constructing a state transition diagram of the work process based on the identified action sequence and equipment status information, the system can model and track the entire filling operation process, realize the visual expression and logical verification of the operation path, and achieve the technical effect of judging the integrity of the process and preventing skipping or misoperation.
[0045] By combining multimodal video data stream acquisition with spatial registration processing, infrared images can provide clear target outlines in environments with extremely poor lighting or smoke interference. With spatial registration, even if the visible light image is blurred, target positioning and identification can still be achieved. This can be used for continuous monitoring in extreme scenarios such as nighttime operations and high-temperature leaks, far exceeding the adaptability of traditional single-modal systems.
[0046] By fusing features through a cross-modal attention mechanism, when the data quality of a certain modality deteriorates (such as a depth camera failure), the system automatically increases the weight of features from other modalities to maintain the stability of overall recognition performance. This improves the system's fault tolerance and adaptability, making it suitable for front-line industrial sites where equipment is easily damaged and the environment is complex. Attached Figure Description
[0047] Figure 1 This is a flowchart illustrating a smart filling process monitoring method based on video analytics. Detailed Implementation
[0048] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and specific examples:
[0049] Combination Figure 1 A smart filling process operation monitoring method based on video analytics specifically includes the following steps:
[0050] S1: Acquire multimodal video data streams of the filling operation area. The multimodal video data streams include visible light image sequences, infrared thermal imaging image sequences, and depth image sequences.
[0051] S2: Perform time synchronization and spatial registration processing on the image sequences of each modality to generate synchronized multimodal image frame groups. The time synchronization and spatial registration processing includes: using a hardware clock synchronization mechanism to perform frame-level alignment of the video streams acquired by multiple cameras; using a calibration board to calibrate the spatial coordinate mapping relationship between different modal cameras; and using an affine transformation algorithm to perform geometric correction on the misaligned images, so that all modal images are presented in the same spatial coordinate system. The synchronization processing during time synchronization and spatial registration uses a hardware clock synchronization mechanism to ensure that the video streams acquired by multiple cameras are aligned at the frame level, using timestamps to mark each frame to ensure they can be correctly aligned to the same time point. During spatial registration, a calibration board is used to determine the spatial coordinate mapping relationship between different modal cameras. Then, based on the above calibration results, an affine transformation is used to perform geometric correction on the misaligned images, so that all modal images are presented in the same spatial coordinate system. The affine transformation can be represented in the following matrix form:
[0052]
[0053] Where a, b, c, and d control scaling, rotation, and shearing operations, t x ,t y It is the translation parameter.
[0054] S3: Local visual features of each modality image frame are extracted based on a convolutional neural network, and multi-scale visual feature maps are obtained through a feature pyramid structure. The acquisition of multi-scale visual feature maps includes: setting multiple convolutional layers with different receptive fields in the convolutional neural network; using a feature pyramid network (FPN) to fuse shallow detail features with deep semantic features; and outputting multi-scale feature maps with high resolution and strong semantic expressive power for subsequent detection tasks. The specific process of the convolutional neural network design is as follows:
[0055] (1) Design a multi-layer convolutional neural network, each layer including components such as convolutional layer, activation function (such as ReLU), pooling layer, etc.
[0056] (2) For each input modality image sequence, local visual features are extracted through forward propagation. For example, at a certain layer, the convolution operation can be represented as:
[0057] z (l) =f(W (l) *x (l-1) +b (l) )
[0058] Here W (l) ,b (l) These are the weights and biases of the l-th layer, x (l-1) It is the output of the previous layer, and f is the activation function.
[0059] The Feature Pyramid Network generates multi-scale feature maps by combining deep semantic information and shallow detail information. The specific process is as follows:
[0060] (1) Set up convolutional layers with different receptive fields at different levels of the network.
[0061] (2) Upsample the high-level semantic feature map and fuse it with the corresponding low-level feature map to enhance the detection capability of small targets.
[0062] (3) Use weighted summation or splicing to combine these features to form a feature pyramid containing rich details and strong semantic information.
[0063] S4: Interactive fusion of multimodal feature maps is performed using a cross-modal attention mechanism to generate a unified fused visual feature representation. The cross-modal attention mechanism in S4 includes: constructing an intermodal correlation matrix, calculating attention weights between features of different modalities, fusing the features using a weighted summation method, and introducing a channel attention mechanism to enhance and suppress channels in the fused feature map. The specific process of the cross-modal attention mechanism is as follows:
[0064] (1) First, construct an intermodal correlation matrix to calculate the similarity or correlation between different modal features.
[0065] (2) The importance of each modality feature is adjusted based on the calculated relevance score as the attention weight.
[0066] (3) The modal features are fused using a weighted summation method, i.e.:
[0067]
[0068] Among them, w i F is the attention weight of the i-th mode. i This is the feature map of this mode.
[0069] The fused feature map is further optimized through an attention mechanism, which enhances useful information and suppresses noise by learning the importance of each channel.
[0070] S5: Input the fused visual features into the target detection module to identify personnel posture, equipment status and operation actions within the work area; the target detection module in S5 includes: using YOLOv5 or Faster R-CNN models to perform target detection on the fused feature map, and outputting target category labels, bounding box coordinates and confidence scores.
[0071] S6: Construct a job process state transition diagram based on the identified action sequence and equipment status information; S6 includes constructing a job process state transition diagram by defining several nodes in the standard job process, each node representing a job stage, each node containing the corresponding target detection result and action recognition result, and nodes being connected by directed edges to represent legal job process transition paths.
[0072] S7: Match and verify the current job process status according to the preset standard process path; The process status matching and verification in S7 includes: using breadth-first search (BFS) to traverse the state transition graph to find whether the current process exists in the standard path set. If the current path does not exist in the standard path set and deviates from the set threshold number of steps, it is determined that the process is abnormal. Otherwise, continue to track the current job path until the process is completed.
[0073] S8: If the current process status is not executed in the standard order or an abnormal status occurs, an early warning signal will be triggered and output to the control terminal.
[0074] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A smart filling process operation monitoring method based on video analytics, characterized in that, Specifically, the following steps are included: S1: Acquire multimodal video data streams of the filling operation area, including visible light image sequences, infrared thermal imaging image sequences, and depth image sequences; S2: Perform time synchronization and spatial registration processing on each modal image sequence to generate a synchronized multimodal image frame group; S3: Extract local visual features of each modality image frame based on a convolutional neural network, and obtain multi-scale visual feature maps through a feature pyramid structure; S4: Interactively fuse multimodal feature maps using a cross-modal attention mechanism to generate a unified fused visual feature representation; S5: Input the fused visual features into the target detection module to identify personnel posture, equipment status and operating actions within the work area; S6: Construct a job flow state transition diagram based on the identified action sequence and equipment status information; S7: Match and verify the current job process status according to the preset standard process path; S8: If the current process status is not executed in the standard order or an abnormal status occurs, an early warning signal will be triggered and output to the control terminal.
2. The intelligent filling process operation monitoring method based on video analysis as described in claim 1, characterized in that, In S2, the synchronization processing during time synchronization and spatial registration uses a hardware clock synchronization mechanism to ensure that the video streams captured by multiple cameras are aligned at the frame level, and uses timestamps to mark each frame. During spatial registration, a calibration board is used to determine the spatial coordinate mapping relationship between cameras of different modalities. Then, based on the calibration results, affine transformation is used to geometrically correct the misaligned images, so that the images of each modality are presented in the same spatial coordinate system. The affine transformation is represented by the following matrix form: Where a, b, c, and d control scaling, rotation, and shearing operations, respectively, and t x ,t y These are the translation parameters.
3. The intelligent filling process operation monitoring method based on video analysis as described in claim 1, characterized in that, The acquisition of multi-scale visual feature maps in S3 involves setting up multiple convolutional layers with different receptive fields in a convolutional neural network, using a feature pyramid network (FPN) to fuse shallow detail features with deep semantic features, and outputting multi-scale feature maps with high resolution and strong semantic expressive power for subsequent detection tasks. The specific process is as follows: Design a multi-layer convolutional neural network, each layer including a convolutional layer, an activation function, and a pooling layer; For each input modality image sequence, local visual features are extracted through forward propagation, and the convolution operation is represented as: z (l) =f(W (l) *x (l-1) +b (l) ) Among them, W (l) ,b (l) These are the weights and biases of the l-th layer, x (l-1) It is the output of the previous layer, and f is the activation function; The Feature Pyramid Network generates multi-scale feature maps by combining deep semantic information and shallow detail information. The specific process is as follows: (1) Set up convolutional layers with different receptive fields at different levels of the network; (2) Upsample the high-level semantic feature map and fuse it with the corresponding low-level feature map to enhance the detection capability of small targets; (3) Use weighted summation or splicing to combine these features to form a feature pyramid containing rich details and strong semantic information.
4. The intelligent filling process operation monitoring method based on video analysis as described in claim 1, characterized in that, The cross-modal attention mechanism in S4 includes: constructing an intermodal correlation matrix, calculating the attention weights between features of different modalities, fusing the features of each modality using a weighted summation method, and introducing a channel attention mechanism to enhance and suppress channels in the fused feature map. The specific process is as follows: (1) First, construct an intermodal correlation matrix to calculate the similarity or correlation between different modal features; (2) Adjust the importance of each modality feature based on the calculated relevance score as the attention weight; (3) The modal features are fused using a weighted summation method, i.e.: Among them, w i F is the attention weight of the i-th mode. i This is the feature map of this mode. The fused feature map is further optimized through an attention mechanism, which enhances useful information and suppresses noise by learning the importance of each channel.
5. The intelligent filling process operation monitoring method based on video analysis as described in claim 1, characterized in that, The object detection module in S5 includes: using YOLOv5 or Faster R-CNN models to perform object detection on the fused feature map, and outputting object class labels, bounding box coordinates, and confidence scores.
6. The intelligent filling process operation monitoring method based on video analysis as described in claim 1, characterized in that, The process of constructing a state transition graph in S6 includes: defining several nodes in the standard work process, each node representing a work stage, each node containing the corresponding target detection result and action recognition result, and nodes being connected by directed edges to represent legal work process transition paths.
7. The intelligent filling process operation monitoring method based on video analysis as described in claim 1, characterized in that, In S7, process status matching verification includes: using breadth-first search (BFS) to traverse the state transition graph to check if the current process exists in the standard path set. If the current path does not exist in the standard path set and deviates from the set threshold number of steps, it is determined to be a process abnormal; otherwise, the current job path is tracked until the process is completed.