A multimodal wellsite video safety analysis method based on attention
By adopting attention-based multimodal analysis method in well site video safety analysis, combining space-time and multi-scale spatial attention mechanisms, the problem of insufficient accuracy in identifying dynamic information and dangerous actions in the existing technology is solved, and more efficient and reliable safety analysis is achieved.
Patent Information
- Application Number
- CN202510323882.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The existing well site video image safety analysis algorithm is insufficient in identifying dynamic information and dangerous actions, has low real-time performance, and cannot fully monitor the well site environment, resulting in misjudgment and omissions.
The attention-based multi-modal well-field video safety analysis method is adopted to extract the low-level fusion characteristics of video frames through an image encoder, combine the spatiotemporal attention mechanism and the multi-scale spatial attention feature extraction module to carry out deep fusion, generate multi-modal fusion characteristics, and finally input into the multi-task decoder for security analysis.
It significantly improves the accuracy and reliability of well site video safety analysis, can more effectively identify dynamic information and small actions, reduce misjudgment, and improve real-time and comprehensiveness.
Smart Images

Figure CN119851185B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent video surveillance and analysis, and particularly to a multi-modal well site video safety analysis method based on attention. Background Art
[0002] As an important place for oil extraction, the safety management and efficient operation of the oil well site are particularly important. However, the operating environment of the oil well site is complex and changeable, involving a large number of mechanical equipment, flammable and explosive substances, and high-risk operation processes, with many potential safety hazards. The traditional safety management method of the oil well site mainly relies on manual inspections and monitoring, which has problems such as management blind spots and slow response speed, and is difficult to meet the safety management requirements of modern oil well sites.
[0003] In recent years, with the rapid development of the Internet of Things (IoT), artificial intelligence (AI), and big data technologies, intelligent oil well scene video surveillance systems have emerged. This system deploys high-definition cameras and edge computing devices at the oil well site, and combines advanced video image analysis algorithms to identify and analyze potential violations, realizing all-round and all-weather monitoring of the oil well site. Existing well site video image safety analysis algorithms mainly focus on two aspects: object recognition-based and behavior posture recognition-based. Object recognition-based methods, such as the YOLO algorithm, identify and locate workers and other key objects in the video based on single-frame images of the well site monitoring video, and analyze current potential safety hazards according to static information such as the presence and location of the objects. Behavior posture recognition-based methods, such as the temporal convolutional algorithm, extract rich dynamic information by using tracking and positioning technologies in multiple frames of images for workers in the video, and identify dangerous actions and behaviors of the workers.
[0004] However, there are many problems with existing safety analysis algorithms, making it difficult to meet the requirements of the actual complex well site environment. Object recognition-based methods only use single-frame images, so the algorithm cannot identify dynamic information, cannot recognize dangerous actions and behaviors, and due to the information limitation brought by single-frame images, the probability of misjudgment of the algorithm is high, which may lead to the neglect of potential safety hazards and false alarms. Behavior posture recognition-based methods have limited behaviors, a high probability of misjudgment, and due to the single nature of the algorithm tasks, the potential safety hazards they focus on are relatively limited and cannot meet the needs of overall well site safety analysis. In addition, due to the complex well site environment, the above algorithms analyze the overall image, which is likely to result in the omission of monitoring small targets and tiny actions. Pose recognition-based methods only focus on the dynamic information of people, cannot judge potential safety hazards in the environment, and due to the need for multi-frame processing, the algorithm has an inherent response time, which will cause a certain delay when facing sudden potential safety hazards. The above two types of algorithms do not make full use of the static information of single frames and the dynamic information of multiple frames of the video to extract information. Summary of the Invention
[0005] In view of the above deficiencies in the prior art, the present invention provides an attention-based multi-modal wellsite video safety analysis method, which solves the problems of insufficient accuracy and low real-time performance in safety monitoring in the prior art.
[0006] To achieve the above invention purpose, the technical solution adopted by the present invention is: an attention-based multi-modal wellsite video safety analysis method, including the following steps:
[0007] S1. Send the frame sequence in the wellsite operation video to be analyzed into a pre-trained image encoder to obtain the low-level fusion features of the video frames;
[0008] S2. Input the obtained low-level fusion features into a video global feature extraction module based on a spatio-temporal attention mechanism to obtain the global spatio-temporal features of the video modality, that is, temporal information and spatial information;
[0009] S3. Input the key frames in the frame sequence described in S1 into a feature extraction module based on multi-scale spatial attention to obtain the pixel-level fine-grained local features of the single-frame image modality;
[0010] S4. Input the extracted global spatio-temporal features of the video modality and the pixel-level fine-grained local features of the single-frame image modality into a multi-modal feature progressive fusion module for deep fusion to obtain multi-modal fusion features;
[0011] S5. Input the multi-modal fusion features into a multi-task decoder module based on multi-modal features to obtain the safety analysis results of the wellsite operation video, including the safety hazard analysis results of three categories: the worker himself, the worker and the environment, and the environment;
[0012] Among them, the image encoder, the video global feature extraction module based on the spatio-temporal attention mechanism, the multi-modal feature progressive fusion module, the multi-task decoder module based on multi-modal features, and the feature extraction module based on multi-scale spatial attention constitute the safety analysis model.
[0013] Furthermore: The image encoder includes a pre-trained vision encoder based on SAM and a pre-trained vision encoder based on CLIP; Take the wellsite operation video frames to be analyzed as the input of the image encoder, and input them into the pre-trained vision encoder based on SAM and the pre-trained vision encoder based on CLIP respectively to obtain two types of features regarding the wellsite operation video; Then perform feature fusion on the obtained two types of features to obtain the low-level fusion features of all video frames, and its expression is:
[0014]
[0015] Among them is the input of the image encoder; is the function representation of the CLIP vision encoder, are the parameters of the CLIP visual encoder, is the output of the CLIP visual encoder; is the functional representation of the SAM visual encoder, are the parameters of the SAM visual encoder, is the output of the SAM visual encoder; H is the low-level fusion feature of all the finally obtained video frames , are the weights corresponding to the SAM visual encoder, is CL the weights corresponding to the IP visual encoder.
[0016] Furthermore: The video global feature extraction module based on the spatio-temporal attention mechanism includes a small kernel depth convolution layer, a multi-head temporal attention mechanism unit, a dilated depth convolution layer, and a 1×1 convolution layer connected in sequence; The video global feature extraction module based on the spatio-temporal attention mechanism further includes a first temporal self-attention mechanism unit, a first average pooling layer, and a first fully connected layer connected in sequence; The low-level fusion feature obtained in step S1 is respectively used as the input of the small kernel depth convolution unit and the first temporal self-attention mechanism unit; The output of the 1×1 convolution layer and the output of the first fully connected layer are subjected to a Kronecker product operation and then a Hadamard product operation with the low-level fusion feature obtained in step S1 to obtain the global spatio-temporal feature of the video modality, and its calculation expression is:
[0017]
[0018] Among them, represents the output of the 1×1 convolution layer, represents a 1×1 convolution, represents a dilated depth convolution, represents a multi-head temporal attention process, represents a small kernel depth convolution, is the low-level fusion feature obtained in step S1; represents the output of the first fully connected layer, represents a fully connected layer, represents a global average pooling, represents a temporal self-attention process; is the output of the video global feature extraction module based on the spatio-temporal attention mechanism, represents a Kronecker product, represents a Hadamard product.
[0019] Furthermore, in step S3, the feature extraction module based on multi-scale spatial attention extracts features of five scales from the key-frame image and adopts fast normalization fusion to obtain multi-scale spatial attention features; the five scales are respectively the original image scale of the key-frame image, one-half of the original image scale, one-fourth of the original image scale, one-eighth of the original image scale, and one-sixteenth of the original image scale.
[0020] Furthermore, the feature extraction module based on multi-scale spatial attention includes a VIT unit, a 2D convolutional unit, a dilated convolutional layer, a multi-head cross-attention mechanism unit, a self-attention mechanism unit, and a fast normalization feature fusion unit connected in sequence; the key-frame image is the input of the feature extraction module based on multi-scale spatial attention, and the output of the fast normalization feature fusion unit is the pixel-level fine-grained local feature of the single-frame image modality; its calculation expression is:
[0021]
[0022]
[0023]
[0024] where is the key-frame image at the i th scale, represents processing, represents 2D convolution, represents dilated convolution, represents the preliminary feature extracted from the key-frame image at the i th scale; represents multi-head cross-attention processing, represents self-attention processing, represents the output of the self-attention mechanism unit; represents the output of the fast normalization feature fusion unit, represents corresponding weight, represents the learning weight corresponding to the j th output of the self-attention mechanism unit, used to normalize ; is an additional parameter to ensure the stability of the result.
[0025] Furthermore, the multi-modal feature progressive fusion module includes a first cross-attention mechanism unit, a multi-layer perceptron, a second average pooling layer, and a separable convolutional layer connected in sequence, and also includes a second cross-attention mechanism unit and a temporal modeling unit connected in sequence; the output of the temporal modeling unit serves as another input to the separable convolutional layer; the inputs of the first cross-attention mechanism unit and the second cross-attention mechanism unit serve as the inputs of the multi-modal feature progressive fusion module; the output of the separable convolutional layer serves as the output of the multi-modal feature progressive fusion module.
[0026] Furthermore, the multi-task decoder module based on multi-modal features includes a first linear layer group, a third average pooling layer, a flattening layer, a second linear layer group, and a second fully connected layer connected in sequence.
[0027] Furthermore, the specific steps for training the safety analysis model include:
[0028] A1. Obtain the wellsite operation video and label the corresponding safety hazards in the wellsite operation video to obtain the training set;
[0029] A2. Send the videos in the training set into the safety analysis model to obtain the safety analysis results of the wellsite operation video, including the safety hazard analysis results of three categories: the workers themselves, the workers and the environment, and the environment;
[0030] A3. Calculate the joint loss according to the safety analysis results of the wellsite operation video. If the joint loss no longer decreases or reaches the set number of training times, end the training; otherwise, return to step A2;
[0031] Among them, the joint loss The calculation expression is:
[0032]
[0033]
[0034]
[0035]
[0036] Among them, represents the loss of judging the hidden danger of workers' violations, represents the number of samples, represents the total number of categories of hidden dangers of workers' violations, represents the th sample contains the hidden danger probability, is the prediction index. When the hidden danger type label in the th sample output by the safety analysis model is the same as the true hidden danger type label is consistent is 1, otherwise is 0; represents the loss of judging the relationship between workers and environmental safety, D represents the total number of categories of the relationship between workers and environmental safety; represents the loss of judging potential hazards in the environment, G represents the total number of categories of potential hazards in the environment; represents a hyperparameter used for weighting various losses; ln represents the natural logarithm.
[0037] Furthermore: When using the multi-modal feature progressive fusion module in step S4 to fuse the global spatio-temporal features of the video modality and the pixel-level fine-grained local features of the single-frame image modality, feature fusion is performed through a back-projection scheme, that is, the output of the multi-modal feature progressive fusion module is used as additional feature information and input into the first cross-attention mechanism unit and the second cross-attention mechanism unit in the multi-modal feature progressive fusion module respectively; in the first cross-attention mechanism unit, the additional feature information is fused with the pixel-level fine-grained local features of the single-frame image modality; in the second cross-attention mechanism unit, the additional feature information and the global spatio-temporal features of the video modality are fused; the fused features are further processed through a separable convolutional layer to obtain fused features, and the obtained fused features are used as additional feature information again; the above back-projection operation is repeatedly executed until all video frames in the sequence are input, and the loop ends to obtain multi-modal fusion features.
[0038] The beneficial effects of the present invention are as follows:
[0039] 1. The present invention utilizes the pre-trained image encoders of the multi-modal large models SAM and CLIP to effectively extract the texture details and category information of video frames. This method not only reduces the cost of data processing, but also can more comprehensively capture the complex environmental information in wellsite operations, thus significantly improving the accuracy and reliability of wellsite video safety analysis.
[0040] 2. The present invention can effectively learn and extract the temporal information and spatial information in the wellsite video modality. This enables the system to maintain high-efficiency recognition and analysis capabilities in dynamic scenarios, accurately monitor continuous actions and complex operations during wellsite operations, and further improve the timeliness and accuracy of video safety analysis.
[0041] 3. The present invention adopts a multi-scale spatial attention feature extraction module to deeply extract the pixel-level fine-grained local features of single-frame video images, enabling the system to accurately identify tiny actions and small targets in wellsite videos, avoiding missing key details, and thus improving the comprehensiveness and accuracy of safety analysis.
[0042] 4. The design of the multi-modal feature progressive fusion module and the multi-task decoder module enables the system to adapt to different well-site video safety analysis tasks. Whether it is detecting the operation status of well-site equipment, identifying dangerous behaviors, or monitoring the standardization of personnel operations, this method can handle them efficiently and has strong adaptability.
[0043] 5. The present invention adopts a modular design, and each functional module can be independently optimized and upgraded, so the system has good scalability. With the changes in the well-site environment and safety requirements, the system can flexibly integrate new feature extraction methods and analysis models to support future technology upgrades and application expansions. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a schematic flowchart of the method proposed by the present invention;
[0045] Figure 2 is a schematic structural diagram of the safety analysis model;
[0046] Figure 3 is a schematic structural diagram of the video global feature extraction module based on the spatio-temporal attention mechanism;
[0047] Figure 4 is a schematic structural diagram of the multi-scale spatial attention feature extraction module;
[0048] Figure 5 is a schematic structural diagram of the multi-modal feature progressive fusion module;
[0049] Figure 6 is a schematic structural diagram of the multi-task decoder module based on multi-modal features;
[0050] Figure 7 is a schematic diagram of the safety analysis result obtained in the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0051] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0052] In an embodiment of the present invention, as Figure 1 shown, a multi-modal well-site video safety analysis method based on attention proposed by the present invention includes the following steps:
[0053] S1. Feed the frame sequence in the wellsite operation video to be analyzed into a pre-trained image encoder to obtain the low-level fusion features of the video frames. Here, the frame sequence is the frame sequence in a GOP of the video, and the wellsite operation video can be obtained through wellsite operation area monitoring devices or other means. CLIP is good at semantic understanding, while SAM focuses on spatial understanding in segmentation tasks. Therefore, an image encoder is composed of a pre-trained vision encoder based on SAM and a pre-trained vision encoder based on CLIP. Take the frame sequence of the wellsite operation video to be analyzed as the input of the image encoder, and input it into the pre-trained vision encoder based on SAM and the pre-trained vision encoder based on CLIP respectively to obtain two types of features regarding the wellsite operation video, including low-level features such as the texture, category, semantic concept, and object relationship of the video frames. Then, perform feature fusion on the two types of obtained features to obtain the low-level fusion features of all video frames, and its expression is:
[0054]
[0055] where is the input of the image encoder; is the function representation of the CLIP vision encoder, are the parameters of the CLIP vision encoder, is the output of the CLIP vision encoder; is the function representation of the SAM vision encoder, are the parameters of the SAM vision encoder, is the output of the SAM vision encoder; H is the finally obtained low-level fusion features of all video frames , are the weights corresponding to the SAM vision encoder, is CL the weights corresponding to the IP vision encoder.
[0056] S2. Feed the obtained low-level fusion features into a video global feature extraction module based on a spatio-temporal attention mechanism to obtain the global spatio-temporal features of the video modality, that is, temporal information and spatial information. As Figure 3As shown in the figure, the video global feature extraction module based on the spatio-temporal attention mechanism includes a small kernel depth convolution layer, a multi-head temporal attention mechanism unit, a dilated depth convolution layer, and a 1×1 convolution layer connected in sequence; the video global feature extraction module based on the spatio-temporal attention mechanism also includes a first temporal self-attention mechanism unit, a first average pooling layer, and a first fully connected layer connected in sequence; the low-level fusion features obtained in step S1 are respectively used as the inputs of the small kernel depth convolution unit and the first temporal self-attention mechanism unit; the output of the 1×1 convolution layer and the output of the first fully connected layer are subjected to a Kronecker product operation and then a Hadamard product operation with the low-level fusion features obtained in step S1 to obtain the global spatio-temporal features of the video modality. During the process of extracting the global spatio-temporal features of the video modality, the finally obtained attention is the product of the dynamic attention and the static attention. The small kernel depth convolution 、multi-head temporal attention mechanism, dilated depth convolution, and 1×1 convolution are used to model the large kernel convolution to obtain the inter-frame dynamic attention SA ; the intra-frame static attention is obtained through the temporal self-attention mechanism, average pooling, and fully connected operations DA . Its calculation expression is:
[0057]
[0058] where represents the output of the 1×1 convolution layer, represents the 1×1 convolution, represents the dilated depth convolution, represents the multi-head temporal attention processing, represents the small kernel depth convolution, is the low-level fusion feature obtained in step S1; represents the output of the first fully connected layer, represents the fully connected, represents the global average pooling, represents the temporal self-attention processing; is the output of the video global feature extraction module based on the spatio-temporal attention mechanism, represents the Kronecker product, represents the Hadamard product.
[0059] S3. Input the key frames (I-frames) in the frame sequence described in S1 into the feature extraction module based on the multi-scale spatial attention to obtain the pixel-level fine-grained local features of the single-frame image modality. As Figure 4As shown in the figure, the feature extraction module based on multi-scale spatial attention includes a VIT unit, a 2D convolutional unit, an atrous convolutional layer, a multi-head cross-attention mechanism unit, a self-attention mechanism unit, and a fast normalized feature fusion unit connected in sequence; the key frame image is the input of the feature extraction module based on multi-scale spatial attention, and the output of the fast normalized feature fusion unit is the pixel-level fine-grained local feature of the single-frame image modality.
[0060] When extracting features from the key frame, feature extraction is performed at five scales and fast normalization fusion is used to obtain multi-scale spatial attention features; the five scales are the original image scale of the key frame image, one-half of the original image scale, one-fourth of the original image scale, one-eighth of the original image scale, and one-sixteenth of the original image scale. First, the key frame images at different scales are used by ViT to extract preliminary features, and the extracted features are sequentially passed through 2D convolution and atrous convolution to obtain features at different scales. Before performing feature fusion at different scales, it is necessary to further capture the self and multi-angle dependence relationships of features at different scales in a fine-grained manner, and perform top-down and bottom-up two-way feature fusion, which is achieved through a multi-head cross-attention mechanism and a self-attention mechanism. In the network, in order to achieve more efficient aggregation of image features at different resolutions, fast normalized fusion (Fast Normalized Fusion) is used to perform adaptive feature fusion. Its calculation expression is:
[0061]
[0062]
[0063]
[0064] Among them, is the key frame image at the i th scale, represents processing, represents 2D convolution, represents atrous convolution, represents the preliminary feature extracted from the key frame image at the i th scale; represents multi-head cross-attention processing, represents self-attention processing, represents the output of the self-attention mechanism unit, represents the output of the fast normalized feature fusion unit, represents the corresponding weight, representing the contribution degree of the corresponding feature; represents thej The learning weights corresponding to the outputs are used to perform normalization processing; Extra parameters to ensure the stability of the results, .
[0065] When performing feature fusion on features of different scales, fast normalization fusion is adopted. The contribution degrees of features of different scales to the total output features are different. Therefore, additional weights are added to the features of each scale, enabling the security analysis model to learn the respective proportions of features of different scales. The range of each normalized weight is between 0 and 1. To ensure , after each , a Relu activation function is added.
[0066] S4. Input the globally spatio-temporal features of the extracted video modality and the pixel-level fine-grained local features of the single-frame image modality into the multi-modal feature progressive fusion module for deep fusion to obtain multi-modal fusion features. As Figure 5 shown, the multi-modal feature progressive fusion module includes a first cross-attention mechanism unit, a multi-layer perceptron, a second average pooling layer, and a separable convolutional layer connected in sequence, and also includes a second cross-attention mechanism unit and a temporal modeling unit connected in sequence; the output of the temporal modeling unit serves as another input to the separable convolutional layer; the inputs of the first cross-attention mechanism unit and the second cross-attention mechanism unit serve as the inputs of the multi-modal feature progressive fusion module; the output of the separable convolutional layer serves as the output of the multi-modal feature progressive fusion module.
[0067] To avoid the problem of loss of shallow features that may occur after feature fusion, when fusing the globally spatio-temporal features of the video modality and the pixel-level fine-grained local features of the single-frame image modality, feature fusion is performed through a back-projection scheme, which helps to better achieve multi-modal feature fusion. The back-projection is to use the output of the multi-modal feature progressive fusion module as additional feature information and input it into the first cross-attention mechanism unit and the second cross-attention mechanism unit in the multi-modal feature progressive fusion module respectively; in the first cross-attention mechanism unit, the additional feature information is fused with the pixel-level fine-grained local features of the single-frame image modality; in the second cross-attention mechanism unit, the additional feature information and the globally spatio-temporal features of the video modality are fused; the fused features are further processed by the separable convolutional layer to obtain fused features, and the obtained fused features are used as additional feature information again; the above back-projection operation is repeatedly executed until all the video frame sequences are input, and the loop ends to obtain multi-modal fusion features. Represent the input of features of different modalities as , represent the input features of the th modality. In this embodiment , indicating there are two modalities; representing the finally obtained multi-modal fusion features as , where represents additional feature information, represents feature fusion. In this way, connections are established between features at different levels, enabling the additional feature information after feature fusion to be utilized by the shallow network, avoiding the problem of "either fusion or loss" in multi-modal feature fusion.
[0068] S5. Input the multi-modal fusion features into the multi-task decoder module based on multi-modal features to obtain the safety analysis results of the wellsite operation video, such as Figure 7 shown, including the safety hazard analysis results of three categories: the worker himself (worker sleeping, using mobile phone, slacking off), the worker and the environment (standing dangerously, leaning dangerously, walking dangerously), and the environment (missing hopper at the material inlet, unclosed cover plate, missing warning signs). The multi-task learning method can learn different types of tasks and conduct result analysis, while avoiding the limitations of the existing video analysis algorithms that can only target a single task, resulting in the singularity of the concerned safety hazards, and the complex problems brought by multiple models. As Figure 6 shown, the multi-task decoder module based on multi-modal features includes a first linear layer group, a third average pooling layer, a flattening layer, a second linear layer group, and a second fully connected layer connected in sequence.
[0069] As Figure 2 shown, the image encoder, the video global feature extraction module based on the spatio-temporal attention mechanism, the multi-modal feature progressive fusion module, the multi-task decoder module based on multi-modal features, and the feature extraction module based on multi-scale spatial attention constitute the safety analysis model.
[0070] In this embodiment, the specific steps for training the safety analysis model include:
[0071] A1. Obtain the wellsite operation video and label the corresponding safety hazards in the wellsite operation video to obtain the training set;
[0072] A2. Send the videos in the training set into the safety analysis model to obtain the safety analysis results of the wellsite operation video, including the safety hazard analysis results of three categories: the worker himself, the worker and the environment, and the environment;
[0073] A3. Calculate the loss of each task according to the safety analysis results of the wellsite operation video, and optimize the model with the joint loss as the overall target loss. If the joint loss no longer decreases or reaches the set number of training times, end the training; otherwise, return to step A2;
[0074] where the joint loss The calculation expression of is:
[0075]
[0076]
[0077]
[0078]
[0079] Among them, represents the loss of judging hidden dangers of workers' rule violations, represents the number of samples, represents the total number of types of hidden dangers of workers' rule violations, represents the th sample contains the probability of hidden dangers ; is a prediction index. When the hidden danger type label in the th sample output by the safety analysis model is consistent with the true hidden danger type label , it is 1, otherwise it is 0; represents the loss of judging the relationship between workers and environmental safety, D represents the total number of types of the relationship between workers and environmental safety; represents the loss of judging hidden dangers in the environment, G represents the total number of types of hidden dangers in the environment; represents a hyperparameter used for weighting various losses; ln represents the natural logarithm.
[0080] Through multi-task learning, the original feature extraction network can learn the shared representations of different video analysis tasks, and utilize the complementarity of different task feature representations to improve the overall performance, and maintain the consistency of feature semantics, enabling the model to implement multiple functions under a unified framework.
[0081] In summary, the present invention has good timeliness and high precision, can handle diverse tasks, and has extremely strong adaptability.
Claims
1. A multimodal well site video safety analysis method based on attention, characterized in that: The following steps are involved: S1, sending the frame sequence in the well site operation video to be analyzed into the pre-trained image encoder to obtain the low-level fusion features of the video frame; S2. Input the obtained low-level fusion features into the video global feature extraction module based on the spatiotemporal attention mechanism to obtain the global spatiotemporal features of the video modality, namely, the temporal information and spatial information; S3, inputting the key frames in the frame sequence described in S1 into a feature extraction module based on multi-scale spatial attention to obtain pixel-level fine-grained local features of a single-frame image modality; S4, inputting the extracted global spatiotemporal features of the video modality and the pixel-level fine-grained local features of the single-frame image modality into the multimodal feature progressive fusion module for deep fusion to obtain multimodal fusion features; S5, inputting the multimodal fusion features into a multi-task decoder module based on multimodal features to obtain safety analysis results of the well site operation video, including safety hazard analysis results of three categories: workers themselves, workers and environment, and environment; The video global feature extraction module based on the spatiotemporal attention mechanism includes a small-core deep convolution layer, a multi-head temporal attention mechanism unit, an expanded deep convolution layer, and a 1×1 convolution layer connected in sequence; the video global feature extraction module based on the spatiotemporal attention mechanism also includes a first temporal self-attention mechanism unit, a first average pooling layer, and a first fully connected layer connected in sequence; the low-level fusion features obtained in step S1 are respectively used as inputs of the small-core deep convolution unit and the first temporal self-attention mechanism unit; the output of the 1×1 convolution layer is subjected to a Kronecker product operation with the output of the first fully connected layer, and then subjected to a Hadamard product operation with the low-level fusion features obtained in step S1 to obtain the global spatiotemporal features of the video modality; The feature extraction module based on multi-scale spatial attention includes a VIT unit, a 2D convolution unit, a hole convolution layer, a multi-head cross attention mechanism unit, a self-attention mechanism unit and a fast normalized feature fusion unit connected in sequence; the key frame image is the input of the feature extraction module based on multi-scale spatial attention, and the output of the fast normalized feature fusion unit is the pixel-level fine-grained local feature of the single-frame image modality; The image encoder, the video global feature extraction module based on the spatiotemporal attention mechanism, the multimodal feature progressive fusion module, the multi-task decoder module based on multimodal features, and the feature extraction module based on multi-scale spatial attention constitute the security analysis model.
2. The attention-based multimodal wellsite video safety analysis method according to claim 1, characterized in that: Image encoders include SAM-based pre-trained visual encoders and CLIP-based pre-trained visual encoders; The frame sequence in the well site operation video to be analyzed is used as the input of the image encoder, and is input into the pre-trained visual encoder based on SAM and the pre-trained visual encoder based on CLIP respectively, to obtain two types of features about the well site operation video; Then the two types of features are fused to obtain the low-level fusion features of all video frames, which are expressed as follows: in is the input of the image encoder; is the function representation of the CLIP visual encoder, are the parameters of the CLIP visual encoder, is the output of the CLIP visual encoder; is the function representation of the SAM visual encoder, are the parameters of the SAM visual encoder, is the output of the SAM visual encoder; H is the low-level fusion feature of all video frames finally obtained, is the weight corresponding to the SAM visual encoder, yes CL The weights corresponding to the IP visual encoder.
3. The attention-based multimodal wellsite video safety analysis method according to claim 1, characterized in that: The calculation expression of the global spatiotemporal characteristics of the video modality is: in, represents the output of the 1×1 convolutional layer, represents a 1×1 convolution, represents a depthwise convolution with dilation, represents multi-headed temporal attention processing, represents a small kernel depth convolution, is the low-level fusion feature obtained in step S1; represents the output of the first fully connected layer, Indicates full connection, represents global average pooling, represents temporal self-attention processing; is the output of the video global feature extraction module based on the spatiotemporal attention mechanism. represents the Kronecker product, Represents the Hadamard product.
4. The attention-based multimodal wellsite video safety analysis method according to claim 1, characterized in that: In step S3, a feature extraction module based on multi-scale spatial attention extracts features of five scales from the key frame image and uses fast normalization fusion to obtain multi-scale spatial attention features; The five scales are respectively the original image scale of the key frame image, half of the original image scale, one quarter of the original image scale, one eighth of the original image scale, and one sixteenth of the original image scale.
5. The attention-based multimodal wellsite video safety analysis method according to claim 4, characterized in that: The calculation expression of the pixel-level fine-grained local features of the single-frame image modality is: in, For the i Key frame images at different scales, express deal with, represents 2D convolution, represents the dilated convolution, Expressing the i Preliminary features extracted from keyframe images at different scales; represents multi-head cross-attention processing, represents self-attention processing, Represents the output of the self-attention mechanism unit; represents the output of the fast normalized feature fusion unit, express The corresponding weight, Represents the self-attention mechanism unit j The learning weights of the outputs are used to Perform normalization processing; Additional parameters to ensure the stability of the results.
6. The attention-based multimodal wellsite video safety analysis method according to claim 1, characterized in that: The multimodal feature progressive fusion module includes a first cross-attention mechanism unit, a multi-layer perceptron, a second average pooling layer, and a separable convolution layer connected in sequence, and also includes a second cross-attention mechanism unit and a timing modeling unit connected in sequence; the output of the timing modeling unit serves as another input of the separable convolution layer; the input of the first cross-attention mechanism unit and the second cross-attention mechanism unit serves as the input of the multimodal feature progressive fusion module; the output of the separable convolution layer serves as the output of the multimodal feature progressive fusion module.
7. The attention-based multimodal wellsite video safety analysis method according to claim 1, characterized in that: The multi-task decoder module based on multimodal features includes a first linear layer group, a third average pooling layer, a flattening layer, a second linear layer group, and a second fully connected layer connected in sequence.
8. The attention-based multimodal wellsite video safety analysis method according to claim 1, characterized in that: The specific steps for training a security analysis model include: A1. Obtain the well site operation video and mark the corresponding safety hazards in the well site operation video to obtain a training set; A2. Send the videos in the training set to the safety analysis model to obtain the safety analysis results of the well site operation videos, including the safety hazard analysis results of the three categories of workers themselves, workers and environment, and environment; A3. Calculate the joint loss based on the safety analysis results of the well site operation video. If the joint loss no longer decreases or reaches the set training times, end the training; otherwise, return to step A2; The combined loss The calculation expression is: in, It means the loss caused by workers' violation of rules and regulations. represents the number of samples, Represents the total number of hidden dangers of workers violating regulations, Indicates Samples contain hidden dangers The probability of As the prediction index, when the safety analysis model outputs The potential hazard type labels in the samples and the real potential hazard type labels When consistent is 1, otherwise is 0; Indicates the loss of judgment of workers on the relationship between safety and the environment, D The total number of categories representing the relationship between workers and environmental safety; Indicates the loss of judgment on hidden dangers in the environment. G Represents the total number of types of hidden dangers in the environment; represents a hyperparameter used to weight various types of losses; ln represents the natural logarithm.
9. The attention-based multimodal wellsite video safety analysis method according to claim 5, characterized in that: When the multimodal feature progressive fusion module is used in step S4 to fuse the global spatiotemporal features of the video modality and the pixel-level fine-grained local features of the single-frame image modality, feature fusion is performed through a back-projection scheme, that is, the output of the multimodal feature progressive fusion module is used as additional feature information, and is input into the first cross-attention mechanism unit and the second cross-attention mechanism unit in the multimodal feature progressive fusion module respectively; in the first cross-attention mechanism unit, the additional feature information is fused with the pixel-level fine-grained local features of the single-frame image modality; in the second cross-attention mechanism unit, the additional feature information is fused with the global spatiotemporal features of the video modality; The fused features are further processed through a separable convolutional layer to obtain fused features, and the obtained fused features are once again used as additional feature information; the above-mentioned back projection operation is repeated until all video frame sequences are input, and the loop ends to obtain multimodal fusion features.
Citation Information
Patent Citations
Motion recognition method of decoupling 3D network based on multi-scale self-attention mechanism
CN117011943A
Posture-adjusted calculation of physiological signals
US20190313915A1