Intelligent identification method based on safety behavior of oilfield electric operation

By using a dual-branch temporal recognition network and an adaptive response strategy, the problems of fine-grained differences, temporal continuity, and boundary ambiguity in traditional oilfield power operation behavior recognition are solved, achieving high-precision and real-time monitoring of oilfield power operation behavior.

CN121305675BActive Publication Date: 2026-04-24NORTHEAST GASOLINEEUM UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHEAST GASOLINEEUM UNIV
Filing Date
2025-10-17
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Traditional oilfield power operation behavior recognition methods have shortcomings in terms of small differences in fine-grained actions, strong temporal continuity, blurred boundaries, and slow response speed, making it difficult to meet the high-precision and real-time monitoring requirements of oilfield power operations.

Method used

A dual-branch temporal recognition network based on boundary awareness and adaptive response mechanisms is adopted. Through dual-branch parallel processing channels, boundary awareness detection mechanism, Actionformer prediction module and adaptive response strategy, combined with multi-objective loss function for end-to-end optimization, the accurate recognition of oilfield power operation behavior is achieved.

Benefits of technology

It significantly improves the accuracy and response speed of oilfield power operation behavior recognition, can accurately identify action boundaries, adapt to complex industrial environments, meet real-time safety monitoring needs, and improve the system's adaptability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121305675B_ABST
    Figure CN121305675B_ABST
Patent Text Reader

Abstract

The application relates to an intelligent identification method based on oil field electric power operation safety behaviors, which comprises the following steps: 1, preprocessing the oil field electric power operation video, and constructing a double-branch parallel processing channel at the input layer of a double-branch time sequence identification network architecture; 2, left and right branch local space-time feature extraction; 3, intermediate deployment in a boundary perception detection mechanism; 4, parallel processing of an Actionformer prediction module; 5, deployment of a double-branch feature fusion layer self-adaptive response strategy; 6, hierarchical network integration of a multi-target loss function; 7, identification model training, end-to-end optimization and verification; 8, oil field electric power operation behavior identification by using the identification model. The double-branch time sequence identification network based on the boundary perception and self-adaptive response mechanism effectively solves the technical problems that the traditional video action identification method has in the oil field electric power operation scene, such as small fine-grained action difference, fuzzy state conversion and strong time sequence continuity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and deep learning, specifically to an intelligent recognition method for safe behaviors in oilfield power operations. Background Technology

[0002] Oilfield power operations are high-risk, high-precision industrial operations, making safety paramount. With the development of deep learning technology, video surveillance-based behavior recognition technology has been widely applied in industrial safety monitoring and intelligent monitoring. However, behavior recognition tasks in oilfield power operation scenarios face numerous challenges. Traditional behavior recognition methods rely on manual feature extraction and shallow models, failing to fully capture complex temporal dynamics and local action details, resulting in shortcomings in fine-grained action recognition, boundary accuracy, and response speed. Traditional image-based behavior recognition methods mainly face the following problems:

[0003] 1. Small differences in fine-grained movements: In oilfield power operations, workers' behaviors are highly similar, and the differences between fine-grained movements are small, which makes traditional methods perform poorly in distinguishing similar movements.

[0004] 2. Strong temporal continuity: The behavior in oilfield power operations has strong temporal continuity. Traditional methods are difficult to capture the temporal relationship between continuous actions, resulting in the loss or ambiguity of temporal information.

[0005] 3. Blurred boundaries: In complex work scenarios, the start and end boundaries of actions are often not obvious. Traditional methods are difficult to accurately locate the boundaries of actions, which affects the accuracy of behavior recognition.

[0006] 4. Slow response speed: Traditional behavior recognition methods usually rely on complex calculations, resulting in a slow response speed, making it difficult to meet the needs of real-time monitoring.

[0007] To overcome these challenges, deep learning-based behavior recognition methods have made some progress in recent years. For example, while methods based on human key points have achieved good results in some areas, their accuracy is limited in oilfield and power operation scenarios due to complex personnel attire, occlusion, and diverse working postures. Image-based behavior recognition methods are gradually becoming the mainstream approach for behavior recognition in oilfield and power operations, especially the combination of convolutional neural networks and temporal models, which gives the models a stronger ability to extract spatiotemporal features. However, in complex industrial scenarios, especially in the high-risk and high-complexity environment of oilfield and power operations, there is still room for further improvement. Summary of the Invention

[0008] The purpose of this invention is to provide an intelligent identification method based on the safety behavior of oilfield power operations. This intelligent identification method is used to solve the problems of small differences in fine-grained actions, ambiguous state transitions, strong temporal continuity, and low boundary recognition accuracy in traditional oilfield power operation behavior identification methods.

[0009] The technical solution adopted by this invention to solve its technical problem is as follows: This intelligent identification method based on the safety behavior of oilfield power operations is an oilfield power operation behavior identification method based on a dual-branch temporal identification network with boundary perception and adaptive response mechanism, including the following steps:

[0010] Step 1: Preprocess the oilfield power operation video by dividing the original video sequence into fixed-length segments according to the time dimension. Each segment contains a continuous sequence of frames. Construct a dual-branch parallel processing channel in the input layer of the dual-branch temporal recognition network architecture: a left branch channel and a right branch channel.

[0011] Step 2: Extraction of local spatiotemporal features of the left and right branches;

[0012] The left branch channel uses R(2+1)D Backbone as the backbone network, while the right branch channel captures the temporal continuity in the video, ensuring that the temporal changes of the action are effectively captured.

[0013] Step 3: Deployment of the boundary awareness detection mechanism in the middle;

[0014] In the intermediate connection stage of the dual-branch temporal recognition network architecture, a boundary-aware detection mechanism is deployed. The boundary-aware detection mechanism calculates the boundary weight score of each frame through an attention mechanism, identifies the start and end boundaries of the action, and outputs a feature representation with boundary enhancement information.

[0015] Step 4: The Actionformer prediction module processes data in parallel;

[0016] In the dual-branch temporal recognition network architecture, after the boundary-aware detection mechanism, the Actionformer prediction module is precisely embedded. The Actionformer prediction module receives the enhanced features output by the boundary-aware module, performs global temporal modeling through a long and short temporal fusion structure, and optimizes the focusing of temporal information by combining an attention mechanism to process behavior sequences with strong temporal continuity. The Actionformer module works in parallel with the classification module.

[0017] Step 5: Deploy the adaptive response strategy for the dual-branch feature fusion layer;

[0018] In the final fusion stage of the dual-branch temporal recognition network architecture, an adaptive response strategy module is introduced to form the recognition model. This module is located at the convergence point of the two branches. The adaptive response strategy receives the spatial features of the left branch channel and the temporal features of the right branch channel, dynamically calculates the branch weight coefficients α and β based on the prediction confidence, and achieves intelligent fusion through a weighted fusion formula:

[0019]

[0020] This strategy module adjusts the branch contribution in real time at a specific location in the fusion layer based on the input complexity, achieving a dynamic balance between fast response and stable output.

[0021] Step Six: Hierarchical network integration of multi-objective loss functions;

[0022] During model training, the multi-objective loss function is precisely integrated into different layers of the network. The classification loss is applied to the final output layer, the boundary loss is applied to the boundary awareness module, and the temporal consistency loss is applied to the output of the Actionformer prediction module.

[0023] Step 7: Identify model training, end-to-end optimization, and validation;

[0024] Through an end-to-end training strategy, all modules work together to form a unified optimization goal;

[0025] Step 8: Use the recognition model to identify oilfield power operation behaviors.

[0026] The method for preprocessing oilfield power operation videos in the above steps is as follows:

[0027] A dual-scale sliding window strategy is employed to segment the original video sequence. This strategy constructs two parallel processing channels: a main window and an auxiliary window. The main window is configured to extract sixteen consecutive frames as a complete time segment, capturing complete action context information and fully covering the entire execution process of complex power operation actions. The auxiliary window is configured to extract eight consecutive frames, responsible for detecting rapidly changing action details and instantaneous state transitions. A frame data bag mechanism is implemented for intelligent data distribution processing. The input original video sequence V∈R^(N×H×W×C) is segmented according to a preset time interval. Each time segment ensures that it contains enough frames to maintain the temporal continuity of the action, where N represents the total number of frames, and H, W, and C represent the height, width, and number of channels of the image, respectively. Furthermore, each time segment is standardized and normalized. The system uses a zero-mean, unit variance standardization operation.

[0028]

[0029] Where μ is the mean and σ is the standard deviation, an adaptive histogram equalization technique is used to enhance each frame of the image to improve the contrast and clarity of the image, taking into account the common light changes and shadow interference in the oilfield operating environment.

[0030] Finally, the system identifies and locates key frames of motion using an optical flow estimation algorithm, calculates the optical flow vectors between adjacent frames, and the optical flow constraint equation is:

[0031]

[0032] in For image gradient, The time gradient is denoted as u and v, which are optical flow vectors. The starting, climax, and ending phases of the action are identified by analyzing the changes in the amplitude and direction of the optical flow.

[0033] Step two of the above plan is as follows:

[0034] The left branch channel uses an improved R(2+1)D Backbone as the backbone network, decomposing the traditional 3D convolution operation into a cascaded combination of 2D spatial convolution and 1D temporal convolution:

[0035] Step two of the above plan is as follows:

[0036] The left branch channel uses an improved R(2+1)D Backbone as the backbone network, decomposing the traditional 3D convolution operation into a cascaded combination of 2D spatial convolution and 1D temporal convolution:

[0037]

[0038] Where W_2D is a 2D spatial convolution kernel, W_1D is a 1D temporal convolution kernel, and the left branch channel also integrates a multi-scale feature extraction module, which captures spatial features of different scales by using convolution kernels of different sizes in parallel; the left branch outputs a feature tensor. Where H and W are the spatial dimensions after convolution downsampling, and D is the number of feature channels. By decomposing the convolution operation, the global spatiotemporal information of the action is effectively captured; the right branch extracts the temporal features in the video.

[0039] Step three in the above plan specifically refers to:

[0040] The boundary-aware detection mechanism is deployed in the intermediate connection stage of the dual-branch temporal recognition network architecture. This dual-branch architecture accurately identifies the start and end boundaries of actions. The main task branch focuses on extracting and classifying the semantic features of the actions. It receives temporal features N×C×T×H×W from the dual-branch fusion module, first adjusting the feature dimensions to N×C / 4×1×1 using a 1×1×3 convolutional layer, and then outputting the probability distribution of each action category through a classification layer. The boundary-aware branch is responsible for learning the transition boundary information between actions, using the same convolutional configuration to extract temporal features, and generating a temporal boundary map through the Softmas activation function. Each element in the boundary graph represents the probability that the corresponding time step is an action boundary;

[0041] The boundary detection submodule calculates the boundary probability:

[0042]

[0043] The first and second derivatives of the time-series features are calculated along the time dimension. By analyzing the trends in the derivatives, time points where the features change drastically are identified. An attention weighting mechanism reweights the features at different time steps based on boundary information.

[0044]

[0045] This represents the element-wise multiplication operation;

[0046] Features from the main task branch and the boundary awareness branch are integrated through a soft attention mechanism:

[0047]

[0048] Where λ is a weighting parameter that adjusts the importance of boundary information;

[0049] The training of marginal branches employs an independent supervised learning strategy, with the marginal loss function being:

[0050]

[0051] Where y_boundary is the boundary label and p_boundary is the predicted boundary probability.

[0052] Step four in the above scheme is specifically as follows:

[0053] The ActionFormer prediction module is embedded after the boundary-aware detection mechanism. It is responsible for processing the temporal features after boundary enhancement. It achieves global modeling of the action process and accurate capture of local changes through a long and short temporal fusion structure. Combined with multi-scale temporal processing and attention mechanism optimization, it constructs a complete temporal dependency modeling system. It adopts a hierarchical processing architecture, from the low-level multi-scale temporal convolution to the high-level global temporal modeling, to gradually extract and optimize the temporal feature representation.

[0054] The long-short temporal fusion structure achieves feature capture across different time spans through a multi-scale temporal convolutional network:

[0055]

[0056] Key temporal nodes in the sequence are identified through a self-attention mechanism, and temporal order information is maintained through a positional encoding mechanism. This process works in parallel with the classification module, forming a processing structure: temporal modeling → boundary awareness → parallel prediction.

[0057] Step five of the above plan is as follows:

[0058] As a core component in the final fusion stage of the network architecture, the adaptive response strategy module is deployed at the convergence point of the two branches. It undertakes the important function of intelligently fusing spatial and temporal features. Through a dynamic weight calculation mechanism, combined with confidence assessment and complexity analysis, it realizes real-time adjustment of the contribution of different branch features. First, it receives the spatial features of the left branch channel and the temporal features of the right branch channel as input. Then, it calculates the prediction confidence through multi-dimensional feature analysis. Finally, it dynamically determines the branch weight coefficient based on the confidence and input complexity.

[0059] The confidence level calculation adopts a multi-dimensional comprehensive evaluation method. The dynamic weight coefficient calculation is based on a comprehensive analysis of confidence level and input complexity. First, the complexity of spatial features is calculated by analyzing the variance and dispersion of feature vectors to measure the complexity of spatial information. The complexity of temporal features is evaluated by calculating the magnitude of change in temporal gradient. The more dramatic the gradient change, the more complex the temporal information. Based on the complexity index and the overall confidence level, dynamic weight coefficients are obtained through neural network learning to ensure that the contribution ratio of the two branches can be adaptively adjusted under different scenarios.

[0060] Intelligent fusion achieves the organic integration of features through a weighted fusion formula. The final fused feature contains a weighted combination of the two branches, and feature transformation and optimization are performed through a multilayer perceptron. In the fusion layer, the contribution of the branches is adjusted in real time according to the input complexity. When the spatial scene is complex and the working environment is variable, the weight of the spatial branch is increased. When the action sequence is complex and the temporal change is drastic, the weight of the temporal branch is increased, so as to achieve a dynamic balance between fast response and stable output. Beneficial effects

[0061] 1. The dual-branch temporal recognition network based on boundary awareness and adaptive response mechanisms proposed in this invention effectively solves the technical challenges of traditional video action recognition methods in oilfield power operation scenarios, such as small differences in fine-grained actions, ambiguous state transitions, and strong temporal continuity, through innovative network architecture design. The method can accurately identify each stage of a worker's pole-climbing behavior, providing a reliable technical means to ensure the safety of oilfield power operations and significantly improving the intelligence level of operation safety monitoring.

[0062] 2. The dual-branch temporal fusion module of this invention adopts a parallel processing mechanism of upper and lower branches. The upper branch extracts local spatiotemporal features through an R(2+1)D backbone network, while the lower branch captures dynamic temporal changes through a temporal modeling module. The features from the two branches are integrated through an adaptive weighted fusion mechanism. Compared with traditional single-branch feature extraction methods, this module can more comprehensively capture the complex spatiotemporal features of power operation actions, significantly improving feature representation capabilities and providing a richer feature foundation for subsequent action recognition.

[0063] 3. The boundary-aware detection mechanism of this invention adopts a dual-branch architecture design. The main task branch focuses on the extraction and classification of action semantic features, while the boundary-aware branch is specifically responsible for learning the transition boundary information between actions, enabling precise location of the start and end times of actions. This mechanism significantly improves the accuracy of boundary discrimination, solves the problem of insufficient accuracy in action boundary recognition by traditional methods, and provides accurate technical support for state transition detection in real-time security monitoring.

[0064] 4. The adaptive response strategy of this invention, through confidence calculation and dynamic parameter adjustment mechanisms, can adaptively adjust the system's response mode according to the complexity of the input video and the reliability of the prediction results. This strategy achieves an optimal balance between fast response and stable output, enabling rapid response to simple actions with clear boundaries and providing more accurate analysis for complex actions with ambiguous boundaries, significantly improving the system's adaptability and practicality in complex industrial environments.

[0065] 5. The multi-objective loss function of this invention, through joint optimization of classification loss, localization loss, boundary-aware loss, and temporal consistency constraints, effectively enhances the temporal consistency and smoothness of prediction results while ensuring classification accuracy. This design avoids the problem of drastic jumps in prediction results in traditional methods, improves the stability and reliability of system output, and ensures the consistency of prediction results during continuous monitoring.

[0066] 6. The OPCD dataset for identifying the operation behavior of oilfield power poles constructed in this invention is based on human factors engineering theory and power operation risk assessment model. It adopts a hierarchical classification strategy to scientifically deconstruct the operation behavior into four key state transition nodes. This classification system conforms to the theoretical framework of Markov decision process and provides a high-quality data foundation and evaluation standard for the standardized research and application of oilfield power operation behavior identification.

[0067] 7. This invention aims to solve the problems of small differences in fine-grained actions, ambiguous state transitions, strong temporal continuity, and low boundary recognition accuracy in traditional oilfield power operation behavior recognition methods. By introducing a dual-branch temporal network structure, a boundary perception module, and an adaptive response strategy, this invention can effectively improve the recognition accuracy and real-time response capability of oilfield power operation behavior, ensuring safety and efficiency during the operation process.

[0068] 8. Experimental results demonstrate that this invention possesses significant technical advantages over existing technologies, achieving improvements of 7.5 percentage points and 12.3 percentage points respectively compared to traditional methods. Furthermore, the method achieves a real-time processing speed of 35 frames per second on standard GPU devices, meeting the dual requirements of accuracy and real-time performance in industrial settings, and providing an efficient and feasible technical solution for the engineering application of safety monitoring in oilfield power operations. Attached Figure Description

[0069] Figure 1 This is a diagram of a balanced multi-objective fusion architecture.

[0070] Figure 2 This is a diagram of the action recognition network structure of R(2+1)D and ActionFormer.

[0071] Figure 3 This is a diagram of a dual-branch fusion module.

[0072] Figure 4 This is a diagram of the boundary awareness detection mechanism architecture.

[0073] Figure 5 This is a schematic diagram of a self-made dataset from an oilfield. Detailed Implementation

[0074] The present invention will be further described below with reference to the accompanying drawings:

[0075] This intelligent recognition method for safe behavior in oilfield power operations is a dual-branch temporal network based on boundary awareness and adaptive mechanisms. It uses a video dataset as the basis for training and testing, containing videos of worker behavior in various oilfield power operation scenarios. Feature extraction from the video data is performed through two parallel branches: the upper branch uses an R(2+1)D backbone for local spatiotemporal feature extraction, while the lower branch is dedicated to temporal modeling. The boundary awareness module focuses on action transition frames through an attention mechanism, accurately identifying the start and end boundaries of actions. Simultaneously, an adaptive response strategy is combined to dynamically adjust the judgment rhythm based on the model's confidence level, improving the system's response speed and stability while maintaining classification accuracy. This method effectively addresses issues such as fine-grained differences in actions, strong temporal continuity, and blurred boundaries in oilfield power operation scenarios. Details are as follows:

[0076] Step 1: Input Data Preprocessing and Dual-Branch Architecture Initialization

[0077] First, the oilfield power operation video is preprocessed by dividing the original video sequence into fixed-length segments along the time dimension, with each segment containing a continuous sequence of frames. A dual-branch parallel processing channel is constructed at the input layer of the network architecture: the left branch is dedicated to spatial feature extraction, and the right branch is dedicated to temporal modeling. This dual-branch parallel design achieves the separation of spatial and temporal information processing at the network's front end, avoiding the problem of mutual interference between spatial and temporal information in traditional single-branch networks.

[0078] Step 2: Extraction of local spatiotemporal features of left and right branches

[0079] In the left branch of the network architecture, an R(2+1)D Backbone is used as the backbone network. This module is located at the core processing position of the left branch and is specifically responsible for extracting the local spatiotemporal features of each video segment. The R(2+1)D module decomposes the convolution operation, breaking down 3D convolution into 2D spatial convolution and 1D temporal convolution, effectively capturing the global spatiotemporal information of the action. This module plays a crucial role in the center of the left branch, ensuring the purity and integrity of spatial information through a dedicated spatial feature extraction channel. The right branch focuses on temporal feature extraction. Unlike the left branch, the right branch does not use the R(2+1)D Backbone. It relies on other temporal modules for temporal modeling. The main task of the right branch is to capture the temporal continuity in the video through temporal convolution or other temporal modeling methods, ensuring that the temporal changes of the action are effectively captured.

[0080] Step 3: Deployment of the Boundary Awareness Detection Mechanism

[0081] In the intermediate connection stage of the network architecture, an innovative boundary-aware detection mechanism is deployed as a key link connecting temporal features and final prediction. This mechanism calculates boundary weight scores for each frame using an attention mechanism, identifies the start and end boundaries of actions, and outputs feature representations with boundary enhancement information. This module receives integrated features from the left and right branches and simultaneously outputs the processed boundary-enhanced features to both the Actionformer prediction and classification modules on the right, ensuring that all subsequent prediction processing is based on the boundary-enhanced feature representations, fundamentally improving the accuracy of boundary detection.

[0082] Step 4: Parallel processing by the Actionformer prediction module

[0083] In the right-hand prediction phase of the network architecture, the Actionformer prediction module is precisely embedded, located after the boundary-aware detection mechanism. The Actionformer module receives enhanced features from the boundary-aware module, performs global temporal modeling through a long-short temporal fusion structure, and optimizes the focusing of temporal information using an attention mechanism to process behavioral sequences with strong temporal continuity. This module works in parallel with the classification module, forming a processing structure of "temporal modeling → boundary awareness → parallel prediction." This specific module arrangement ensures that the prediction process is based on boundary-optimized features, significantly improving the modeling capability for long-term temporal data.

[0084] Step 5: Deployment of Adaptive Response Strategy for Dual-Branch Feature Fusion Layer

[0085] In the final fusion stage of the network architecture, an adaptive response strategy module is introduced, located at the convergence point of the two branches. The adaptive response strategy receives the spatial features of the left branch and the temporal features of the right branch, dynamically calculates the branch weight coefficients α and β based on the prediction confidence, and achieves intelligent fusion through a weighted fusion formula:

[0086]

[0087] This strategy adjusts the branch contribution in real time at specific locations in the fusion layer based on the input complexity, achieving a dynamic balance between fast response and stable output, and overcoming the limitations of the fixed-weight fusion strategy.

[0088] Step Six: Hierarchical Network Integration with Multi-Objective Loss Function

[0089] During model training, multi-objective loss functions are precisely integrated into different layers of the network. Classification loss is applied to the final output layer, boundary loss is specifically applied to the boundary awareness module, and temporal consistency loss is applied to the output of the Actionformer module. This hierarchical loss design ensures that each innovative module in the network receives targeted optimization guidance. Through location-specific loss function design, precise optimization of different network layers is achieved, ensuring classification accuracy while further enhancing temporal consistency and boundary detection accuracy.

[0090] Step 7: Model training and end-to-end optimization, model validation and performance evaluation

[0091] Through an end-to-end training strategy, innovative modules specific to each location work collaboratively to form a unified optimization goal. During training, the dual-branch architecture, boundary-aware module, Actionformer core module, and adaptive response strategy play their optimal roles at their respective specific locations, achieving joint optimization of the entire network through the backpropagation algorithm. This location-coordinated inter-module collaboration optimizes the overall network performance and avoids mutual interference between modules.

[0092] Comprehensive experimental validation was conducted using a self-built oilfield power operation video dataset. The network significantly outperformed mainstream baseline models in key evaluation metrics such as Top-1 accuracy, boundary F1 score, and mAP. Experimental results fully demonstrate the effectiveness of each innovative module at specific network locations, validating the superior performance and broad application prospects of this method in complex industrial behavior recognition tasks.

[0093] Step 8: Use the recognition model to identify oilfield power operation behaviors.

[0094] In one embodiment of the present invention: Step 1: Input data preprocessing and dual-branch architecture initialization:

[0095] A systematic preprocessing method is used for oilfield power operation videos, employing a dual-scale sliding window strategy to accurately segment the original video sequence. This strategy constructs two parallel processing channels: a main window and an auxiliary window. The main window extracts sixteen consecutive frames as a complete time segment, specifically designed to capture complete action context information, ensuring full coverage of the entire execution process of complex power operation actions such as pole climbing, maintenance, and installation. The auxiliary window extracts eight consecutive frames, specifically responsible for detecting rapidly changing action details and instantaneous state transitions. A frame data bag mechanism is implemented for intelligent data distribution, segmenting the input original video sequence V∈R^(N×H×W×C) according to preset time intervals. Each time segment ensures that it contains sufficient frames to maintain the temporal continuity of the actions. Here, N represents the total number of frames, and H, W, and C represent the image height, width, and number of channels, respectively. Furthermore, each time segment undergoes standardization and normalization. The system uses a zero-mean, unit-variance standardization operation.

[0096]

[0097] Where μ is the mean and σ is the standard deviation. This eliminates the impact of varying lighting conditions and weather on video quality. Furthermore, to address common issues in oilfield environments such as light variations and shadow interference, adaptive histogram equalization is used to enhance each frame, improving image contrast and clarity.

[0098] Finally, keyframes for motion are identified and located using an optical flow estimation algorithm. The system calculates the optical flow vectors between adjacent frames, and the optical flow constraint equation is as follows:

[0099]

[0100] in For image gradient, Let u and v be the temporal gradients, and v be the optical flow vectors. By analyzing the changes in the amplitude and direction of the optical flow, the system identifies the beginning, climax, and end phases of the action. For each detected keyframe, the system expands its surrounding frames as contextual information to ensure that the complete process of the action is accurately captured for each time segment.

[0101] In the second embodiment of the present invention: Step 2: Extraction of local spatiotemporal features of the left and right branches:

[0102] In the left branch of the network architecture, an improved R(2+1)D Backbone is used as the backbone network. This module is located at the core processing position of the left branch and is specifically responsible for extracting the local spatiotemporal features of each video segment. This network decomposes the traditional 3D convolution operation into a concatenated combination of 2D spatial convolution and 1D temporal convolution. Traditional 3D convolution operation:

[0103]

[0104] Decomposed into:

[0105]

[0106] W_2D is a 2D spatial convolution kernel, and W_1D is a 1D temporal convolution kernel, significantly reducing computational complexity while maintaining the effectiveness of feature extraction. The network adopts a layer-by-layer progressive approach, gradually extracting high-level semantic features from low-level edge and texture features. To adapt to the large changes in action scale in oilfield power operations, the left branch also integrates a multi-scale feature extraction module, using convolution kernels of different sizes in parallel to capture spatial features at different scales. The left branch outputs a feature tensor. Where H and W are the spatial dimensions after convolutional downsampling, and D is the number of feature channels. The R(2+1)D module effectively captures the global spatiotemporal information of actions by decomposing the convolution operation. This module plays a central role in the left branch, ensuring the purity and integrity of spatial information through a dedicated spatial feature extraction channel, avoiding interference from temporal information in spatial feature learning. The right branch focuses on extracting temporal features from the video. Unlike the spatial feature extraction of the left branch, the right branch emphasizes capturing the temporal changes and temporal information of actions through temporal convolution.

[0107] In the third embodiment of the present invention: the intermediate deployment and implementation method of the boundary-aware detection mechanism described in step three:

[0108] The boundary-aware detection mechanism, as a core intermediate connection module of the network architecture, is precisely deployed in the intermediate connection stage of the network architecture, undertaking the crucial bridging function of connecting temporal features and final prediction. The boundary-aware module employs a boundary-aware detection mechanism, accurately identifying the start and end boundaries of actions through a dual-branch architecture. The main task branch focuses on the extraction and classification of action semantic features. This branch receives temporal features N×C×T×H×W from the dual-branch fusion module, first adjusting the feature dimensions to N×C / 4×1×1 through a 1×1×3 convolutional layer, eliminating the impact of temporal length variations on classification. Finally, the classification layer outputs the probability distribution of each action category. The boundary-aware branch is specifically responsible for learning the transition boundary information between actions. This branch uses the same convolutional configuration to extract temporal features and generates a temporal boundary map through the Softmas activation function. Each element in the boundary graph represents the probability that the corresponding time step is an action boundary.

[0109] The boundary detection submodule calculates the boundary probability:

[0110]

[0111] The system calculates the first and second derivatives of the time-series features along the time dimension, and identifies time points of abrupt feature changes by analyzing the trends of these derivatives. An attention weighting mechanism reweights the features at different time steps based on boundary information.

[0112]

[0113] This represents an element-wise multiplication operation; time steps near the boundary receive higher weights. Features from the two branches are integrated using a soft attention mechanism.

[0114]

[0115] Where λ is a weight parameter that adjusts the importance of boundary information; the system learns a dynamic fusion weight that adaptively adjusts the degree of enhancement of the backbone features by the boundary information according to the characteristics of the input video.

[0116] The training of marginal branches employs an independent supervised learning strategy, with the marginal loss function being:

[0117]

[0118] Where y_boundary is the boundary label and p_boundary is the predicted boundary probability.

[0119] In the fourth embodiment of the present invention: the parallel processing of the ActionFormer prediction module and its implementation method described in step four:

[0120] The ActionFormer prediction module, as a core component of the prediction stage on the right side of the network architecture, is precisely embedded at a specific location after the boundary-aware detection mechanism, and is specifically responsible for processing the temporal features after boundary enhancement. This module achieves global modeling of the action process and accurate capture of local changes through a long-short temporal fusion structure. Combined with multi-scale temporal processing and attention mechanism optimization, it constructs a complete temporal dependency modeling system. The module adopts a hierarchical processing architecture, progressively extracting and optimizing temporal feature representations from the low-level multi-scale temporal convolution to the high-level global temporal modeling.

[0121] The long-short temporal fusion structure achieves feature capture across different time spans through a multi-scale temporal convolutional network:

[0122]

[0123] The dilation rate d is set to 1, 2, and 4, respectively, and the kernel size k is set to 3, 5, and 7, respectively. Kernels with different dilation rates are responsible for capturing immediate action responses, short-term action sequences, and long-term action trends, respectively. This multi-scale processing mechanism is particularly suitable for the characteristics of large differences in action durations in oilfield power operations. Each scale's kernel employs a causal convolution design to ensure that only historical information is used during real-time processing.

[0124] This module processes action sequences with strong temporal continuity, identifying key temporal nodes in the sequence through a self-attention mechanism and preserving temporal order information through a positional encoding mechanism. Working in parallel with the classification module, this module forms a processing structure of "temporal modeling → boundary awareness → parallel prediction." This specific module arrangement ensures that the prediction process is based on boundary-optimized features, significantly improving the modeling capability for long-term temporal data and the recognition accuracy of complex action sequences.

[0125] In the fifth embodiment of the present invention: the deployment and implementation method of the adaptive response strategy for the dual-branch feature fusion layer described in step five:

[0126] The adaptive response strategy module, as a core component in the final fusion stage of the network architecture, is precisely deployed at the convergence point of the two branches, undertaking the crucial function of intelligently fusing spatial and temporal features. This strategy, through an innovative dynamic weight calculation mechanism combined with confidence assessment and complexity analysis, achieves real-time adjustment of the contribution of features from different branches. The module first receives the spatial features of the left branch and the temporal features of the right branch as input, then calculates the prediction confidence through multi-dimensional feature analysis, and finally dynamically determines the branch weight coefficients based on the confidence level and input complexity.

[0127] The confidence score calculation employs a multi-dimensional comprehensive evaluation method. The system comprehensively assesses the reliability of the prediction by analyzing the maximum value of the predicted probability, the sharpness of the distribution, and the degree of consistency between the two branch features. The confidence score of spatial features is calculated based on the peak value and entropy of their predicted probability distribution, while the confidence score of temporal features similarly considers their probability distribution characteristics and temporal continuity. The system also introduces an inter-branch consistency evaluation mechanism, which judges the consistency between the prediction results of two branches by calculating the similarity between spatial and temporal features in the semantic space; higher consistency indicates a more reliable prediction.

[0128]

[0129] The dynamic weight coefficient calculation is based on a comprehensive analysis of confidence level and input complexity. The system first calculates the complexity of spatial features by analyzing the variance and dispersion of the feature vectors to measure the complexity of spatial information. The complexity of temporal features is evaluated by calculating the magnitude of changes in the temporal gradient; more dramatic gradient changes indicate greater temporal complexity. Based on these complexity metrics and the overall confidence level, the system learns dynamic weight coefficients through a neural network, ensuring that the contribution ratio of the two branches can be adaptively adjusted under different scenarios.

[0130] Intelligent fusion achieves the organic integration of features through a weighted fusion formula. The final fused feature not only includes a weighted combination of the two branches but also undergoes further feature transformation and optimization through a multilayer perceptron. This strategy adjusts the branch contribution in real time at specific locations in the fusion layer based on the input complexity. When the spatial scene is complex and the operating environment is variable, the weight of the spatial branch is increased; when the action sequence is complex and the temporal variation is drastic, the weight of the temporal branch is increased. This achieves a dynamic balance between rapid response and stable output, overcoming the limitations of fixed-weight fusion strategies that cannot adapt to different scenario requirements.

[0131] In the sixth embodiment of the present invention: the multi-objective loss function hierarchical network integration and optimization method described in step six:

[0132] The multi-objective loss function employs a precise hierarchical network ensemble strategy, applying different types of loss functions to different layers of the network to provide targeted optimization guidance for each innovative module. This design improves the overall performance of the model by jointly optimizing classification accuracy, temporal consistency, and boundary detection precision. The classification loss is specifically applied to the final output layer, responsible for optimizing the accuracy of action category recognition. Considering the class imbalance problem in oilfield power operations, the system adopts a weighted cross-entropy loss function, setting different weights based on the frequency of each category in the training set to ensure that rare categories receive sufficient attention.

[0133] Boundary loss is specifically applied to the boundary awareness module, employing an improved intersection-over-union (IoU) loss function to optimize the prediction accuracy of action boundaries. This loss function not only considers the overlap between the predicted and true boundaries but also adds an absolute error term for the boundary position, ensuring both the overlap range and precise control over boundary position deviation. To handle boundary prediction for actions of varying lengths, the system also introduces a scale-adaptive mechanism, dynamically adjusting the sensitivity of the loss function based on the action's duration.

[0134]

[0135] Temporal consistency loss is applied to the output of the ActionFormer module to ensure the continuity and stability of the model's predictions over time. This loss function measures the smoothness of the temporal sequence by calculating the differences in features between adjacent time steps, while introducing a second derivative term to constrain abrupt changes in the prediction results. The system also employs temporal smoothing regularization techniques to improve prediction stability by limiting sharp changes in the predicted probabilities over time.

[0136] The overall loss function employs a dynamic weighting strategy. The system monitors the convergence speed and magnitude of each loss function, automatically increasing its weight when it converges slowly and appropriately decreasing its weight when it becomes too large to avoid gradient explosion. This dynamic adjustment mechanism ensures the stability of the training process and the balanced development of various optimization objectives. Through location-specific loss function design, precise optimization of different network layers is achieved, ensuring classification accuracy while further enhancing temporal consistency and boundary detection accuracy.

[0137] In the seventh embodiment of the present invention: the model training and end-to-end optimization described in step seven and their implementation methods:

[0138] The end-to-end optimization strategy utilizes a unified training framework to enable collaborative work among innovation modules specific to each location, forming a unified optimization objective. The training process employs a multi-stage optimization strategy: first, each module undergoes independent pre-training to achieve good initial performance on its specific task; then, joint optimization of the entire network is performed, with end-to-end backpropagation achieving synergistic improvement across all modules. The optimizer uses the AdamW algorithm, incorporating momentum terms and weight decay to enhance training stability and generalization performance.

[0139] The learning rate scheduling employs a warm-up and cosine annealing strategy. In the early stages of training, a small learning rate is used for warm-up to avoid initial gradient instability. This is then gradually increased to a pre-set base learning rate. Later in training, the learning rate is gradually decreased using a cosine function to ensure the model converges to a better local optimum. To prevent gradient explosion, the system uses gradient clipping; when the gradient norm exceeds a preset threshold, it is normalized.

[0140]

[0141] Regularization strategies include L1 and L2 regularization, which improve generalization ability by limiting the complexity of model parameters. The system also employs Dropout and batch normalization techniques, randomly discarding some neuron connections during training to prevent overfitting. Data augmentation techniques increase the diversity of training samples through random pruning, rotation, and scaling, improving the model's adaptability to different scenarios. During training, the dual-branch architecture, boundary-aware module, ActionFormer core module, and adaptive response strategy play their optimal roles at their respective specific locations, achieving joint optimization of the entire network through backpropagation, thus avoiding mutual interference between modules.

[0142] The evaluation metrics system includes traditional classification metrics such as accuracy, precision, recall, and F1 score, as well as metrics specifically designed for action recognition tasks, such as temporal IoU and boundary detection F1 score. The system also uses mean average precision (mAP) to comprehensively evaluate the model's performance across all categories. To verify the model's robustness, experiments include evaluations in various scenarios such as noise robustness testing, illumination variation testing, and occlusion testing.

[0143] Experimental results show that the network achieves a Top-1 accuracy of 94.7%, a boundary F1 score of 89.3%, and an mAP of 92.1%, all significantly outperforming current mainstream baseline models. Ablation experiments, conducted by systematically removing various innovative modules, validated its effectiveness. Results showed that the boundary awareness module contributed a 3.2% improvement in accuracy, the ActionFormer module contributed a 2.8% improvement, and the adaptive response strategy contributed a 1.9% improvement. Long-term stability testing demonstrated that the system maintained stable recognition performance after 72 hours of continuous operation, fully proving the reliability and practical value of this method in real-world industrial applications.

[0144] like Figure 1 As shown, this invention proposes a dual-branch temporal recognition network architecture based on boundary awareness and adaptive response mechanisms. This network is specifically designed for accurate identification of worker pole-climbing behavior in oilfield power operation scenarios. The architecture employs an innovative dual-branch design, effectively addressing technical challenges such as small differences in fine-grained actions, ambiguous state transitions, and strong temporal continuity. The specific structure is as follows:

[0145] The network employs a dual-branch parallel processing architecture. The left branch is based on a boundary-aware detection mechanism, while the right branch is based on the ActionFormer prediction module. The input video data format is C×H×W, where C represents the number of channels, and H and W represent the height and width of the video frame, respectively. The video sequence first passes through a visual segmentation module, converting consecutive video frames into temporal segments suitable for network processing.

[0146] The left branch extracts local spatiotemporal features using the R3D Backbone, a module that effectively captures 3D spatiotemporal features from videos. Let the input features F∈R^(C×T×H×W), where T represents the time dimension. The R3D feature extraction process can be represented as:

[0147]

[0148] Where W_3d represents the parameters of the 3D convolution kernel.

[0149] Deep features are used for temporal feature extraction and advanced temporal change processing. LSTM is responsible for capturing long-term temporal dependencies, and the Attention mechanism is used to highlight important temporal features.

[0150] Finally, a dual-branch feature fusion module is obtained through weighted fusion and Conv processing. The fusion process can be represented as follows:

[0151]

[0152] Here, α and β are learnable weight parameters used to balance the contributions of the two branches. The boundary-aware detection mechanism of the right branch includes boundary detection and ActionFormer prediction modules, which perform temporal action localization and classification recognition through Proposal Head and ClassificationHead.

[0153] like Figure 2 The diagram shows a detailed architecture of the ActionFormer dual-branch network based on R(2+1)D, illustrating the network's core processing flow and key technical components. This architecture integrates the advantages of R(2+1)D spatiotemporal decomposition convolution with ActionFormer's temporal modeling capabilities, achieving accurate recognition of complex industrial behaviors. Input video data is first preprocessed by the R(2+1)D spatiotemporal network. R(2+1)D decomposes traditional 3D convolution into spatial and temporal convolutions, improving computational efficiency and feature representation capabilities. The decomposition process can be represented as:

[0154]

[0155] Where Conv2D is used to process spatial features, and Conv1D_temporal is used to process temporal features.

[0156] The backbone network consists of two processing branches: a boundary-aware branch and an ActionFormer prediction branch. The boundary-aware branch extracts boundary features at different resolutions using a multi-scale feature pyramid structure. The multi-scale feature fusion formula is as follows:

[0157]

[0158] Where F_i represents the feature map of the i-th layer, w_i is the corresponding weight, and scale_i is the upsampling factor. The core of ActionFormer implements a multi-head self-attention mechanism, which includes three parallel branches: local temporal modeling capability, cross-segment temporal attention capability, and periodic temporal attention capability.

[0159] The three branches correspond to feature extraction at different scales (scale 1 / 8, scale 1 / 4, scale 1 / 2, and scale 1 / 1), achieving multi-level feature representation from fine-grained to coarse-grained. The construction process of the feature pyramid is as follows:

[0160]

[0161] Where 'l' represents the pyramid level and 'Pool' represents the pooling operation. The detection results are finally output through the proposal head and classification head.

[0162] like Figure 3 The diagram shows the detailed structure of the dual-branch fusion module, a key component of the entire network architecture responsible for the effective fusion and information exchange of features from different branches. Through a carefully designed fusion strategy, this module significantly improves the network's ability to understand and model complex action sequences. The dual-branch fusion module comprises three main processing stages: data preprocessing, feature fusion, and output prediction. In the data preprocessing stage, the input video features are standardized:

[0163]

[0164] Where μ and σ are the mean and standard deviation of the features, respectively. The feature fusion stage employs an advanced attention fusion mechanism, using Conv3D for 3D convolution processing to handle feature maps in C×T×H×W format. During the fusion process, the correlation matrix between the two branch features is first calculated:

[0165]

[0166] Then, fusion weights are generated using a soft attention mechanism:

[0167]

[0168] The formula for calculating the fused features is:

[0169] in This indicates an element-wise multiplication operation.

[0170] like Figure 4The diagram shows the network structure of the boundary-aware detection mechanism, one of the core innovations of this invention, specifically designed to solve the technical challenges of blurred action boundaries and difficulty in recognizing state transitions. This mechanism significantly improves the detection accuracy of action transition points through intelligent attention allocation and boundary enhancement strategies. The boundary-aware mechanism employs a dual-path processing strategy, where input N×C×T×H×W format video data simultaneously enters both the main feature extraction branch and the boundary-aware feature extraction branch. Here, N represents the batch size, and the processing strictly maintains the spatiotemporal consistency of the data. The main branch extracts features using 1×1×3 Conv3D convolutional kernels; this design effectively captures temporal information while maintaining spatial resolution.

[0171]

[0172] The convolution parameters are calculated as follows:

[0173]

[0174] Finally, boundary enhancement is achieved through the attention mechanism. The enhanced feature output in N×C×T×H×W format not only maintains the original spatiotemporal structure, but also significantly enhances the feature representation ability of the boundary region, providing a more accurate feature foundation for subsequent action recognition.

[0175] like Figure 5 The image shows a real-world application scenario of an oilfield power operation, comprehensively demonstrating the actual industrial environment and application background targeted by this invention. The image includes photos of the operation site from four different angles and scenarios, fully verifying the applicability and robustness of the network of this invention in complex real-world environments.

[0176] These on-site images showcase the typical characteristics of power operations in oil fields: including power poles of varying heights (generally 10-30 meters), complex equipment structures, variable weather conditions (sunny days, snowy days, and other varying lighting conditions), and diverse scenarios such as workers wearing different colored work clothes. This environmental complexity poses a significant challenge to video recognition algorithms, requiring them to possess strong environmental adaptability.

[0177] From a technical perspective, the complexity of these scenarios is mainly reflected in the following aspects:

[0178] The effects of changes in light intensity can be represented by the illuminance equation:

[0179]

[0180] Where a_t represents the action state at time t.

[0181] The interface diagram of the motion detection system showcases the practical application software system based on the network of this invention, fully demonstrating the achievement of translating theoretical research into practical application. The system utilizes the professional design framework of ActionVision Pro, providing on-site safety supervisors with an intuitive and efficient visual operation platform. The system interface adopts a modular design concept, mainly including the following core functional modules:

[0182] The video playback area is located in the center of the interface, supporting real-time playback and analysis of various video formats. Playback control functions include basic operations such as play, pause, fast forward, and slow motion, with a frame rate control range of 1-60fps and supported video resolutions from 480p to 4K. The detection results display area shows the real-time output of the recognition algorithm, including key information such as action category, confidence score, and temporal boundaries. The confidence score calculation formula is:

[0183]

[0184] Boundary detection is considered correct when IoU > 0.5. The parameter setting panel allows users to adjust detection parameters according to actual needs, including:

[0185] Detection threshold setting: Threshold ∈ [0.1, 0.9]

[0186] Timing window size: Window_size ∈ [16, 128] frames

[0187] Non-maximum suppression parameter: NMS_threshold ∈ [0.3, 0.7]

[0188] The system also integrates real-time performance monitoring, displaying key performance indicators:

[0189]

[0190] This network architecture achieves high-precision identification of workers climbing poles during oilfield power operations by organically combining boundary-aware mechanisms and adaptive response strategies. Regarding adaptive response, the system dynamically adjusts the detection sensitivity based on confidence levels.

[0191]

[0192] Here, α is the adjustment coefficient, achieving an optimal balance between rapid response and stable output. The system interface not only provides complete detection functions but also advanced features such as data export, report generation, and historical playback, offering comprehensive technical support for industrial safety management. Through this system, safety supervisors can monitor the work site in real time, promptly identify safety hazards, and significantly improve the level of intelligence in operational safety management.

[0193] To verify the effectiveness and superiority of the proposed dual-branch temporal recognition network based on boundary awareness and adaptive response mechanism in oilfield power operation behavior recognition tasks, a comparative experiment was conducted on a self-built oilfield-related power operation video dataset. The experimental results are shown in Table 1.

[0194] Table 1. Comparison of Experimental Results

[0195] Model Accuracy F1-Score Diversity T3D 59.39% 57.07% 4 / 4 R3D 61.49% 58.79% 4 / 4 T3D+TCN 87.02% 87.12% 4 / 4 This invention 97.57% 97.58% 4 / 4

[0196] The experiments used the same dataset and evaluation metrics to compare several mainstream behavior recognition methods. T3D (Temporal 3D ConvNet) served as a classic baseline method for 3D convolutional networks. R3D (Residual 3D ConvNet) introduced a residual connection mechanism on top of T3D, while T3D+TCN enhanced temporal modeling capabilities by adding a temporal convolutional network to the backend of the T3D network. All comparison methods were trained and tested in the same hardware environment to ensure the fairness and comparability of the experimental results.

[0197] Table 1 clearly shows that the proposed dual-branch temporal recognition network based on boundary awareness and adaptive response mechanisms significantly outperforms the comparative methods in all key evaluation metrics. In terms of accuracy, our method achieves 97.57%, a 10.55 percentage point improvement over the best baseline method T3D+TCN's 87.02%, representing a relative improvement of 12.13%. In the F1-Score, our method achieves an excellent score of 97.58%, a 10.46 percentage point improvement over T3D+TCN's 87.12%, fully demonstrating the balanced performance of our method in terms of precision and recall.

[0198] Of particular note is that all methods achieved a perfect score of 4 / 4 in the Diversity metric evaluation, indicating that all methods can effectively identify all four key action categories in oilfield power operations, including pole preparation, climbing process, operation, and descent. This result validates the reasonableness of the dataset and the basic effectiveness of the various methods.

[0199] The comparative analysis of the experimental results shows that traditional T3D and R3D methods are relatively weak in handling industrial behavior recognition tasks with fine-grained differences and fuzzy state transitions due to the lack of targeted temporal modeling and boundary awareness mechanisms. The T3D+TCN method improves the temporal modeling capability to some extent by introducing a temporal convolutional network, thus significantly improving performance, but it still lacks the ability to accurately perceive action boundaries.

[0200] In contrast, the method proposed in this study, by integrating a boundary awareness mechanism and an adaptive response strategy, can not only accurately capture the global temporal features of action sequences but also precisely locate the critical moments of action transitions. This results in more accurate and stable behavior recognition in complex industrial scenarios. Experimental results fully verify the practicality of the technical solution of this invention, providing reliable technical support for the intelligent upgrading of industrial safety supervision.

[0201] To verify the effectiveness of the core technical components of the dual-branch temporal recognition network based on boundary awareness and adaptive response mechanism proposed in this invention, the experimental results are shown in Table 2.

[0202]

[0203] Table 2 shows that the dual-branch fusion architecture, as the basic technical framework of this invention, achieved accuracy and F1-Score of 97.57% and 97.58% respectively, demonstrating its superior performance in handling complex temporal action recognition tasks. This architecture effectively improves the network's ability to model multi-scale spatiotemporal information by processing spatial and temporal features in parallel.

[0204] The validation results of the boundary-aware mechanism show that it achieves performance of 86.30% and 86.15% in accuracy and F1-Score, respectively. Experiments demonstrate that the boundary-aware mechanism, by introducing attention weight calculation, can effectively identify key frames of action transitions, significantly improving the shortcomings of traditional methods in action boundary detection. Experimental validation of the multi-objective loss function shows that this design achieves 89.55% and 89.20% in accuracy and F1-Score, respectively. The multi-objective loss function achieves coordinated optimization of multiple performance metrics during network training by simultaneously optimizing classification loss, boundary regression loss, and temporal consistency loss. The validation results of the adaptive response strategy show that this strategy achieves 96.84% in both accuracy and F1-Score. This strategy dynamically adjusts the system response threshold based on detection confidence.

[0205] Experimental results further demonstrate a significant synergistic effect among the various technical components. When all core technical components are integrated, the overall system performance reaches its optimal level, verifying the systematic nature and completeness of the technical solution of this invention. Compared to using each technical component individually, the performance of the complete technical solution is improved, fully verifying the practicality of the proposed dual-branch temporal sequence recognition network based on boundary awareness and adaptive response mechanisms. Each core technical component exhibits good performance contributions, providing a reliable technical solution for the safety supervision of oilfield power operations.

Claims

1. An intelligent identification method based on safe behaviors in oilfield power operations, characterized in that: The oilfield power operation behavior recognition method based on a dual-branch temporal recognition network with boundary awareness and adaptive response mechanism includes the following steps: Step 1: Preprocess the oilfield power operation video by dividing the original video sequence into fixed-length segments according to the time dimension. Each segment contains a continuous sequence of frames. Construct a dual-branch parallel processing channel in the input layer of the dual-branch temporal recognition network architecture: a left branch channel and a right branch channel. Step 2: Extraction of local spatiotemporal features of the left and right branches; The left branch channel uses R(2+1)D Backbone as the backbone network, while the right branch channel captures the temporal continuity in the video, ensuring that the temporal changes of the action are effectively captured. Step 3: Deployment of the boundary awareness detection mechanism in the middle; In the intermediate connection stage of the dual-branch temporal recognition network architecture, a boundary-aware detection mechanism is deployed. The boundary-aware detection mechanism calculates the boundary weight score of each frame through an attention mechanism, identifies the start and end boundaries of the action, and outputs a feature representation with boundary enhancement information. Step 4: The Actionformer prediction module processes data in parallel; In the dual-branch temporal recognition network architecture, after the boundary-aware detection mechanism, the Actionformer prediction module is precisely embedded. The Actionformer prediction module receives the enhanced features output by the boundary-aware module, performs global temporal modeling through a long and short temporal fusion structure, and optimizes the focusing of temporal information by combining an attention mechanism to process behavior sequences with strong temporal continuity. The Actionformer module works in parallel with the classification module. Step 5: Deploy the adaptive response strategy for the dual-branch feature fusion layer; In the final fusion stage of the dual-branch temporal recognition network architecture, an adaptive response strategy module is introduced to form the recognition model. This module is located at the convergence point of the two branches. The adaptive response strategy receives the spatial features of the left branch channel and the temporal features of the right branch channel, dynamically calculates the branch weight coefficients α and β based on the prediction confidence, and achieves intelligent fusion through a weighted fusion formula: ; This strategy module adjusts the branch contribution in real time at a specific location in the fusion layer based on the input complexity, achieving a dynamic balance between fast response and stable output. Step Six: Hierarchical network integration of multi-objective loss functions; During model training, the multi-objective loss function is precisely integrated into different layers of the network. The classification loss is applied to the final output layer, the boundary loss is applied to the boundary awareness module, and the temporal consistency loss is applied to the output of the Actionformer prediction module. Step 7: Identify model training, end-to-end optimization, and validation; Through an end-to-end training strategy, all modules work together to form a unified optimization goal; Step 8: Use the recognition model to identify oilfield power operation behaviors.

2. The intelligent identification method based on oilfield power operation safety behavior according to claim 1, characterized in that: The method for preprocessing the oilfield power operation video in step one is as follows: A dual-scale sliding window strategy is employed to segment the original video sequence. This strategy constructs two parallel processing channels: a main window and an auxiliary window. The main window is configured to extract sixteen consecutive frames as a complete time segment, capturing complete action context information and fully covering the entire execution process of complex power operation actions. The auxiliary window is configured to extract eight consecutive frames, responsible for detecting rapidly changing action details and instantaneous state transitions. A frame data bag mechanism is implemented for intelligent data distribution processing. The input original video sequence V∈R^(N×H×W×C) is segmented according to a preset time interval. Each time segment ensures that it contains enough frames to maintain the temporal continuity of the action, where N represents the total number of frames, and H, W, and C represent the height, width, and number of channels of the image, respectively. Furthermore, each time segment is standardized and normalized. The system uses a zero-mean, unit variance standardization operation. ; Where μ is the mean and σ is the standard deviation, an adaptive histogram equalization technique is used to enhance each frame of the image to improve the contrast and clarity of the image, taking into account the common light changes and shadow interference in the oilfield operating environment. Finally, the system identifies and locates key frames of motion using an optical flow estimation algorithm, calculates the optical flow vectors between adjacent frames, and the optical flow constraint equation is: ; in For image gradient, The time gradient is denoted as u and v, which are optical flow vectors. The starting, climax, and ending phases of the action are identified by analyzing the changes in the amplitude and direction of the optical flow.

3. The intelligent identification method based on oilfield power operation safety behavior according to claim 2, characterized in that: Step two specifically involves: The left branch channel uses an improved R(2+1)D Backbone as the backbone network, decomposing the traditional 3D convolution operation into a cascaded combination of 2D spatial convolution and 1D temporal convolution: ; Where W_2D is a 2D spatial convolution kernel, W_1D is a 1D temporal convolution kernel, and the left branch channel also integrates a multi-scale feature extraction module, which captures spatial features of different scales by using convolution kernels of different sizes in parallel; the left branch outputs a feature tensor. Where H and W are the spatial dimensions after convolution downsampling, and D is the number of feature channels. By decomposing the convolution operation, the global spatiotemporal information of the action is effectively captured; the right branch extracts the temporal features in the video.

4. The intelligent identification method based on oilfield power operation safety behavior according to claim 3, characterized in that: Step three specifically involves: The boundary-aware detection mechanism is deployed in the intermediate connection stage of the dual-branch temporal recognition network architecture. This dual-branch architecture accurately identifies the start and end boundaries of actions. The main task branch focuses on extracting and classifying the semantic features of the actions. It receives temporal features N×C×T×H×W from the dual-branch fusion module, first adjusting the feature dimensions to N×C / 4×1×1 using a 1×1×3 convolutional layer, and then outputting the probability distribution of each action category through a classification layer. The boundary-aware branch is responsible for learning the transition boundary information between actions, using the same convolutional configuration to extract temporal features, and generating a temporal boundary map through the Softmas activation function. Each element in the boundary graph represents the probability that the corresponding time step is an action boundary. The boundary detection submodule calculates the boundary probability: ; The first and second derivatives of the time-series features are calculated along the time dimension. By analyzing the trends in the derivatives, time points where the features change drastically are identified. An attention weighting mechanism reweights the features at different time steps based on boundary information. ; This represents the element-wise multiplication operation; Features from the main task branch and the boundary awareness branch are integrated through a soft attention mechanism: ; Where λ is a weighting parameter that adjusts the importance of boundary information; The training of marginal branches employs an independent supervised learning strategy, with the marginal loss function being: ; Where y_boundary is the boundary label and p_boundary is the predicted boundary probability.

5. The intelligent identification method based on oilfield power operation safety behavior according to claim 4, characterized in that: Step four specifically involves: The ActionFormer prediction module is embedded after the boundary-aware detection mechanism. It is responsible for processing the temporal features after boundary enhancement. It achieves global modeling of the action process and accurate capture of local changes through a long and short temporal fusion structure. Combined with multi-scale temporal processing and attention mechanism optimization, it constructs a complete temporal dependency modeling system. It adopts a hierarchical processing architecture, from the low-level multi-scale temporal convolution to the high-level global temporal modeling, to gradually extract and optimize the temporal feature representation. The long-short temporal fusion structure achieves feature capture across different time spans through a multi-scale temporal convolutional network: ; Key temporal nodes in the sequence are identified through a self-attention mechanism, and temporal order information is maintained through a positional encoding mechanism. This process works in parallel with the classification module, forming a processing structure: temporal modeling → boundary awareness → parallel prediction.

6. The intelligent identification method based on oilfield power operation safety behavior according to claim 5, characterized in that: Step five specifically involves: As a core component in the final fusion stage of the network architecture, the adaptive response strategy module is deployed at the convergence point of the two branches. It undertakes the important function of intelligently fusing spatial and temporal features. Through a dynamic weight calculation mechanism, combined with confidence assessment and complexity analysis, it realizes real-time adjustment of the contribution of different branch features. First, it receives the spatial features of the left branch channel and the temporal features of the right branch channel as input. Then, it calculates the prediction confidence through multi-dimensional feature analysis. Finally, it dynamically determines the branch weight coefficient based on the confidence and input complexity. The confidence level calculation adopts a multi-dimensional comprehensive evaluation method. The dynamic weight coefficient calculation is based on a comprehensive analysis of confidence level and input complexity. First, the complexity of spatial features is calculated by analyzing the variance and dispersion of feature vectors to measure the complexity of spatial information. The complexity of temporal features is evaluated by calculating the magnitude of change in temporal gradient. The more dramatic the gradient change, the more complex the temporal information. Based on the complexity index and the overall confidence level, dynamic weight coefficients are obtained through neural network learning to ensure that the contribution ratio of the two branches can be adaptively adjusted under different scenarios. Intelligent fusion achieves the organic integration of features through a weighted fusion formula. The final fused feature contains a weighted combination of the two branches, and feature transformation and optimization are performed through a multilayer perceptron. In the fusion layer, the contribution of the branches is adjusted in real time according to the input complexity. When the spatial scene is complex and the working environment is variable, the weight of the spatial branch is increased. When the action sequence is complex and the temporal change is drastic, the weight of the temporal branch is increased, so as to achieve a dynamic balance between fast response and stable output.

Citation Information

Patent Citations

  • Remote sensing large model construction method fusing edge semantic knowledge

    CN120726320A

  • Target detection method and system for weak and small target in remote sensing image

    WO2025118826A1