Intelligent Alarm Method and System Based on Environmental Perception
Features are extracted through three-dimensional lightweight convolution and hierarchical attention modules, and the three-dimensional rotary transformer modules are used to fusion features, solving the problems of large amount of calculation and poor feature fusion in the prior art, real-time and high-accurate behavior recognition warning is achieved.
Patent Information
- Application Number
- CN202510429252.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-08
AI Technical Summary
In the existing behavior recognition and early warning technology, the three-dimensional convolutional network has a large amount of computation, high memory occupancy, which affects real-time performance, and the three-dimensional transformer model cannot effectively integrate spatio-temporal features, resulting in a decrease in accuracy.
Features are extracted using three-dimensional lightweight convolution modules and three-dimensional hierarchical attention modules, and feature fusion is performed by combining three-dimensional rotation transformer modules. Through the multi-head self-attention mechanism of the three-dimensional sliding window and the shift sliding window, the calculation efficiency is optimized and global and local features are captured.
It improves the real-time and accuracy of behavior recognition warnings, can efficiently identify human behaviors in the video and issue warnings in a timely manner to prevent dangers from occurring.
Smart Images

Figure CN119942654B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence applications, and particularly to an intelligent alarm method and system based on environmental perception. Background Art
[0002] Environmental perception refers to the process of perceiving and understanding dynamic information in the surrounding environment through sensors and algorithms. Behavior recognition refers to the process of inferring the intention, state, or characteristics of a human or other entity by analyzing and understanding their actions, behaviors, and related environmental information. Behavior recognition early warning refers to the process in which, based on behavior recognition, an early warning system detects potential risks, anomalies, or dangerous situations in advance by judging and analyzing certain specific behaviors or actions, and issues an early warning in a timely manner. Behavior recognition early warning technology mainly relies on algorithms such as machine learning, deep learning, and pattern recognition to analyze and judge the collected data, and can identify abnormal behaviors such as running fast.
[0003] The key issues in behavior recognition are how to extract local and global features in space-time simultaneously and how to fuse features. In recent years, researchers have proposed using 3D convolution and 3D transformers to extract and fuse local and global features in space-time. The convolutional network can effectively extract local information from the relevant feature maps through neighborhood convolution operations. However, the limited receptive field of the convolutional network makes it difficult to capture global context information. Based on the weighted average operation of the context information of the input features by the attention transformer, global information can be effectively captured. However, blind similarity matching of adjacent image patches may lead to a high false matching rate.
[0004] Behavior recognition early warning requires real-time performance and accuracy. Generally, the 3D convolutional network and 3D attention network used for feature extraction have a large computational amount and high memory occupancy, which will affect the real-time performance of behavior recognition early warning. Generally, the 3D transformer used for feature fusion fuses local and global features in space-time. However, the 3D transformer model is large, has a large computational amount, and cannot fuse global feature information and local feature information in space-time well, which will reduce the accuracy of behavior recognition early warning. Summary of the Invention
[0005] This application aims to solve at least one of the technical problems in the related technologies to some extent. To this end, an object of this application is to propose an intelligent alarm method and system based on environmental perception to improve the real-time performance and accuracy of behavior recognition early warning.
[0006] The first aspect disclosed in this application provides an intelligent alarm method based on environmental perception, and the specific steps include:
[0007] Obtain the video captured by the environmental perception camera, initialize the video, and divide the video into T frame pictures;
[0008] Feature extraction is performed on the T-frame images based on a behavior recognition feature extraction network to obtain hierarchical features, where the behavior recognition feature extraction network consists of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module;
[0009] Based on a three-dimensional rotator module, the hierarchical features are fused to obtain fused features;
[0010] The fused features are input into a behavior classification network for behavior recognition and classification to obtain a behavior classification result, and the behavior classification result is input into an early warning network;
[0011] A threshold is set in the early warning network. When the behavior classification result is greater than or equal to the threshold, it is determined as a fast running behavior, that is, an early warning is triggered. When the behavior classification result is less than the threshold, the next time series video is returned for acquisition.
[0012] The steps of acquiring the video captured by the environmental perception camera, initializing the video, and dividing the video into T-frame images include:
[0013] Collect the human behavior video in the dynamic scene through the camera and initialize it;
[0014] The video is divided into T-frame images, and the size of the T-frame images is T×H×W×3, where T represents the time dimension, H represents the height, W represents the width, and 3 represents the number of color channels.
[0015] The steps of performing feature extraction on the T-frame images based on the behavior recognition feature extraction network to obtain hierarchical features include:
[0016] The T-frame images are sequentially input into the three-dimensional lightweight convolution module in the behavior recognition feature extraction network;
[0017] Use Convolution calculation is performed to obtain the features of the T-frame images , where the features of the T-frame images include local features and global features ;
[0018] The features of the T-frame images are input into the three-dimensional hierarchical attention module for extraction of hierarchical features to obtain hierarchical features.
[0019] The three-dimensional rotator module includes a first block and a second block. The step of fusing the hierarchical features based on the three-dimensional rotator module to obtain fused features includes:
[0020] The hierarchical features first pass through the first block and then through the second block to obtain fused features; both the first block and the second block first go through a layer normalization, then through a 3D multi-head self-attention, then through another layer normalization, then through a 3D lightweight depth convolution, and finally through a feed-forward neural network; the difference between the first block and the second block is that the first block uses 3D sliding window multi-head self-attention and the second block uses 3D shifted sliding window multi-head self-attention.
[0021] The specific process of the first block using 3D sliding window multi-head self-attention includes:
[0022] Take the hierarchical features output by the action recognition feature extraction network as input data;
[0023] Perform a linear transformation on the normalized data and project it into the query, key, and value spaces respectively to obtain the query space mapping value, key space mapping value, and value space mapping value;
[0024] Divide the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum;
[0025] After completing the attention calculation and weighted sum operations of all heads, concatenate the output results of multiple heads along the channel dimension to obtain the output of the 3D sliding window multi-head self-attention.
[0026] The specific process of the second block using 3D shifted sliding window multi-head self-attention includes:
[0027] Take the output of the first block after normalization as the input data of the second block;
[0028] Perform a linear transformation on the normalized data and project it into the query, key, and value spaces respectively to obtain the query space mapping value, key space mapping value, and value space mapping value;
[0029] Divide the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum;
[0030] After completing the attention calculation and weighted sum operations of all heads, concatenate the output results of multiple heads along the channel dimension to obtain the output of the 3D shifted sliding window multi-head self-attention.
[0031] The process of the 3D transformer module fusing hierarchical features to obtain fused features also includes:
[0032] Define two parameters: window size adjustment factor and stride adjustment factor;
[0033] Calculate the displacement of pixels between adjacent frames by the optical flow method, and then obtain the average motion speed;
[0034] Measure the feature complexity by calculating the entropy of the feature map;
[0035] Adjust the window size and step size of the three-dimensional sliding window and the three-dimensional shifted sliding window in the three-dimensional transformer module according to the average motion speed and the feature complexity.
[0036] The method for fusing hierarchical features by the three-dimensional transformer module further includes:
[0037] Fuse the input of the first block and the output of the first block by addition or concatenation to obtain fused input data;
[0038] Replace the input of the second block with the fused input data.
[0039] The method for fusing hierarchical features by the three-dimensional transformer module further includes:
[0040] Perform convolutional processing on the spatial dimension of the input data of the first block by two-dimensional convolution to extract spatial local feature data;
[0041] Replace the three-dimensional sliding window multi-head self-attention with a two-dimensional sliding window multi-head self-attention, and replace the three-dimensional shifted sliding window multi-head self-attention with a one-dimensional shifted sliding window multi-head self-attention;
[0042] Use the spatial local feature data as the input data of the first block again;
[0043] Perform a convolutional operation on the output of the first block in the time dimension to extract time series features;
[0044] Fuse the output of the first block and the output of the second block to obtain a fused feature.
[0045] The second aspect of the present application discloses an intelligent alarm system based on environmental perception, and the system includes:
[0046] A video acquisition and initialization module, configured to acquire a video captured by an environmental perception camera, initialize the video, and divide the video into T frame pictures;
[0047] A behavior recognition feature extraction module, configured to extract hierarchical features from the T frame pictures based on a behavior recognition feature extraction network, where the behavior recognition feature extraction network is composed of a three-dimensional lightweight convolutional module and a three-dimensional hierarchical attention module;
[0048] A three-dimensional rotation transformer feature fusion module is used to fuse hierarchical features based on a three-dimensional rotation transformer module to obtain fused features;
[0049] A classification module is used to input the fused features into the classification module for recognition and classification to obtain a behavior classification result, and input the behavior classification result into an early warning network;
[0050] An early warning module is used to set a threshold in the early warning network. When the behavior classification result is greater than or equal to the threshold, it is determined as a fast running behavior, that is, an early warning is triggered. When the behavior classification result is less than the threshold, it returns to obtain the next time series video.
[0051] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0052] An intelligent alarm method and system based on environmental perception provided in this application solve the problems of large computational complexity of the behavior feature extraction network, inability to accurately capture local feature information, and inability of the feature fusion network to better fuse global feature information and local feature information in space and time in existing behavior recognition and early warning. It efficiently realizes the recognition of human behaviors in videos and issues an early warning for fast running behaviors to prevent danger from occurring. Description of the Drawings
[0053] Figure 1 is a schematic flowchart of an intelligent alarm method based on environmental perception provided by an embodiment of this application;
[0054] Figure 2 is a schematic structural diagram of a behavior recognition feature extraction network framework in an intelligent alarm method based on environmental perception provided by an embodiment of this application;
[0055] Figure 3 is a schematic structural diagram of a three-dimensional rotation transformer module for feature fusion in an intelligent alarm method based on environmental perception provided by an embodiment of this application;
[0056] Figure 4 is a schematic diagram of an intelligent alarm system based on environmental perception provided by an embodiment of this application. Detailed Embodiments
[0057] To better understand this application, more detailed descriptions will be made for various aspects of this application with reference to the accompanying drawings. It should be understood that these detailed descriptions are only descriptions of exemplary embodiments of this application and do not limit the scope of this application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.
[0058] In the accompanying drawings, for ease of illustration, the sizes, dimensions, and shapes of the elements have been slightly adjusted. The accompanying drawings are for illustrative purposes only and are not drawn to an exact scale. As used herein, terms such as "substantially", "approximately", and similar terms are used as terms of approximation and not as terms of degree, and are intended to account for the inherent deviations in measured or calculated values that would be recognized by a person of ordinary skill in the art.
[0059] It should also be understood that expressions such as "comprising", "including", "having", "containing", and / or "including having" are open-ended rather than closed-ended expressions in this specification, which mean that there are the stated features, elements, and / or components, but do not exclude the presence of one or more other features, elements, components, and / or combinations thereof. In addition, when an expression such as "at least one of..." appears after a list of listed features, it modifies the entire list of features rather than just individual elements in the list. In addition, when describing embodiments of the present application, the use of "may" means "one or more embodiments of the present application". And the term "exemplary" is intended to refer to an example or illustration.
[0060] Unless otherwise defined, all terms used herein (including engineering terms and scientific and technical terms) have the same meaning as commonly understood by a person of ordinary skill in the art to which this application belongs. It should also be understood that unless clearly stated in this application, words defined in a common dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense.
[0061] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0062] Embodiment 1
[0063] Figure 1 FIG. is a schematic flow chart of an intelligent alarm method based on environmental perception provided by an embodiment of the present application. As Figure 1 shown, the method includes:
[0064] Obtain the video captured by the environmental perception camera, initialize the video, and divide the video into T frame pictures. Collect the human behavior video in the dynamic scene through the camera and initialize it, and divide the video into T frame pictures. The size of the T frame pictures is T×H×W×3, where T represents the time dimension, H represents the height, W represents the width, and 3 represents the number of color channels. Input the T frame pictures into the behavior feature extraction network in sequence.
[0065] The hierarchical features are extracted from the T-frame pictures based on the behavior recognition feature extraction network, where the behavior recognition feature extraction network consists of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module.
[0066] Figure 2 It is a schematic structural diagram of the framework of the behavior recognition feature extraction network in the intelligent alarm method based on environmental perception provided by an embodiment of the present application. As Figure 2 shown, the behavior recognition feature extraction network consists of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module. The T-frame pictures are input into the three-dimensional lightweight convolution module, and convolution calculation is used to obtain the features of the T-frame pictures , where the features of the T-frame pictures include local features and global features . The three-dimensional lightweight convolution module is a convolution operation used in a three-dimensional convolutional neural network. Its main purpose is to reduce the model complexity, thereby improving the calculation efficiency and the model inference speed. Compared with the traditional three-dimensional convolution operation, the three-dimensional lightweight convolution module adopts a channel attention mechanism, which can effectively reduce the number of parameters and the amount of calculation and maintain the accuracy of the model.
[0067] The features of the T-frame pictures are used as the input of the three-dimensional hierarchical attention module. The three-dimensional hierarchical attention can decompose the global self-attention with three-dimensional complexity into multi-layer attention with lower computational cost. A hierarchical module is introduced in the three-dimensional hierarchical attention to obtain the three-dimensional hierarchical attention module, which is used to summarize the spatio-temporal local features and spatio-temporal global features and enrich the features.
[0068] The specific steps for obtaining the hierarchical features through the three-dimensional hierarchical attention module are as follows: First, a pooling operation is performed on the features of the T-frame pictures to obtain hierarchical features, specifically as follows:
[0069]
[0070] where represents the pooling operation. The pooling operation is a commonly used operation in convolutional neural networks. Its main function is to downsample or reduce the dimension of the input feature map, reduce the model calculation amount and memory occupancy, and can extract the main information in the input feature map. represents the hierarchical features, represents the features of the input T-frame pictures.
[0071] Then, an attention operation is performed on the hierarchical features, specifically as follows:
[0072]
[0073]
[0074] Among them, represents layer normalization, which plays a role in normalizing the input features in a neural network, can accelerate the convergence of the network, improve the generalization ability and enhance the robustness of the network, thereby improving the performance and effect of the model. represents 3D multi-head self-attention, which can effectively model spatial relationships, temporal relationships, and relationships in point cloud data when processing 3D data. and are learnable cross-channel scale multipliers, which are used for 3D multi-head self-attention and feed-forward network respectively, and their function is to enhance the feature expression ability and reduce the computational amount to obtain better 3D lightweight convolution results. represents the feed-forward network.
[0075] Merge the local features , global features and hierarchical features of the features of the T-frame pictures, and perform an attention operation, specifically as follows:
[0076]
[0077]
[0078]
[0079] Among them, the first represents the result after merging the local features , global features and hierarchical features , represents layer normalization, represents 3D multi-head self-attention, and are learnable cross-channel scale multipliers, represents the feed-forward network, (·) represents the merging operation.
[0080] Finally, hierarchical features are output through the 3D hierarchical attention module, specifically:
[0081]
[0082] Among them, represents the hierarchical features output by the 3D hierarchical attention module, (·) represents the upsampling operation.
[0083] Fusing the hierarchical features based on the three-dimensional rotator module to obtain the fused features includes:
[0084] Figure 3 This is a schematic diagram of the three-dimensional rotator module structure for feature fusion in the intelligent alarm method based on environmental perception provided by an embodiment of the present application. As Figure 3 shown, the three-dimensional rotator module includes a first block and a second block, which are used together. Both the first block and the second block first go through a layer normalization, then a three-dimensional multi-head self-attention, then a layer normalization, then a three-dimensional lightweight depth convolution, and finally a feed-forward neural network, where the feed-forward neural network consists of a two-layer multi-layer perceptron and the activation function GELU;
[0085] The difference between the first block and the second block is that the first block is a three-dimensional sliding window multi-head self-attention for local feature information exchange within the window, and the second block is a three-dimensional shifted sliding window multi-head self-attention for global feature information exchange between windows.
[0086] Input the hierarchical features into the first block of the three-dimensional rotator module. The features obtained through layer normalization and three-dimensional sliding window multi-head self-attention are represented by , and the features obtained through the feed-forward neural network are represented by Specifically, as follows:
[0087]
[0088]
[0089] Among them, represents the three-dimensional sliding window multi-head self-attention. This attention is a mechanism for calculating the interaction relationship between sequence data. It divides the input sequence into blocks, and then calculates the attention weights within each block to obtain the correlation between each element and other elements, thereby realizing the modeling of sequence data; the three-dimensional sliding window attention can process longer input sequences, and it does not need to calculate the attention for the entire sequence, only needs to calculate the time steps within each block. Since the number of time steps within each block is the same, efficient matrix multiplication can be used for calculation, thereby improving the running efficiency. represents layer normalization, represents the feed-forward neural network, Represents a three-dimensional depth lightweight convolution, which is a combination based on depthwise separable convolution and per-channel convolution. Depthwise separable convolution is a method of splitting a standard convolution into a depth convolution and a pointwise convolution, which can reduce the number of model parameters and computational complexity. Per-channel convolution is a convolution operation performed on the input signal in the channel dimension, which can more effectively retain the local correlation of the signal.
[0090] The output of the first block in the three-dimensional rotator module As the input of the second block, the features obtained through layer normalization and three-dimensional shifted sliding window multi-head self-attention are denoted by The features obtained through the feed-forward neural network are denoted by Specifically as follows:
[0091]
[0092]
[0093] Among them, Denotes three-dimensional shifted sliding window multi-head self-attention. Compared with the traditional three-dimensional sliding window attention, this attention can better capture the temporal and spatial relationships in the sequence data. In the traditional three-dimensional sliding window attention, the attention weights between time steps within each block are obtained through a convolution operation in three-dimensional space, while the three-dimensional shifted sliding window attention adopts the idea similar to dilated convolution and adds a translation operation within the convolution kernel to achieve interaction between different positions. Specifically, the three-dimensional shifted sliding window attention divides the input sequence into multiple blocks of the same size, and each block can be regarded as a matrix, where the rows represent time steps and the columns represent features; within each block, for each time step, the attention weights between this time step and other time steps are calculated; the role of the three-dimensional shifted sliding window attention is to better capture the temporal and spatial relationships in the sequence data, thereby improving the modeling ability of the sequence data. Output the final fused features .
[0094] Input the fused features into the behavior classification network for recognition and classification to obtain the behavior classification result, and input the behavior classification result into the warning network. The behavior classification result is a specific numerical value, ranging from 1.00 to 10.00; set the threshold to 6.00. When the behavior classification result is greater than or equal to the threshold 6.00, it is determined as a fast running behavior, that is, the warning is triggered. If the behavior classification result is less than the threshold 6.00, return to obtain the next time series video.
[0095] Furthermore, the hierarchical features first pass through the first block and then through the second block to obtain the fused features.
[0096] Specifically, the use of three-dimensional sliding window multi-head self-attention in the first block specifically includes:
[0097] Taking the hierarchical features output by the action recognition feature extraction network as input data;
[0098] Performing a linear transformation on the normalized data, and projecting it into the query, key, and value spaces respectively to obtain the query space mapping value, key space mapping value, and value space mapping value;
[0099] Dividing the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum;
[0100] After completing the attention calculation and weighted sum operations of all heads, concatenate the output results of multiple heads along the channel dimension to obtain the output of the three-dimensional sliding window multi-head self-attention.
[0101] Specifically, the use of three-dimensional shifted sliding window multi-head self-attention in the second block specifically includes:
[0102] Taking the output of the first block after normalization as the input data of the second block;
[0103] Performing a linear transformation on the normalized data, and projecting it into the query, key, and value spaces respectively to obtain the query space mapping value, key space mapping value, and value space mapping value;
[0104] Dividing the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum;
[0105] After completing the attention calculation and weighted sum operations of all heads, concatenate the output results of multiple heads along the channel dimension to obtain the output of the three-dimensional shifted sliding window multi-head self-attention.
[0106] Specifically, the data input and layer normalization steps in the first block are as follows: Taking the hierarchical features output by the action recognition feature extraction network as input data, denoted as Z b . First, perform layer normalization on the input data Z b . Perform layer normalization on all features of each sample, and the formula is:
[0107]
[0108] where represents the layer normalization result, μ is the mean of Z b on the feature dimension, is the variance, is a very small value that can prevent the denominator from being zero, and are learnable parameters. Through layer normalization in this application, the input data has better distribution characteristics when entering subsequent operations, accelerating network convergence and improving model performance.
[0109] 3D Sliding Window Multi-Head Self-Attention Calculation: The data after layer normalization enters the 3D Sliding Window Multi-Head Self-Attention (3DW-MSA) module. In this module, the input data is first projected into three different spaces of Query, Key, and Value respectively to obtain Q, K, and V. Assuming the dimension of the input data is [T, H, W, C] (T is the time dimension, H is the height, W is the width, and C is the number of channels), then:
[0110]
[0111]
[0112]
[0113] Among them, , , are learnable weight matrices. Then, Q, K, and V are divided into multiple heads, and each head calculates the attention score and weighted sum respectively.
[0114] Taking one head as an example, the calculation method of the attention score is:
[0115]
[0116] Where is the dimension of the key. The weighted sum is O: . Finally, the outputs of multiple heads are concatenated to obtain the output of the 3D sliding window multi-head self-attention .
[0117] Repeated Layer Normalization: The output of the 3D sliding window multi-head self-attention is subjected to repeated layer normalization operation again. The calculation method is the same as that of the layer normalization in the first step, and this application will not elaborate on it.
[0118] 3D Depth Lightweight Depth Convolution: 3D Depth Lightweight Convolution (3DDWConv) is used to perform convolution operation on the data after repeated layer normalization. 3D lightweight depth convolution combines the advantages of depthwise separable convolution and per-channel convolution. First, depth convolution is performed, and convolution operations are performed on each channel respectively, and the convolution kernel size is [k T , k H , k W . Assuming the input data is X, the output after depth convolution is:
[0119]
[0120] Among them, is the convolution kernel of depth convolution, , , are the coordinate indices of the feature map in different dimensions respectively, is the channel index. Then, channel-wise convolution is performed, and convolution operation is carried out on the output of depth convolution in the channel dimension to obtain the final convolution result .
[0121] Feedforward neural network processing: The output of the three-dimensional lightweight depth convolution is input into a feedforward neural network (FFN). The feedforward neural network consists of a two-layer multi-layer perceptron (MLP) and an activation function GELU. Let the weight matrix of the first layer of the MLP be W1, the bias be b1, the weight matrix of the second layer be W2, and the bias be b2. Then the calculation process of the FFN is as follows:
[0122]
[0123] After being processed by the feedforward neural network, the final output of the first block is obtained, and this output will be used as the input of the second block.
[0124] Implementation steps of the second block:
[0125] Data input and layer normalization: The output of the first block is used as the input data of the second block. Similarly, first perform layer normalization operation on it, and the calculation method is the same as that of the layer normalization in the first block to stabilize the data distribution.
[0126] Three-dimensional shifted sliding window multi-head self-attention calculation: The data after layer normalization enters the three-dimensional shifted sliding window multi-head self-attention (3DWS-MSA) module. This module is similar to the three-dimensional sliding window multi-head self-attention module, but a shifting operation is introduced in window partitioning. First, the input data is projected into the query, key, and value spaces in the same way to obtain Q, K, and V. Then, when partitioning the window, each window has a certain shift (such as a shift step of 1) in the time, height, and width dimensions relative to the previous window. Taking one head as an example, the method of calculating the attention score and weighted sum is the same as that of the three-dimensional sliding window multi-head self-attention module, which will not be elaborated in this application. The outputs of multiple heads are concatenated to obtain the output of the three-dimensional shifted sliding window multi-head self-attention.
[0127] Layer normalization again: Perform layer normalization on the output of the three-dimensional shifted sliding window multi-head self-attention to ensure the stability of the data, and the calculation method is the same as the previous layer normalization step.
[0128] 3D lightweight depth convolution: Consistent with the 3D lightweight depth convolution operation in the first block, perform 3D lightweight depth convolution on the data that has undergone layer normalization again. First, perform depth convolution, and then perform channel-wise convolution to obtain the convolution result.
[0129] Feed-forward neural network processing: Input the output of the 3D lightweight depth convolution into the feed-forward neural network. The structure and calculation method of the feed-forward neural network are the same as those in the first block. After being processed by the feed-forward neural network, the final output of the second block, that is, the fused feature, is obtained. .
[0130] It should be noted that the traditional 3D transformer model has problems of large computational complexity and poor effect of fusing spatio-temporal features. The present application innovatively introduces a lightweight convolution structure into the 3D rotation transformer module, especially adopts a grouped convolution optimization method in the 3D lightweight depth convolution link.
[0131] Exemplarily, in traditional 3D convolution, for the input feature map , where T is the time dimension, H is the height, W is the width, and C is the number of input channels) and the convolution kernel , where is the number of output channels, the convolution calculation is Y = W X, and the computational complexity is ). While grouped convolution divides the number of input channels C into G groups, with the number of input channels in each group . For the convolution kernel of the g-th group, represents the number of output channels in each group, and the output feature map Y g of the g-th group is calculated as (X g is the g-th group data after the input feature map X is divided by group). The final output feature map Y is composed of the concatenation of the outputs of each group. At this time, the computational complexity becomes ). It can be seen that the present application significantly reduces the computational complexity without losing too much accuracy. The significant reduction in computational complexity directly improves the running efficiency of the model. In the real-time behavior recognition and warning scenario, the system can process video data faster, shortening the time interval from video acquisition to behavior judgment and warning. Since grouped convolution can maintain a certain accuracy while reducing the computational complexity, the feature fusion effect is not significantly affected. The improved module structure can better fuse spatio-temporal features, accurately capture the features of a person running fast, and improve the accuracy of behavior classification and warning.
[0132] In addition, in the embodiments of the present application, to further improve the attention mechanism performance of the 3D rotation transformer module, a dynamic window adjustment mechanism is introduced.
[0133] Specifically, the three-dimensional rotator module fuses the hierarchical features to obtain fused features, and further includes:
[0134] Define two parameters: window size adjustment factor and step size adjustment factor;
[0135] Calculate the displacement of pixels between adjacent frames by the optical flow method, and then obtain the average motion speed;
[0136] Measure the feature complexity by calculating the entropy of the feature map;
[0137] Adjust the window size and step size of the three-dimensional sliding window and the three-dimensional shifted sliding window in the three-dimensional rotator module according to the average motion speed and the feature complexity.
[0138] Specifically, first, calculate the displacement of pixels between adjacent frames by the optical flow method to obtain the average motion speed v. Assume that between the t-th frame and the (t + 1)-th frame, the total pixel displacement is S and the total number of pixels is N, then v = S / N. Evaluate the feature complexity by calculating the entropy of the feature map. p(k) represents the probability that the pixel value in the feature map is k, then the entropy .
[0139] After that, the window size and step size of the sliding window can be dynamically adjusted according to the motion speed and the feature complexity. When it is detected that the motion speed of the person in the video is fast and the action changes frequently (the judgment basis is that v is greater than the set threshold), increase the window step size to quickly capture key features; when the motion of the person in the scene is relatively slow and the feature details are rich, reduce the window size to focus on local detailed features.
[0140] In this application, by introducing a dynamic window adjustment mechanism, the attention mechanism is made more flexible and efficient. When processing video data in different scenarios, it can adaptively adjust the window parameters according to the actual situation to better capture different behavior features. When the person's motion speed is fast, increasing the step size can quickly scan key actions and avoid missing important information; when the motion is slow, reducing the window can more carefully analyze local features, improving the ability to capture behavior features in complex scenarios, thereby enhancing the accuracy of behavior classification and early warning, reducing the occurrence of misjudgment and missed judgment, and enhancing the reliability and practicality of the entire behavior recognition and early warning system.
[0141] Further, in other embodiments of this application, the three-dimensional rotator module fuses the hierarchical features to obtain fused features, and further includes:
[0142] Fuse the input of the first block and the output of the first block by addition fusion or splicing fusion to obtain fused input data;
[0143] Replace the input of the second block with the fused input data.
[0144] In the three-dimensional rotation transformer module, a skip connection is constructed based on the first block and the second block. The input of the first block is Z input , after passing through layer normalization, three-dimensional sliding window multi-head self-attention, layer normalization again, three-dimensional lightweight depth convolution, and a feed-forward neural network in the first block, the output is Z first-block .
[0145] To implement the skip connection, Z input is directly fused with Z first-block . The fusion method is either addition or concatenation. If addition fusion is used, the input of the second block is:
[0146]
[0147] If concatenation fusion is used, assuming the dimensions of Z input and Z first-block are [T, H, W, C1] and [T, H, W, C2] respectively, then the dimension after concatenation is [T, H, W, C1 + C2], that is, the input of the second block is:
[0148]
[0149] where Concat(·) represents the concatenation function.
[0150] During the processing of the second block, whether the result of addition or concatenation fusion is used as the input, the features of the early layer (the input of the first block) can directly participate in the processing of the second block. During backpropagation, the gradient can flow more smoothly from the deep layer back to the shallow layer through the skip connection, alleviating the problem of gradient disappearance, while retaining the key information in the original features, which helps the model learn richer spatio-temporal features.
[0151] Furthermore, in a preferred embodiment of the present application, a branch fusion scheme is also provided. Specifically, the three-dimensional rotation transformer module fuses hierarchical features to obtain fused features, and further includes:
[0152] Performing convolution processing on the spatial dimension of the input data of the first block using a two-dimensional convolution to extract spatial local feature data;
[0153] Replacing the three-dimensional sliding window multi-head self-attention with a two-dimensional sliding window multi-head self-attention, and replacing the three-dimensional shifted sliding window multi-head self-attention with a one-dimensional shifted sliding window multi-head self-attention;
[0154] Regarding the spatial local feature data as the input data of the first block again;
[0155] Performing a time dimension convolution operation on the output of the first block to extract time series features;
[0156] Fuse the output of the first block and the output of the second block to obtain the fused features.
[0157] The following provides a detailed description of the branch fusion.
[0158] (I) Spatial Feature Branch
[0159] Data Input and Preprocessing: Input the input Z of the three-dimensional rotation transformer module input into the spatial feature branch. First, perform a preprocessing operation specifically for spatial feature extraction. Use a two-dimensional convolution to perform convolution processing on the spatial dimensions (height and width) to extract spatial local features. Assume that the two-dimensional convolution kernel used is W 2D-spatial , then the output after convolution is:
[0160]
[0161] where k H and k W are the sizes of the two-dimensional convolution kernel in the height and width directions. The two colons in mean taking all elements in the corresponding dimension. The first two colons in mean taking all elements in the corresponding dimension. The colon in and means starting from the and positions and selecting all subsequent elements in the corresponding dimension.
[0162] Spatial Feature Extraction and Processing: Z spatial-pre undergoes layer normalization to stabilize the data distribution. Then it enters the spatial attention module, which can be based on a two-dimensional attention mechanism, such as two-dimensional multi-head self-attention (2DMHSA). Calculate the attention scores and weighted sums to obtain an output that is more focused on spatial features. After that, it passes through a lightweight three-dimensional convolution to further fuse spatial features and a small amount of temporal dimension information (here the size of the temporal dimension convolution kernel can be set to 1, that is, only perform convolution operations in the spatial dimension) to obtain the final output Z spatial-output of the spatial feature branch.
[0163] (II) Temporal Feature Branch
[0164] Data Input and Preprocessing: Similarly, input Z input into the temporal feature branch. First, perform preprocessing for temporal feature extraction. Use a one-dimensional convolution to perform convolution operations on the temporal dimension (T) to extract time series features. Assume that the one-dimensional convolution kernel is , then the output after convolution is:
[0165]
[0166] where k T is the size of the one-dimensional convolutional kernel in the time dimension.
[0167] Time feature extraction and processing: After layer normalization of the output result of convolution, it enters the time attention module, such as one-dimensional multi-head self-attention (1DMHSA), calculates the attention scores and weighted sums, and obtains the output focusing on time features. Then, through a lightweight three-dimensional convolution (the spatial dimension convolutional kernel size is set to 1, mainly performing convolution operations in the time dimension), the time features and a small amount of spatial dimension information are further fused to obtain the final output Z of the time feature branch temporal-output .
[0168] (III) Branch fusion
[0169] Fuse the output Z spatial-output of the spatial feature branch and the output Z temporal-output of the time feature branch. Using the weighted fusion method, let the weight of the spatial feature branch be α and the weight of the time feature branch be 1 - α, then the fused output is:
[0170]
[0171] The weight α can be a fixed value set according to experimental experience, or it can be a learnable parameter, and the model is trained to automatically adjust it to adapt to different behavior data and scenario requirements. In this way, by respectively focusing on processing spatial and time features through the dual-branch parallel structure and then fusing the results, the comprehensive processing ability of the model for spatio-temporal features can be improved, and it can better adapt to complex behavior patterns.
[0172] Embodiment 2
[0173] Figure 4 is an intelligent alarm system based on environmental perception provided by an embodiment of the present application. The system includes:
[0174] A video acquisition and initialization module, configured to acquire the video captured by the environmental perception camera, initialize the video, and divide the video into T frame pictures;
[0175] A behavior recognition feature extraction module, configured to:
[0176] Extract hierarchical features from the T frame pictures based on the behavior recognition feature extraction network;
[0177] A three-dimensional rotator feature fusion module, configured to:
[0178] Fuse the hierarchical features based on the three-dimensional rotator module to obtain fused features;
[0179] A classification module, configured to:
[0180] Input the fused features into the classification module for recognition and classification to obtain the behavior classification result, and input the behavior classification result into the warning network;
[0181] The warning module is used for:
[0182] Set a threshold value in the warning network. When the behavior classification result is greater than or equal to the set threshold value, it is determined as a fast running behavior, that is, the warning is triggered. When the behavior classification result is less than the threshold value, return to obtain the video of the next time period.
[0183] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0184] The above-described embodiments only express several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
[0185] The above order of the steps for the method is only for illustration. The steps of the method of this application are not limited to the above specifically described order, unless otherwise specifically stated. In addition, in some embodiments, this application can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the method according to this application. Therefore, this application also covers a recording medium storing a program for executing the method according to this application.
[0186] As described above in the specific implementation manners, the purpose, technical solution and beneficial effects of the present invention are further described in detail. It should be understood that the above is only the specific implementation manner of the present invention and is not used to limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An intelligent alarm method based on environmental perception, characterized in that, Including: Obtain the video captured by the environmental perception camera, initialize the video, and divide the video into T frame pictures; Extract hierarchical features from the T frame pictures based on the behavior recognition feature extraction network, where the behavior recognition feature extraction network consists of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module; Fuse the hierarchical features based on the three-dimensional rotary transformer module to obtain fused features; Input the fused features into the behavior classification network for behavior recognition and classification to obtain a behavior classification result, and input the behavior classification result into the warning network; Set a threshold in the warning network. When the behavior classification result is greater than or equal to the threshold, it is determined as a fast running behavior, that is, a warning is triggered. When the behavior classification result is less than the threshold, return to obtain the next time series video; Among them, the three-dimensional rotary transformer module includes a first block and a second block. The step of fusing the hierarchical features based on the three-dimensional rotary transformer module to obtain fused features includes: The hierarchical features first pass through the first block and then through the second block to obtain fused features; both the first block and the second block first pass through a layer normalization, then through a three-dimensional multi-head self-attention, then through a layer normalization, then through a three-dimensional lightweight depth convolution, and finally through a feed-forward neural network; the difference between the first block and the second block is that the first block uses three-dimensional sliding window multi-head self-attention, and the second block uses three-dimensional shifted sliding window multi-head self-attention.
2. The intelligent alarm method based on environmental perception according to claim 1, wherein The step of obtaining the video captured by the environmental perception camera, initializing the video, and dividing the video into T frame pictures includes: Collect the human behavior video in the dynamic scene through the camera and initialize it; Divide the video into T frame pictures. The size of the T frame pictures is T×H×W×3, where T represents the time dimension, H represents the height, W represents the width, and 3 represents the number of color channels.
3. The intelligent alarm method based on environmental perception according to claim 1, characterized in that, The step of extracting hierarchical features from the T frame pictures based on the behavior recognition feature extraction network includes: Input the T frame pictures into the three-dimensional lightweight convolution module in the behavior recognition feature extraction network in sequence; Use The features of T-frame pictures are obtained through convolution calculation , where the features of the T-frame pictures include local features and global features ; Input the features of the T frame pictures into the three-dimensional hierarchical attention module for extracting hierarchical features to obtain hierarchical features.
4. The intelligent alarm method based on environmental perception according to claim 1, characterized in that The specific method of the first block using three-dimensional sliding window multi-head self-attention includes: Use the hierarchical features output by the behavior recognition feature extraction network as input data; Perform a linear transformation on the normalized data, project it into the query, key, and value spaces respectively to obtain the query space mapping value, key space mapping value, and value space mapping value; Divide the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum; After completing the attention calculation and weighted sum operations of all heads, splice the output results of multiple heads along the channel dimension to obtain the output of the three-dimensional sliding window multi-head self-attention.
5. The intelligent alarm method based on environmental perception according to claim 1, characterized in that The specific method of the second block using three-dimensional shifted sliding window multi-head self-attention includes: Use the output of the first block after normalization processing as the input data of the second block; Perform a linear transformation on the normalized data, project it into the query, key, and value spaces respectively, and obtain the query space mapping value, key space mapping value, and value space mapping value; Divide the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum; After completing the attention calculation and weighted sum operations of all heads, splice the output results of multiple heads along the channel dimension to obtain the output of the three-dimensional shifted sliding window multi-head self-attention.
6. The intelligent alarm method based on environmental perception according to claim 1, wherein The method of fusing hierarchical features based on the three-dimensional transformer module further includes: Define two parameters: window size adjustment factor and stride adjustment factor; Calculate the displacement of pixels between adjacent frames by the optical flow method, and then obtain the average motion speed; Measure the feature complexity by calculating the entropy of the feature map; Adjust the window size and stride of the three-dimensional sliding window and three-dimensional shifted sliding window in the three-dimensional transformer module according to the average motion speed and feature complexity.
7. The intelligent alarm method based on environmental perception according to claim 1, wherein The method of fusing hierarchical features based on the three-dimensional transformer module further includes: Fuse the input of the first block and the output of the first block by addition fusion or splicing fusion to obtain the fused input data; Replace the input of the second block with the fused input data.
8. The intelligent alarm method based on environmental perception according to claim 7, characterized in that The method of fusing hierarchical features based on the three-dimensional transformer module further includes: Perform convolution processing on the spatial dimension of the input data of the first block by two-dimensional convolution to extract spatial local feature data; Replace the three-dimensional sliding window multi-head self-attention with two-dimensional sliding window multi-head self-attention, and replace the three-dimensional shifted sliding window multi-head self-attention with one-dimensional shifted sliding window multi-head self-attention; Use the spatial local feature data as the input data of the first block again; Perform a time dimension convolution operation on the output of the first block to extract time series features; Fuse the output of the first block and the output of the second block to obtain the fused feature.
9. An intelligent alarm system based on environmental perception, characterized in that The system includes: A video acquisition and initialization module, configured to acquire a video captured by an environmental perception camera, initialize the video, and divide the video into T frame pictures; A behavior recognition feature extraction module, configured to extract hierarchical features from the T frame pictures based on a behavior recognition feature extraction network, wherein the behavior recognition feature extraction network is composed of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module; A three-dimensional transformer feature fusion module, configured to fuse hierarchical features based on a three-dimensional transformer module to obtain a fused feature; A classification module, configured to input the fused feature into a behavior classification network for behavior recognition and classification to obtain a behavior classification result, and input the behavior classification result into a warning network; A warning module, configured to set a threshold in the warning network. When the behavior classification result is greater than or equal to the threshold, it is determined as a fast running behavior, that is, a warning is triggered. When the behavior classification result is less than the threshold, it returns to acquire the next time series video; Wherein, the three-dimensional transformer module includes a first block and a second block, and the method of fusing hierarchical features based on the three-dimensional transformer module includes: The hierarchical features first pass through the first block and then through the second block to obtain fused features; both the first block and the second block first go through a layer normalization, then through a 3D multi-head self-attention, then through another layer normalization, then through a 3D lightweight depth convolution, and finally through a feed-forward neural network; the difference between the first block and the second block is that the first block uses 3D sliding window multi-head self-attention, and the second block uses 3D shifted sliding window multi-head self-attention.
Citation Information
Patent Citations
Signal searching method based on wavelet convolution and Former multi-scale feature fusion
CN117830787A
Mobile terminal alarm device and method based on artificial intelligence
CN119155378A