Intelligent alarm method and system based on environment perception

By introducing three-dimensional lightweight convolution module, three-dimensional hierarchical attention module and three-dimensional rotation transformer module into behavior recognition warning technology, the problem of large amount of feature extraction and fusion calculation in the existing technology is solved, and more efficient behavior recognition and fast running behavior warning are achieved.

CN119942654AActive Publication Date: 2025-05-06XIAN XUYANG COMM EQUIP CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510429252.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

In the existing behavior recognition and warning technology, the feature extraction network has a large amount of computation, making it difficult to capture global context information, and the feature fusion network has a large amount of computation, making it difficult to integrate global feature information and local feature information in time and space, affecting real-time and accuracy.

Method used

An intelligent alarm method based on environmental perception is proposed, using a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module for feature extraction, and combining a three-dimensional rotation transformer module for feature fusion to improve the real-time and accuracy of behavior recognition.

Benefits of technology

By optimizing the feature extraction and fusion process, the computing volume and memory usage are reduced, the real-time and accuracy of behavior recognition warnings are improved, and the video data can be processed faster and the characteristics of the human body's fast running behavior can be accurately captured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942654A_ABST
    Figure CN119942654A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent alarm method and system based on environmental perception, and relates to the technical field of artificial intelligence application, and the method comprises the steps: obtaining a video of a dynamic scene through a camera, dividing the video into T frames of pictures, and inputting the T frames of pictures into a behavior recognition feature extraction network; the behavior recognition feature extraction network is composed of a three-dimensional lightweight convolution module and a three-dimensional layered attention module; the three-dimensional lightweight convolution module is used for outputting features of T frames of pictures, and outputting hierarchical features of the features of the T frames of pictures through the three-dimensional hierarchical attention module; the three-dimensional rotary converter module fuses the layered features to obtain fused features; a behavior classification result is obtained according to the fusion features through a behavior classification network, and an early warning network achieves rapid running behavior recognition early warning for dynamic scene perception according to the classification result. The invention aims to improve the feature expression capability and the feature fusion capability of behavior recognition early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence application technology, and in particular to an intelligent alarm method and system based on environmental perception. Background Art

[0002] Environmental perception refers to the perception and understanding of dynamic information in the surrounding environment through sensors and algorithms. Behavior recognition refers to the process of inferring the intention, state or characteristics of humans or other entities by analyzing and understanding their actions, behaviors and related environmental information. Behavior recognition warning refers to the early detection of potential risks, abnormalities or dangerous situations in advance and timely warning by the early warning system based on behavior recognition by judging and analyzing certain specific behaviors or actions. Behavior recognition warning technology mainly relies on algorithms such as machine learning, deep learning and pattern recognition to analyze and judge the collected data, and can identify abnormal behaviors such as fast running.

[0003] The key issue of behavior recognition is how to extract local and global features of space and time at the same time and how to fuse features. In recent years, researchers have proposed the use of 3D convolution and 3D transformer to extract and fuse local and global features of space and time. Convolutional networks can effectively extract local information from relevant feature maps through neighborhood convolution operations. However, the limited receptive field of convolutional networks makes it difficult to capture global contextual information. The weighted average operation of the contextual information of the input features based on the attention transformer can effectively capture global information, but blind similarity matching of adjacent image blocks may lead to a high mismatch rate.

[0004] Behavior recognition warning requires real-time and accuracy. The three-dimensional convolutional network and three-dimensional attention network generally used for feature extraction have large computational complexity and high memory usage, which will affect the real-time performance of behavior recognition warning. The three-dimensional transformer generally used for feature fusion fuses local and global features of space and time. However, the three-dimensional transformer model is large, computationally intensive, and cannot well fuse the global and local feature information of space and time, which will reduce the accuracy of behavior recognition warning. Summary of the invention

[0005] The present application aims to solve at least one of the technical problems in the related art to a certain extent. To this end, one purpose of the present application is to propose an intelligent alarm method and system based on environmental perception to improve the real-time and accuracy of behavior recognition warning.

[0006] The first aspect disclosed in the present application provides an intelligent alarm method based on environmental perception, and the specific steps include:

[0007] Get the video captured by the environment perception camera, initialize the video, and divide the video into T frames;

[0008] Extracting features from the T-frame images based on a behavior recognition feature extraction network to obtain hierarchical features, wherein the behavior recognition feature extraction network is composed of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module;

[0009] The hierarchical features are fused based on the three-dimensional rotation transformer module to obtain fused features;

[0010] Inputting the fusion features into a behavior classification network to perform behavior recognition and classification, obtaining a behavior classification result, and inputting the behavior classification result into an early warning network;

[0011] A threshold is set in the early warning network. When the behavior classification result is greater than or equal to the threshold, it is determined to be a fast running behavior, and an early warning is triggered. When the behavior classification result is less than the threshold, it returns to obtain the next time series video.

[0012] The step of obtaining a video captured by an environment perception camera, initializing the video, and dividing the video into T frames of pictures includes:

[0013] Use the camera to collect and initialize human behavior videos in dynamic scenes;

[0014] The video is divided into T frame pictures, and the size of the T frame pictures is T×H×W×3, where T represents the time dimension, H represents the height, W represents the width, and 3 represents the number of color channels.

[0015] The step of extracting features from the T-frame images based on the behavior recognition feature extraction network to obtain hierarchical features includes:

[0016] Inputting the T frame images sequentially into the three-dimensional lightweight convolution module in the behavior recognition feature extraction network;

[0017] use Convolution calculation obtains the features of T frame images , wherein the characteristics of the T frame picture Including local features and global features ;

[0018] The features of the T frame image are input into a three-dimensional hierarchical attention module to extract hierarchical features to obtain hierarchical features.

[0019] The three-dimensional rotation transformer module includes a first block and a second block, and the fusion feature obtained by fusing the hierarchical features based on the three-dimensional rotation transformer module includes:

[0020] The hierarchical features first pass through the first block and then pass through the second block to obtain the fused features; both the first block and the second block first pass through a layer normalization, then a three-dimensional multi-head self-attention, then a layer normalization, then a three-dimensional lightweight deep convolution, and finally a feedforward neural network; the difference between the first block and the second block is that the first block uses a three-dimensional sliding window multi-head self-attention, and the second block uses a three-dimensional shifted sliding window multi-head self-attention.

[0021] The first block uses a three-dimensional sliding window multi-head self-attention and specifically includes:

[0022] The hierarchical features output by the behavior recognition feature extraction network are used as input data;

[0023] Perform linear transformation on the normalized data and project them into query, key and value spaces respectively to obtain query space mapping value, key space mapping value and value space mapping value;

[0024] Divide the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum;

[0025] After completing the attention calculation and weighted sum operation of all heads, the output results of multiple heads are spliced ​​together according to the channel dimension to obtain the output of the three-dimensional sliding window multi-head self-attention.

[0026] The second block uses a three-dimensional shift sliding window multi-head self-attention and specifically includes:

[0027] Normalize the output of the first block and use it as the input data of the second block;

[0028] Perform linear transformation on the normalized data and project them into query, key and value spaces respectively to obtain query space mapping value, key space mapping value and value space mapping value;

[0029] Divide the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum;

[0030] After completing the attention calculation and weighted sum operation of all heads, the output results of multiple heads are spliced ​​together according to the channel dimension to obtain the output of the three-dimensional shifted sliding window multi-head self-attention.

[0031] The step of fusing the hierarchical features based on the three-dimensional rotation transformer module to obtain the fused features further includes:

[0032] Define two parameters: window size adjustment factor and step size adjustment factor;

[0033] The displacement of pixels between adjacent frames is calculated by the optical flow method, and then the average motion speed is obtained;

[0034] The feature complexity is measured by calculating the entropy of the feature map;

[0035] The window size and step size of the 3D sliding window and the 3D shift sliding window in the 3D rotation transformer module are adjusted according to the average motion speed and feature complexity.

[0036] The step of fusing the hierarchical features based on the three-dimensional rotation transformer module to obtain the fused features further includes:

[0037] The input of the first block and the output of the first block are added or concatenated to obtain fused input data;

[0038] The input of the second block is replaced by the fused input data.

[0039] The step of fusing the hierarchical features based on the three-dimensional rotation transformer module to obtain the fused features further includes:

[0040] Use two-dimensional convolution to perform convolution processing on the spatial dimension of the input data of the first block to extract spatial local feature data;

[0041] Use two-dimensional sliding window multi-head self-attention to replace three-dimensional sliding window multi-head self-attention, and use one-dimensional shift sliding window multi-head self-attention to replace three-dimensional shift sliding window multi-head self-attention;

[0042] Re-using the spatial local feature data as input data of the first block;

[0043] Perform a time dimension convolution operation on the output of the first block to extract time series features;

[0044] The output of the first block is fused with the output of the second block to obtain the fused feature.

[0045] The second aspect disclosed in the present application provides an intelligent alarm system based on environmental perception, the system comprising:

[0046] The video acquisition and initialization module is used to acquire the video captured by the environment perception camera, initialize the video, and divide the video into T frame images;

[0047] A behavior recognition feature extraction module, used for extracting features from the T-frame images based on a behavior recognition feature extraction network to obtain hierarchical features, wherein the behavior recognition feature extraction network is composed of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module;

[0048] A three-dimensional rotation transformer feature fusion module, used for fusing hierarchical features based on the three-dimensional rotation transformer module to obtain fused features;

[0049] A classification module, used for inputting the fusion features into the classification module for identification and classification, obtaining a behavior classification result, and inputting the behavior classification result into the early warning network;

[0050] The early warning module is used to set a threshold in the early warning network. When the behavior classification result is greater than or equal to the threshold, it is determined to be a fast running behavior, that is, triggering an early warning. When the behavior classification result is less than the threshold, it returns to obtain the next time series video.

[0051] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0052] The present application provides an intelligent alarm method and system based on environmental perception, which solves the problems in existing behavior recognition and warning, such as large computational complexity of the behavior feature extraction network, inability to accurately capture local feature information, and inability of the feature fusion network to better integrate the global feature information and local feature information of time and space. It efficiently realizes the recognition of human behavior in the video and issues warnings for fast running behavior to prevent danger. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 An embodiment of the present application provides a flow chart of an intelligent alarm method based on environment perception;

[0054] Figure 2 An embodiment of the present application provides a structural diagram of a behavior recognition feature extraction network framework in an intelligent alarm method based on environment perception;

[0055] Figure 3 A schematic diagram of the structure of a three-dimensional rotation transformer module for feature fusion in an intelligent alarm method based on environment perception is provided for one embodiment of the present application;

[0056] Figure 4 An embodiment of the present application provides a schematic diagram of an intelligent alarm system based on environmental perception. DETAILED DESCRIPTION

[0057] In order to better understand the present application, a more detailed description will be made of various aspects of the present application with reference to the accompanying drawings. It should be understood that these detailed descriptions are only descriptions of exemplary embodiments of the present application, and are not intended to limit the scope of the present application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.

[0058] In the accompanying drawings, the size, dimensions and shapes of elements have been slightly adjusted for ease of illustration. The drawings are for illustration only and are not drawn strictly to scale. As used herein, the terms "substantially," "approximately," and similar terms are used as terms of approximation, not degree, and are intended to account for the inherent deviations in measurements or calculations that would be recognized by one of ordinary skill in the art.

[0059] It should also be understood that expressions such as "including", "comprising", "having", "containing" and / or "comprising" are open rather than closed expressions in this specification, which indicate the presence of the stated features, elements and / or components, but do not exclude the presence of one or more other features, elements, components and / or combinations thereof. In addition, when expressions such as "at least one of..." appear after a list of listed features, they modify the entire list of features rather than just the individual elements in the list. In addition, when describing embodiments of the present application, "may" is used to mean "one or more embodiments of the present application". And, the term "exemplary" is intended to refer to an example or illustration.

[0060] Unless otherwise defined, all words (including engineering terms and scientific and technological terms) used in this article have the same meaning as those commonly understood by ordinary technicians in the field to which this application belongs. It should also be understood that unless there is a clear explanation in this application, words defined in commonly used dictionaries should be interpreted as having the same meaning as their meaning in the context of the relevant technology, and should not be interpreted in an idealized or overly formal sense.

[0061] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0062] Example 1

[0063] Figure 1 An embodiment of the present application provides a flow chart of an intelligent alarm method based on environmental perception, such as Figure 1 As shown, the method includes:

[0064] Get the video shot by the environment perception camera, initialize the video, and divide the video into T frames. Use the camera to collect the human behavior video in the dynamic scene and initialize it, divide the video into T frames, the size of the T frame is T×H×W×3, where T represents the time dimension, H represents the height, W represents the width, and 3 represents the number of color channels. Input the T frame images into the behavior feature extraction network in sequence.

[0065] Feature extraction is performed on the T-frame image based on a behavior recognition feature extraction network to obtain hierarchical features, wherein the behavior recognition feature extraction network is composed of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module.

[0066] Figure 2 is a structural diagram of a behavior recognition feature extraction network framework in an intelligent alarm method based on environment perception provided by an embodiment of the present application, such as Figure 2 As shown in Figure 2, the behavior recognition feature extraction network consists of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module. The T-frame image is input into the three-dimensional lightweight convolution module and used Convolution calculation obtains the features of T frame images , wherein the characteristics of the T frame picture Including local features and global features The 3D lightweight convolution module is a convolution operation used in 3D convolutional neural networks. Its main purpose is to reduce model complexity, thereby improving computational efficiency and model inference speed. Compared with traditional 3D convolution operations, the 3D lightweight convolution module uses a channel attention mechanism, which can effectively reduce the number of parameters and computational complexity while maintaining the accuracy of the model.

[0067] Features of the T-frame picture As the input of the 3D hierarchical attention module, the 3D hierarchical attention can decompose the global self-attention with 3D complexity into multiple layers of attention with lower computational cost. A hierarchical module is introduced into the 3D hierarchical attention to obtain the 3D hierarchical attention module, which is used to summarize the spatiotemporal local features and the spatiotemporal global features and enrich the features.

[0068] The specific steps of extracting hierarchical features by using the three-dimensional hierarchical attention module to obtain hierarchical features are as follows: first, the features of the T-frame images are pooled to obtain hierarchical features, as follows:

[0069]

[0070] in, Represents the pooling operation, which is a commonly used operation in convolutional neural networks. Its main function is to downsample or reduce the dimension of the input feature map, reduce the model calculation amount and memory usage, and extract the main information in the input feature map. Represents hierarchical features, Represents the features of the input T frame image.

[0071] Then, the attention operation is performed on the hierarchical features, as follows:

[0072]

[0073]

[0074] in, It stands for layer normalization. Layer normalization plays the role of normalizing input features in neural networks, which can accelerate network convergence, improve generalization ability and enhance network robustness, thereby improving the performance and effect of the model. Represents 3D multi-head self-attention, which can effectively model spatial relationships, temporal relationships, and relationships in point cloud data when processing 3D data. and It is a learnable cross-channel scale multiplier, which is used for three-dimensional multi-head self-attention and feedforward networks respectively. Its function is to enhance the feature expression ability and reduce the amount of calculation to obtain better three-dimensional lightweight convolution results. represents a feed-forward network.

[0075] The local features of the features of the T frame picture , global features With hierarchical features Combine them and perform attention operations as follows:

[0076]

[0077]

[0078]

[0079] Among them, the first Represents local features , global features and hierarchical features The result after merging is Representation layer normalization, represents three-dimensional multi-head self-attention, and is a learnable cross-channel scale multiplier, represents a feed-forward network, (·) indicates a merge operation.

[0080] Finally, the hierarchical features are output through the three-dimensional hierarchical attention module, specifically:

[0081]

[0082] in, represents the hierarchical features output by the 3D hierarchical attention module, (·) indicates an upsampling operation.

[0083] The fused features obtained by fusing the hierarchical features based on the three-dimensional rotation transformer module include:

[0084] Figure 3 is a schematic diagram of the structure of a three-dimensional rotation transformer module for feature fusion in an intelligent alarm method based on environment perception provided in an embodiment of the present application, such as Figure 3 As shown, the three-dimensional rotation transformer module includes a first block and a second block, which are connected together for use. The first block and the second block are first normalized by a layer, then a three-dimensional multi-head self-attention, then a layer normalized, then a three-dimensional lightweight deep convolution, and finally a feedforward neural network, wherein the feedforward neural network is composed of a two-layer multi-layer perceptron and an activation function GELU;

[0085] The difference between the first block and the second block is that the first block is a three-dimensional sliding window multi-head self-attention, which is used for local feature information exchange within the window, and the second block is a three-dimensional shifted sliding window multi-head self-attention, which is used for global feature information exchange between windows.

[0086] The hierarchical features are input to the first block in the 3D rotation transformer module, and the features obtained by layer normalization and 3D sliding window multi-head self-attention are used Represented by, and then the features obtained by the feedforward neural network are expressed as Indicates, as follows:

[0087]

[0088]

[0089] in, Represents three-dimensional sliding window multi-head self-attention, which is a mechanism for calculating the interaction relationship between sequence data. It divides the input sequence into blocks and then calculates the attention weight in each block to obtain the correlation between each element and other elements, thereby realizing the modeling of sequence data. The three-dimensional sliding window attention can process longer input sequences and does not need to perform attention calculation on the entire sequence. It only needs to calculate the time steps in each block. Since the number of time steps in each block is the same, efficient matrix multiplication can be used for calculation, thereby improving operating efficiency. Representation layer normalization, represents a feed-forward neural network, Represents three-dimensional deep lightweight convolution, which is a combination of depthwise separable convolution and channel-by-channel convolution. Depthwise separable convolution is a method of splitting standard convolution into depthwise convolution and point-by-point convolution, which can reduce the number of parameters and computational complexity of the model. Channel-by-channel convolution is a convolution operation performed on the input signal in the channel dimension, which can more effectively preserve the local correlation of the signal.

[0090] The output of the first block in the 3D Rotation Transformer module As the input of the second block, the features obtained by layer normalization and 3D shift sliding window multi-head self-attention are used Indicates that the features obtained by the feedforward neural network are Indicates, as follows:

[0091]

[0092]

[0093] in, Represents the three-dimensional shifted sliding window multi-head self-attention, which can better capture the temporal and spatial relationships in sequence data compared to the traditional three-dimensional sliding window attention. In the traditional three-dimensional sliding window attention, the attention weights between time steps in each block are obtained through convolution operations in three-dimensional space, while the three-dimensional shifted sliding window attention adopts an idea similar to the dilated convolution and adds a translation operation to the convolution kernel to achieve interaction between different positions. Specifically, the three-dimensional shifted sliding window attention divides the input sequence into multiple blocks of the same size, each block can be regarded as a matrix, where rows represent time steps and columns represent features; within each block, for each time step, the attention weight between the time step and other time steps is calculated; the role of the three-dimensional shifted sliding window attention is to better capture the temporal and spatial relationships in sequence data, thereby improving the modeling ability of sequence data. Output the final fusion feature .

[0094] The fusion feature is input into the behavior classification network for identification and classification to obtain the behavior classification result, and the behavior classification result is input into the early warning network. The behavior classification result is a specific numerical value ranging from 1.00 to 10.00. The threshold is set to 6.00. When the behavior classification result is greater than or equal to the threshold 6.00, it is determined to be a fast running behavior, that is, a warning is triggered. If the behavior classification result is less than the threshold 6.00, it returns to obtain the next time series video.

[0095] Furthermore, the hierarchical features first pass through the first block and then pass through the second block to obtain the fused features.

[0096] Specifically, the first block uses a three-dimensional sliding window multi-head self-attention to specifically include:

[0097] The hierarchical features output by the behavior recognition feature extraction network are used as input data;

[0098] Perform linear transformation on the normalized data and project them into query, key and value spaces respectively to obtain query space mapping value, key space mapping value and value space mapping value;

[0099] Divide the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum;

[0100] After completing the attention calculation and weighted sum operation of all heads, the output results of multiple heads are spliced ​​together according to the channel dimension to obtain the output of the three-dimensional sliding window multi-head self-attention.

[0101] Specifically, the second block uses a three-dimensional shift sliding window multi-head self-attention and specifically includes:

[0102] Normalize the output of the first block and use it as the input data of the second block;

[0103] Perform linear transformation on the normalized data and project them into query, key and value spaces respectively to obtain query space mapping value, key space mapping value and value space mapping value;

[0104] Divide the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum;

[0105] After completing the attention calculation and weighted sum operation of all heads, the output results of multiple heads are spliced ​​together according to the channel dimension to obtain the output of the three-dimensional shifted sliding window multi-head self-attention.

[0106] Specifically, the data input and layer normalization steps in the first block are: the hierarchical features output from the behavior recognition feature extraction network are used as input data, denoted as Z b First, the input data Z b Perform layer normalization to normalize all features of each sample. The formula is:

[0107]

[0108] in Represents the normalized result of the layer, μ is Z b The mean value in the feature dimension, is the variance, is a minimum value that prevents the denominator from being zero. and This application uses layer normalization to make the input data have better distribution characteristics when entering subsequent operations, accelerate network convergence and improve model performance.

[0109] 3D sliding window multi-head self-attention calculation: The data after layer normalization enters the 3D sliding window multi-head self-attention (3DW-MSA) module. In this module, the input data is first projected into three different spaces: query, key, and value, to obtain Q, K, and V. Assuming that the dimension of the input data is [T, H, W, C] (T is the time dimension, H is the height, W is the width, and C is the number of channels), then:

[0110]

[0111]

[0112]

[0113] in, , , is a learnable weight matrix. Then, Q, K, and V are divided into multiple heads, and each head calculates the attention score and weighted sum.

[0114] Taking one head as an example, the attention score is calculated as:

[0115]

[0116] in is the dimension of the key. The weighted sum is O: Finally, the outputs of multiple heads are concatenated to obtain the output of the three-dimensional sliding window multi-head self-attention .

[0117] Normalize the layers again: Output of multi-head self-attention for 3D sliding windows The layer normalization operation is performed again, and its calculation method is the same as the layer normalization in the first step, which is not described in detail in this application.

[0118] 3D Deep Lightweight Convolution: 3D deep lightweight convolution (3DDWConv) is used to perform convolution operations on the re-normalized data. 3D lightweight convolution combines the advantages of depthwise separable convolution and channel-by-channel convolution. First, depthwise convolution is performed, and convolution operations are performed on each channel separately. The convolution kernel size is [k T ,k H ,k W ]. Assuming the input data is X, the output after depth convolution is:

[0119]

[0120] in, is the convolution kernel of the depthwise convolution, , , They are the coordinate indexes of the feature maps in different dimensions, is the channel index. Then, convolution is performed channel by channel, and the output of the depth convolution is convolved in the channel dimension to obtain the final convolution result. .

[0121] Feedforward neural network processing: The output of the three-dimensional lightweight deep convolution is input into the feedforward neural network (FFN). The feedforward neural network consists of a two-layer multilayer perceptron (MLP) and an activation function GELU. Assume that the weight matrix of the first layer of MLP is W1, the bias is b1, the weight matrix of the second layer is W2, and the bias is b2, then the calculation process of FFN is:

[0122]

[0123] After being processed by the feedforward neural network, the final output of the first block is obtained , which will be the input to the second block.

[0124] The implementation steps of the second block are:

[0125] Data input and layer normalization: The output of the first block As the input data of the second block, it is first subjected to layer normalization, and the calculation method is consistent with the layer normalization in the first block to stabilize the data distribution.

[0126] Three-dimensional shifted sliding window multi-head self-attention calculation: The data after layer normalization enters the three-dimensional shifted sliding window multi-head self-attention (3DWS-MSA) module. This module is similar to the three-dimensional sliding window multi-head self-attention module, but introduces a shift operation in window division. First, the input data is also projected into the query, key, and value spaces to obtain Q, K, and V. Then when dividing the window, each window has a certain shift in time, height, and width dimensions relative to the previous window (such as a shift step of 1). Taking one head as an example, the method of calculating the attention score and weighted sum is the same as that of the three-dimensional sliding window multi-head self-attention module, and this application will not go into details. The outputs of multiple heads are spliced ​​together to obtain the output of the three-dimensional shifted sliding window multi-head self-attention.

[0127] Normalize the layer again: Normalize the output of the three-dimensional shift sliding window multi-head self-attention layer to ensure the stability of the data. The calculation method is the same as the previous layer normalization step.

[0128] 3D lightweight depth convolution: Consistent with the 3D lightweight depth convolution operation in the first block, a 3D lightweight depth convolution operation is performed on the re-normalized data. First, depth convolution is performed, and then channel-by-channel convolution is performed to obtain the convolution result.

[0129] Feedforward neural network processing: The output of the three-dimensional lightweight deep convolution is input into the feedforward neural network. The structure and calculation method of the feedforward neural network are the same as those in the first block. After the feedforward neural network processing, the final output of the second block is obtained, that is, the fused feature .

[0130] It should be noted that the traditional three-dimensional transformer model has the problems of large computational complexity and poor fusion of spatiotemporal features. The present application innovatively introduces a lightweight convolution structure into the three-dimensional rotation transformer module, especially adopts a grouped convolution optimization method in the three-dimensional lightweight deep convolution link.

[0131] For example, in traditional 3D convolution, for the input feature map , T is the time dimension, H is the height, W is the width, and C is the number of input channels) and the convolution kernel ,in is the number of output channels, and the convolution calculation is Y=W X, the calculation amount is ), while the grouped convolution divides the input channel number C into G groups, each with a channel number , for the convolution kernel of group g , Represents the number of output channels in each group, and the output feature map Y of the g-th group g Calculated as (X g is the g-th group of data after the input feature map X is divided into groups). The final output feature map Y is concatenated from the outputs of each group. At this time, the amount of calculation becomes ), it can be seen that the present application has greatly reduced the computational complexity without losing too much accuracy. The significant reduction in the amount of calculation directly improves the operating efficiency of the model. In the real-time behavior recognition and warning scenario, the system can process video data faster and shorten the time interval from video acquisition to behavior judgment and warning. Since group convolution can maintain a certain accuracy while reducing the amount of calculation, the feature fusion effect is not significantly affected. The improved module structure can better integrate spatiotemporal features, accurately capture the characteristics of human fast running behavior, and improve the accuracy of behavior classification and warning.

[0132] In addition, in order to further improve the attention mechanism performance of the three-dimensional rotation transformer module in the embodiment of the present application, a dynamic window adjustment mechanism is introduced.

[0133] Specifically, the step of fusing the hierarchical features to obtain the fused features based on the three-dimensional rotation transformer module further includes:

[0134] Define two parameters: window size adjustment factor and step size adjustment factor;

[0135] The displacement of pixels between adjacent frames is calculated by the optical flow method, and then the average motion speed is obtained;

[0136] The feature complexity is measured by calculating the entropy of the feature map;

[0137] The window size and step size of the 3D sliding window and the 3D shift sliding window in the 3D rotation transformer module are adjusted according to the average motion speed and feature complexity.

[0138] Specifically, first, the displacement of pixels between adjacent frames is calculated by the optical flow method to obtain the average motion speed v. Assuming that the total pixel displacement between frame t and frame t+1 is S and the total number of pixels is N, then v=S / N. The feature complexity is evaluated by calculating the entropy of the feature map. p(k) represents the probability of pixel value k in the feature map. The entropy is .

[0139] The size and step size of the sliding window can then be dynamically adjusted based on the speed of movement and feature complexity. When it is detected that the person in the video is moving fast and the movements are changing frequently (the judgment basis is that v is greater than the set threshold), the window step size is increased to quickly capture key features; when the person in the scene moves slowly and the feature details are rich, the window size is reduced to focus on local detail features.

[0140] This application introduces a dynamic window adjustment mechanism to make the attention mechanism more flexible and efficient. When processing video data of different scenes, the window parameters can be adaptively adjusted according to the actual situation to better capture different behavioral characteristics. When the person moves fast, increasing the step length can quickly scan key actions and avoid missing important information; when the movement is slow, reducing the window can more carefully analyze local features, improve the ability to capture behavioral characteristics in complex scenes, thereby improving the accuracy of behavior classification and warning, reducing the occurrence of misjudgment and missed judgment, and enhancing the reliability and practicality of the entire behavior recognition and warning system.

[0141] Furthermore, in other embodiments of the present application, the step of fusing the hierarchical features to obtain the fused features based on the three-dimensional rotation transformer module further includes:

[0142] The input of the first block and the output of the first block are added or concatenated to obtain fused input data;

[0143] The input of the second block is replaced by the fused input data.

[0144] In the 3D rotation transformer module, a skip connection is constructed based on the first block and the second block. The input of the first block is Z input After layer normalization, 3D sliding window multi-head self-attention, layer normalization again, 3D lightweight deep convolution and feedforward neural network processing in the first block, the output is Z first-block .

[0145] To achieve skip connection, Z input Directly with Z first-block The fusion is performed by addition or concatenation. If addition fusion is used, the input of the second block is:

[0146]

[0147] If splicing fusion is used, assuming Z input With Z first-block The dimensions are [T,H,W,C1] and [T,H,W,C2] respectively, then the dimensions after concatenation are [T,H,W,C1+C2], that is, the input of the second block is:

[0148]

[0149] Wherein, Concat(·) represents the concatenation function.

[0150] In the second block processing, whether the result of addition or concatenation is used as input, the features of the early layer (the first block input) can be directly involved in the processing of the second block. During back propagation, the gradient can flow back to the shallow layer more smoothly from the deep layer through the jump connection, alleviating the gradient vanishing problem while retaining the key information in the original features, which helps the model learn richer spatiotemporal features.

[0151] Furthermore, in a preferred embodiment of the present application, a branch fusion solution is also provided. Specifically, the three-dimensional rotation transformer module is used to fuse the hierarchical features to obtain the fused features, and further includes:

[0152] Use two-dimensional convolution to perform convolution processing on the spatial dimension of the input data of the first block to extract spatial local feature data;

[0153] Use two-dimensional sliding window multi-head self-attention to replace three-dimensional sliding window multi-head self-attention, and use one-dimensional shift sliding window multi-head self-attention to replace three-dimensional shift sliding window multi-head self-attention;

[0154] Re-using the spatial local feature data as input data of the first block;

[0155] Perform a time dimension convolution operation on the output of the first block to extract time series features;

[0156] The output of the first block is fused with the output of the second block to obtain the fused feature.

[0157] The branch fusion is described in detail below.

[0158] 1. Spatial feature branch

[0159] Data input and preprocessing: The input Z of the three-dimensional rotation transformer module input Input to the spatial feature branch. First, a preprocessing operation specifically for spatial feature extraction is performed, using two-dimensional convolution to perform convolution processing on the spatial dimensions (height and width) to extract spatial local features. Assume that the two-dimensional convolution kernel used is W 2D-spatial , then the output after convolution is:

[0160]

[0161] where k H and k W is the size of the two-dimensional convolution kernel in height and width, The two colons in represent taking all elements in the corresponding dimension. The first two colons in represent taking all elements in the corresponding dimension. and The colon in the and Starting from position , select all subsequent elements in the corresponding dimension.

[0162] Spatial feature extraction and processing: Z spatial-pre Perform layer normalization to stabilize data distribution. Then enter the spatial attention module, which can be based on a two-dimensional attention mechanism, such as two-dimensional multi-head self-attention (2DMHSA). Calculate the attention score and weighted sum to obtain an output that is more focused on spatial features. After that, it goes through a lightweight three-dimensional convolution to further integrate spatial features and a small amount of time dimension information (here the convolution kernel size of the time dimension can be set to 1, that is, convolution operation is performed only in the spatial dimension), and obtain the final output Z of the spatial feature branch. spatial-output .

[0163] 2. Temporal feature branch

[0164] Data input and preprocessing: Z input Input to the time feature branch. First, preprocess the time feature extraction, use one-dimensional convolution to perform convolution operation on the time dimension (T) to extract time series features. Suppose the one-dimensional convolution kernel is , then the output after convolution is:

[0165]

[0166] where k T is the size of the one-dimensional convolution kernel in the time dimension.

[0167] Temporal feature extraction and processing: After the output result after convolution is layer-normalized, it enters the temporal attention module, such as one-dimensional multi-head self-attention (1DMHSA), calculates the attention score and weighted sum, and obtains the output focused on the temporal feature. Then, through a lightweight three-dimensional convolution (the convolution kernel size of the spatial dimension is set to 1, and the convolution operation is mainly performed in the temporal dimension), the temporal feature and a small amount of spatial dimension information are further integrated to obtain the final output Z of the temporal feature branch. temporal-output .

[0168] 3. Branch Fusion

[0169] The output Z of the spatial feature branch spatial-output And the output Z of the temporal feature branch temporal-output Fusion. Using weighted fusion, assuming the weight of the spatial feature branch is α, and the weight of the temporal feature branch is 1-α, the fused output is:

[0170]

[0171] The weight α can be a fixed value set according to experimental experience, or it can be a learnable parameter, which allows the model to automatically adjust through training to adapt to different behavioral data and scenario requirements. In this way, the dual-branch parallel structure focuses on processing spatial and temporal features respectively, and then integrates the results, which can improve the model's comprehensive processing capabilities for spatiotemporal features and better adapt to complex behavioral patterns.

[0172] Example 2

[0173] Figure 4 An intelligent alarm system based on environmental perception is provided in one embodiment of the present application, and the system includes:

[0174] The video acquisition and initialization module is used to acquire the video captured by the environment perception camera, initialize the video, and divide the video into T frame images;

[0175] Behavior recognition feature extraction module, used for:

[0176] Based on the behavior recognition feature extraction network, feature extraction is performed on T-frame images to obtain hierarchical features;

[0177] 3D Rotation Transformer Feature Fusion Module for:

[0178] The hierarchical features are fused based on the three-dimensional rotation transformer module to obtain fused features;

[0179] Classification module for:

[0180] The fusion features are input into the classification module for identification and classification to obtain the behavior classification results, and the behavior classification results are input into the early warning network;

[0181] Early warning module, used to:

[0182] A threshold is set in the early warning network. When the behavior classification result is greater than or equal to the set threshold, it is judged as fast running behavior, which triggers an early warning. When the behavior classification result is less than the threshold, it returns to obtain the video of the next time period.

[0183] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0184] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

[0185] The above order of steps for the method is for illustration only, and the steps of the method of the present application are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present application may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present application. Thus, the present application also covers a recording medium storing a program for executing the method according to the present application.

[0186] The specific implementation modes as described above further describe the purpose, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation mode of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An intelligent alarm method based on environmental perception, characterized in that: include: Get the video captured by the environment perception camera, initialize the video, and divide the video into T frames; Extracting features from the T-frame images based on a behavior recognition feature extraction network to obtain hierarchical features, wherein the behavior recognition feature extraction network is composed of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module; The hierarchical features are fused based on the three-dimensional rotation transformer module to obtain fused features; Inputting the fusion features into a behavior classification network to perform behavior recognition and classification, obtaining a behavior classification result, and inputting the behavior classification result into an early warning network; A threshold is set in the early warning network. When the behavior classification result is greater than or equal to the threshold, it is determined to be a fast running behavior, and an early warning is triggered. When the behavior classification result is less than the threshold, it returns to obtain the next time series video.

2. The intelligent alarm method based on environmental perception according to claim 1 is characterized in that: The step of obtaining a video captured by an environment perception camera, initializing the video, and dividing the video into T frames of pictures includes: Use the camera to collect and initialize human behavior videos in dynamic scenes; The video is divided into T frame pictures, and the size of the T frame pictures is T×H×W×3, where T represents the time dimension, H represents the height, W represents the width, and 3 represents the number of color channels.

3. The intelligent alarm method based on environmental perception according to claim 1 is characterized in that: The step of extracting features from the T-frame images based on the behavior recognition feature extraction network to obtain hierarchical features includes: Inputting the T frame images sequentially into the three-dimensional lightweight convolution module in the behavior recognition feature extraction network; use Convolution calculation obtains the features of T frame images , wherein the characteristics of the T frame picture Include local features and global features ; The features of the T frame image are input into a three-dimensional hierarchical attention module to extract hierarchical features to obtain hierarchical features.

4. The intelligent alarm method based on environmental perception according to claim 1 is characterized in that: The three-dimensional rotation transformer module includes a first block and a second block, and the fusion feature obtained by fusing the hierarchical features based on the three-dimensional rotation transformer module includes: The hierarchical features first pass through the first block and then pass through the second block to obtain the fused features; both the first block and the second block first pass through a layer normalization, then a three-dimensional multi-head self-attention, then a layer normalization, then a three-dimensional lightweight deep convolution, and finally a feedforward neural network; the difference between the first block and the second block is that the first block uses a three-dimensional sliding window multi-head self-attention, and the second block uses a three-dimensional shifted sliding window multi-head self-attention.

5. The intelligent alarm method based on environmental perception according to claim 4 is characterized in that: The first block uses a three-dimensional sliding window multi-head self-attention and specifically includes: The hierarchical features output by the behavior recognition feature extraction network are used as input data; Perform linear transformation on the normalized data and project them into query, key and value spaces respectively to obtain query space mapping value, key space mapping value and value space mapping value; Divide the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum; After completing the attention calculation and weighted sum operation of all heads, the output results of multiple heads are spliced ​​together according to the channel dimension to obtain the output of the three-dimensional sliding window multi-head self-attention.

6. The intelligent alarm method based on environmental perception according to claim 4 is characterized in that: The second block uses a three-dimensional shift sliding window multi-head self-attention and specifically includes: Normalize the output of the first block and use it as the input data of the second block; Perform linear transformation on the normalized data and project them into query, key and value spaces respectively to obtain query space mapping value, key space mapping value and value space mapping value; Divide the query space mapping value, key space mapping value, and value space mapping value into multiple heads, and each head independently calculates the attention score and weighted sum; After completing the attention calculation and weighted sum operation of all heads, the output results of multiple heads are spliced ​​together according to the channel dimension to obtain the output of the three-dimensional shifted sliding window multi-head self-attention.

7. The intelligent alarm method based on environmental perception according to claim 4 is characterized in that: The step of fusing the hierarchical features based on the three-dimensional rotation transformer module to obtain the fused features further includes: Define two parameters: window size adjustment factor and step size adjustment factor; The displacement of pixels between adjacent frames is calculated by the optical flow method, and then the average motion speed is obtained; The feature complexity is measured by calculating the entropy of the feature map; The window size and step size of the 3D sliding window and the 3D shift sliding window in the 3D rotation transformer module are adjusted according to the average motion speed and feature complexity.

8. The intelligent alarm method based on environmental perception according to claim 4 is characterized in that: The step of fusing the hierarchical features based on the three-dimensional rotation transformer module to obtain the fused features further includes: The input of the first block and the output of the first block are added or concatenated to obtain fused input data; The input of the second block is replaced by the fused input data.

9. The intelligent alarm method based on environmental perception according to claim 8 is characterized in that: The step of fusing the hierarchical features based on the three-dimensional rotation transformer module to obtain the fused features further includes: Use two-dimensional convolution to perform convolution processing on the spatial dimension of the input data of the first block to extract spatial local feature data; Use two-dimensional sliding window multi-head self-attention to replace three-dimensional sliding window multi-head self-attention, and use one-dimensional shift sliding window multi-head self-attention to replace three-dimensional shift sliding window multi-head self-attention; Re-using the spatial local feature data as input data of the first block; Perform a time dimension convolution operation on the output of the first block to extract time series features; The output of the first block is fused with the output of the second block to obtain the fused feature.

10. An intelligent alarm system based on environmental perception, characterized in that: The system comprises: The video acquisition and initialization module is used to acquire the video captured by the environment perception camera, initialize the video, and divide the video into T frame images; A behavior recognition feature extraction module, used for extracting features from the T-frame images based on a behavior recognition feature extraction network to obtain hierarchical features, wherein the behavior recognition feature extraction network is composed of a three-dimensional lightweight convolution module and a three-dimensional hierarchical attention module; A three-dimensional rotation transformer feature fusion module, used for fusing hierarchical features based on the three-dimensional rotation transformer module to obtain fused features; A classification module, used for inputting the fusion features into a behavior classification network for behavior recognition and classification, obtaining a behavior classification result, and inputting the behavior classification result into an early warning network; The early warning module is used to set a threshold in the early warning network. When the behavior classification result is greater than or equal to the threshold, it is determined to be a fast running behavior, that is, triggering an early warning. When the behavior classification result is less than the threshold, it returns to obtain the next time series video.

Citation Information

Patent Citations

  • Video action recognition method

    CN110765854A

  • Signal type identification method and system based on fusion feature and group convolution ViT network

    CN117743946A

  • Signal searching method based on wavelet convolution and Former multi-scale feature fusion

    CN117830787A

  • Behavior recognition model training method and system fusing 3DCNN and attention mechanism

    CN118334752A

  • Mobile terminal alarm device and method based on artificial intelligence

    CN119155378A