A method for abnormal behavior recognition based on dual-stream attention graph convolution

By introducing time, space and channel attention modules into the abnormal behavior recognition algorithm and combining the dual-stream fusion network, the problem of failing to fully utilize the simple and crude human behavior characteristics and model integration methods in the existing technology is solved, and a more efficient abnormal behavior recognition and fusion effect is achieved.

CN115171206BActive Publication Date: 2025-05-06SOUTH CHINA UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210161226.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-22
Publication Date
2025-05-06
Estimated Expiration
2042-02-22

AI Technical Summary

Technical Problem

The existing abnormal behavior recognition algorithms fail to fully utilize the particularity of human behavior and cannot exert the superiority of attention mechanisms. The model integration method is simple and crude, and the optimal fusion result cannot be guaranteed.

Method used

A method of abnormal behavior recognition based on dual-stream attention graph convolution is proposed. By modifying the graph convolution network, the time attention module, spatial attention module and channel attention module are introduced, and combined with the dual-stream fusion network, the recognition effect of abnormal behavior of human bodies is improved.

Benefits of technology

By introducing attention modules and dual-stream fusion networks, the model can more effectively extract human behavior characteristics, improve the accuracy and fusion effect of abnormal behavior recognition, and enhance the protection of personnel safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171206B_ABST
    Figure CN115171206B_ABST
Patent Text Reader

Abstract

The present invention discloses an abnormal behavior recognition method based on dual-stream attention graph convolution. Aiming at the movement characteristics of the human body, a time attention module, a space attention module, and a channel attention module are proposed on the basis of the existing graph convolution network. The above modules can be directly inserted into any graph convolution to enhance the model performance of the graph convolution. During reasoning, the joint point information and the skeleton information of the human body are respectively input into the corresponding graph convolution network for joint point feature extraction and the graph convolution network for skeleton feature extraction to obtain preliminary classification results, and then the results are sent to the trained dual-stream fusion network to calculate the optimal fusion parameters, and then the final classification results are obtained. The present invention can improve the feature extraction performance of the graph convolution network model, enhance the effect of dual-stream result fusion, and effectively improve the accuracy of abnormal human behavior recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of abnormal behavior recognition, and in particular to an abnormal behavior recognition method based on dual-stream attention graph convolution. Background Art

[0002] With the rapid development of science and technology, the level of social intelligence has been greatly improved, especially with the popularization of video surveillance devices. Various intelligent algorithms based on image processing have emerged one after another, which greatly reduced the workload of personnel. Abnormal behavior recognition, as one of the tasks in video processing, is of great significance to personnel safety. Especially in public places with a large number of people, such as subways, shopping malls, and squares, when people fall or fight, if they cannot be discovered in time, it is very easy to cause regional chaos, trampling and other accidents, which seriously threaten life safety. Therefore, it is particularly important to detect abnormal behavior of people in a timely manner. At present, most of the mainstream abnormal behavior recognition algorithms are based on skeleton graph convolution classification networks, which have the advantages of small computational complexity and high generalization. However, they do not make full use of the particularity of human behavior and cannot give play to the superiority of the attention mechanism. In addition, in the process of model integration, it is generally adopted to train multiple models with different inputs and average the prediction results of all models to achieve fusion. This method is simple and crude and cannot guarantee the optimal fusion result. Therefore, there is still a lot of room for improvement in the model effect of human abnormal behavior classification. Summary of the invention

[0003] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art and propose an abnormal behavior recognition method based on dual-stream attention graph convolution. This method modifies the existing graph convolution network and proposes a dual-stream fusion network, which can effectively improve the recognition effect of abnormal human behavior.

[0004] To achieve the above object, the technical solution provided by the present invention is: an abnormal behavior recognition method based on dual-stream attention graph convolution, comprising the following steps:

[0005] 1) Obtain video segments captured by cameras that contain abnormal human behavior and label the abnormal behavior category of each person;

[0006] 2) Use the key point extraction network to obtain the human joint point sequence in the video, combine it with the corresponding abnormal behavior category to produce a behavior classification joint point dataset, convert the human joint point sequence into a human skeleton sequence to produce a behavior classification skeleton dataset, and divide the behavior classification joint point dataset and the behavior classification skeleton dataset into training set and validation set in the same proportion;

[0007] 3) Use the training sets of the behavior classification joint point dataset and the behavior classification skeleton dataset to train two improved MS-G3D networks, and use the corresponding validation sets to verify the model accuracy to select the optimal model parameters;

[0008] 4) The human joint point sequence and human skeleton sequence of the same person are predicted using the trained corresponding improved MS-G3D network, and the two prediction results and the abnormal behavior category of the person are combined to produce a two-stream fusion dataset, which is divided into a training set and a validation set;

[0009] 5) Use the training set of the two-stream fusion data set to train the two-stream fusion network, and use the corresponding validation set to verify the model accuracy to select the optimal model parameters;

[0010] 6) The passenger skeleton is extracted and tracked for each frame of the video to be detected to obtain the human joint point sequence and the human bone sequence. The two sequences are respectively sent to the trained corresponding improved MS-G3D network for prediction. The two prediction results are sent to the trained dual-stream fusion network for prediction. The prediction result is the final abnormal behavior category of the passenger.

[0011] Further, in step 3), the improved parts of the improved MS-G3D network are specifically as follows: three modules are added on the basis of the MS-G3D network, namely, a temporal attention module, a spatial attention module, and a channel attention module; the temporal attention module includes a frame difference information extraction submodule, a multi-time scale feature extraction module and an output module, and the spatial attention module includes a node difference information extraction submodule, a multi-node scale feature extraction module and an output module.

[0012] Furthermore, the channel attention module includes an average pooling layer, a 1×1 convolution layer, a ReLU activation function, a Sigmoid activation function, a BN layer and two cross connections, wherein the average pooling layer is used to compress the original input Input of the channel attention module oC The temporal and spatial information of the input is obtained by using a 1×1 convolution layer and a ReLU activation function to reduce the channel dimension and reduce the amount of calculation. Then, a 1×1 convolution layer and a Sigmoid activation function are used to increase the channel dimension and obtain the weight coefficient of each channel. The first cross connection is used to connect the Input oC Multiply it with the weight coefficient and connect the result to the Input using the second cross connection oC Finally, the BN layer and ReLU activation function are connected to adjust the data distribution and increase the nonlinear fitting ability of the model to obtain the final channel attention module output result Output C .

[0013] Furthermore, the frame difference information extraction submodule includes two average pooling layers, a square layer and a cross connection, wherein the first average pooling layer is used to obtain the original input Input of the temporal attention module oT The average feature of each node in all frames is used as the average skeleton, and then connected to the input oT The difference is then calculated, and each element is squared through the square layer. Finally, the spatial dimension is averaged through the second average pooling layer to obtain the output result of the frame difference information extraction submodule Output sT , that is, the average value of all node features of a frame is taken as the importance of the frame, so that the greater the difference from the average skeleton, the higher the importance of the frame.

[0014] Furthermore, the multi-time scale feature extraction module includes a dilated convolution layer, a concat layer, an average pooling layer and a sigmoid layer, wherein multiple 13×1 dilated convolution layers with expansion coefficients of 1, 5, 9, 13, 17, 21 and 24 are connected in parallel, and the input is also the output result of the frame difference information extraction submodule Output sT , use the Concat layer to concatenate the outputs of all the dilated convolutional layers according to the channel dimension, then use the average pooling layer to compress the channel dimension, and pass the Sigmoid layer to obtain the importance coefficients P of different frames. T .

[0015] Furthermore, the output module includes a cross-connection layer, a BN layer and a ReLU activation function, wherein the cross-connection layer converts the importance coefficients P of different frames into T Or the importance coefficient P of different nodes V The original input of the temporal attention module oT Or the original input of the spatial attention module oV Multiply and then add to Input oT or Input oV Add them together, and finally connect the BN layer and the ReLU activation function to obtain the final output result of the temporal attention module or the spatial attention module.

[0016] Furthermore, the node difference information extraction submodule includes two average pooling layers, a square layer and a cross connection, wherein the first average pooling layer is used to obtain the original input Input of the spatial attention module oV The average feature of each node in all frames is then connected to the input oV The difference is then calculated by the square layer to square each element, and finally the second average pooling layer is used to average the time dimension to obtain the output result of the node difference information extraction submodule Output sV, that is, the average value of all frame features of a node is taken as the importance of the node, so that the more drastic the change, the higher the importance of the node.

[0017] Furthermore, the multi-node scale feature extraction module includes a hole convolution layer, a Concat layer, an average pooling layer and a Sigmoid activation function, wherein multiple 3×1 hole convolution layers with expansion coefficients of 1, 3, 5, 7, 9, 11, and 12 are connected in parallel, and the input is also the output result of the node difference information extraction submodule Output sV , use the Concat layer to concatenate the outputs of all the dilated convolutional layers according to the channel dimension, then use the average pooling layer to compress the channel dimension, and use the Sigmoid activation function to obtain the importance coefficients P of different nodes. V .

[0018] Further, in step 5), the dual-stream fusion network includes a channel importance extraction module and a stream importance extraction module;

[0019] The channel importance extraction module includes a BN layer, a Concat layer, a fully connected layer, a ReLU activation function, and a Sigmoid activation function. First, the dual-stream inputs with the number of channels C are normalized using the BN layer, and then the Concat layer is used for feature concatenation. Then, the output with the number of channels 2C is obtained through a fully connected layer, a ReLU activation function, a fully connected layer, and a Sigmoid activation function, corresponding to the importance of each channel of the dual-stream input;

[0020] The stream importance extraction module includes an average pooling layer, a maximum pooling layer, a Concat layer, a fully connected layer, a ReLU activation function and a Sigmoid activation function. First, the dual-stream inputs with the number of channels C are connected to the average pooling layer and the maximum pooling layer respectively, and then the Concat layer is used to perform feature concatenation to obtain a feature vector with a channel number of 4, and then the fully connected layer, the ReLU activation function, the fully connected layer, the ReLU activation function, the fully connected layer, and the Sigmoid activation function are passed to obtain an output with a channel number of 2, corresponding to the overall importance of different streams.

[0021] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0022] 1. The temporal attention module is introduced to obtain the average skeleton based on the motion features of multiple frames. It is used to measure the difference between the skeletons of different frames and the average skeleton. The greater the difference, the more important the frame is to the final behavior classification. This makes the model more focused on important frames, which is conducive to improving the model effect.

[0023] 2. Introduce the spatial attention module to obtain the average features of the nodes based on the motion features of multiple frames. This is used to measure the degree of change of different nodes over a period of time. The more frequent the change, the more important the node is. This allows the model to focus on more important nodes and improve the model's feature extraction capabilities.

[0024] 3. Introduce the channel attention module to adaptively weaken useless channel features and increase the importance of effective channel features, so that the model relies more on important channels, alleviate the impact of useless features, and improve the recognition effect of the model.

[0025] 4. The dual-stream fusion network can adaptively learn the optimal fusion model based on the training data. Compared with the direct average fusion method, it can make full use of the distribution characteristics of different data streams and greatly improve the fusion effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is the network structure diagram of the channel attention module.

[0027] Figure 2 This is the network structure diagram of the temporal attention module.

[0028] Figure 3 This is the network structure diagram of the frame difference information extraction submodule.

[0029] Figure 4 This is the network structure diagram of the spatial attention module.

[0030] Figure 5 This is the network structure diagram of the node difference information extraction submodule.

[0031] Figure 6 This is the dual-stream fusion network structure diagram.

[0032] Figure 7 The figure is a logical flow chart of the method of the present invention. DETAILED DESCRIPTION

[0033] The present invention is further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.

[0034] This embodiment discloses an abnormal behavior recognition method based on dual-stream attention graph convolution, the details of which are as follows:

[0035] Step 1: Obtain video segments captured by the camera containing abnormal human behavior and label the abnormal behavior category of each person.

[0036] Step 2: Use the key point extraction network to obtain the human joint point sequence of each person in the video, combine it with the corresponding abnormal behavior category to produce a behavior classification joint point dataset, convert the human joint point sequence into a human skeleton sequence by subtracting the coordinates of adjacent human joint points to produce a behavior classification skeleton dataset, and divide the behavior classification joint point dataset and the behavior classification skeleton dataset into training set and validation set in a ratio of 9:1.

[0037] Step 3: Use the training sets of the behavior classification joint point dataset and the behavior classification skeleton dataset to train two improved MS-G3D networks. The specific improvements of the network are as follows:

[0038] Three modules are added to the original MS-G3D network, namely the temporal attention module, the spatial attention module, the channel attention module and the multi-stream fusion network. Figure 2 As shown in Fig. , the temporal attention module includes a frame difference information extraction submodule, a multi-time scale feature extraction module (module 1) and an output module (module 2). Figure 3 As shown, the spatial attention module includes a node difference information extraction submodule, a multi-node scale feature extraction module (module 3) and an output module (module 2).

[0039] like Figure 1 As shown in Figure 2, the channel attention module contains an average pooling layer (AvgPool), a 1×1 convolution layer (Conv1×1), a ReLU activation function, a Sigmoid activation function, a BN layer, and two cross connections. Assume that the original input of the channel attention module is oC The dimension is C×T×V, and the average pooling layer is used to compress the Input oC The temporal and spatial information of the network changes the dimension to C×1×1, and then a 1×1 convolutional layer and ReLU activation function are used to reduce the channel dimension and reduce the amount of calculation, and the dimension becomes Then use the 1×1 convolution layer and the Sigmoid activation function to increase the channel dimension and obtain the weight coefficient of each channel. The dimension is C×1×1. Use the first cross connection to connect the Input oC Multiply it with the weight coefficient and connect the result to the Input using the second cross connection oC Finally, the BN layer and ReLU activation function are connected to adjust the data distribution and increase the nonlinear fitting ability of the model to obtain the final channel attention module output result Output C , the dimensions are C×T×V.

[0040] like Figure 3 As shown, the frame difference information extraction submodule includes two average pooling layers, a square layer (Square) and a cross connection. Assume that the original input of the temporal attention module is oTThe dimension is C×T×V. The first average pooling layer is used to obtain the average feature of each node in all frames, that is, average pooling is performed in the time dimension (dim=2), and the feature is used as the average skeleton. The average skeleton dimension is C×1×V, and then it is connected to the Input oT The difference is then calculated, and each element is squared through the square layer. At this time, the dimension is C×T×V. Finally, the second average pooling layer is used to average the spatial dimension (dim=3) to obtain the output result of the frame difference information extraction submodule Output sT , whose dimension is C×T×1, that is, the average value of all node features of a frame is taken as the importance of the frame, so that the greater the difference from the average skeleton, the higher the importance of the frame.

[0041] like Figure 2 As shown in the figure, module 1 is a multi-time scale feature extraction module, which includes a dilated convolution layer, a concat layer, an average pooling layer and a sigmoid activation function. Multiple dilated convolution layers with a 13×1 convolution kernel size and dilation coefficients of 1, 5, 9, 13, 17, 21 and 24 are connected in parallel, and the input is also the output result of the frame difference information extraction submodule Output sT , the output dimensions are all 1×T×1, and the Concat layer is used to concatenate the outputs of all the hole convolution layers according to the channel dimension, so that the dimension becomes 7×T×1, and then the average pooling layer is used to compress the channel dimension, and the importance coefficients P of different frames are obtained through the Sigmoid layer. T , whose dimension is 1×T×1.

[0042] like Figure 2 and Figure 4 Module 2 is the output module, which includes a cross-connection layer, a BN layer, and a ReLU activation function. The cross-connection layer converts the importance coefficients P of different frames into T Or the importance coefficient P of different nodes V With the original input of the temporal attention module oT Or the original input of the spatial attention module oV Multiply and then add to Input oT or Input oV Add them together, and finally connect the BN layer and the ReLU activation function to obtain the output result of the final temporal attention module or spatial attention module, whose dimension is C×T×V.

[0043] like Figure 5 As shown in the figure, the node difference information extraction submodule contains two average pooling layers, a square layer and a cross connection. Assume that the original input of the spatial attention module is oVThe dimension is C×T×V. The first average pooling layer is used to obtain the average features of each node in all frames, that is, average pooling is performed on the time dimension to obtain a result with a dimension of C×1×V, and then the result is connected to the input through a cross connection. oV The difference is then calculated, and each element is squared through the square layer, so that the dimension becomes C×T×V. Finally, the second average pooling layer is used to average the time dimension to obtain the output result of the node difference information extraction submodule Output sV , whose dimension is C×1×V, that is, the average value of all frame features of a node is taken as the importance of the node, so that the more drastic the change, the higher the importance of the node.

[0044] like Figure 4 As shown in the figure, module 3 is a multi-node scale feature extraction module, which includes a hole convolution layer, a Concat layer, an average pooling layer and a Sigmoid activation function. Among them, multiple 3×1 hole convolution layers with expansion coefficients of 1, 3, 5, 7, 9, 11, and 12 are connected in parallel, and the input is also the output result of the node difference information extraction submodule Output sV , the output dimensions are all 1×1×V, and the Concat layer is used to concatenate the outputs of all the dilated convolutional layers according to the channel dimension, so that the dimension becomes 7×1×V. Then the average pooling layer is used to compress the channel dimension, and the importance coefficients P of different nodes are obtained through the Sigmoid layer. V , whose dimensions are 1×1×V.

[0045] Two improved MS-G3D networks were trained using the training sets of the behavior classification joint point dataset and the behavior classification skeleton dataset respectively. The Adam optimization method and the initial learning rate of 0.001 were used for training. The performance of the current model was verified using the validation set every 10 epochs, and the model parameters with the highest validation set accuracy in the first 100 epochs were selected as the prediction model.

[0046] Step 4: Use the trained corresponding improved MS-G3D network to predict the human joint point sequence and human skeleton sequence of the same person, and combine the two prediction results and the abnormal behavior category of the person to produce a two-stream fusion data set, and divide the training set and the validation set into a certain proportion;

[0047] Step 5: Use the training set of the two-stream fusion dataset to train the two-stream fusion network. Its specific structure is as follows:

[0048] like Figure 6As shown, the dual-stream fusion network includes a stream importance extraction module (module 4) and a channel importance extraction module (module 5). The inputs of the dual-stream fusion network are the network model output results Input1 with joint points as input and Input2 with bones as input. The labels are the real labels of the training set. The skeleton information to be predicted is sent to the above two networks, and the output is sent to the dual-stream fusion network to obtain the final prediction result. Module 4 is a channel importance extraction module, which includes a BN layer, a Concat layer, a fully connected layer, a ReLU activation function, and a Sigmoid activation function. First, the dual-stream inputs with a channel number of C are normalized using the BN layer, and then the Concat layer is used for feature splicing. After the fully connected layer, the ReLU activation function, the fully connected layer, and the Sigmoid activation function, the output P with a channel number of 2C is obtained. L , each feature corresponds to the importance of each channel of the dual-stream input. Module 5 is the stream importance extraction module, which includes an average pooling layer, a maximum pooling layer (AvgPool), a Concat layer, a fully connected layer, a ReLU activation function, and a Sigmoid activation function. First, the dual-stream input with the number of channels C is connected to the average pooling layer and the maximum pooling layer respectively, and then the Concat layer is used for feature concatenation to obtain a feature vector with a channel number of 4. After the fully connected layer, the ReLU activation function, the fully connected layer, the ReLU activation function, the fully connected layer, the Sigmoid activation function, the output P with a channel number of 2 is obtained. G , corresponding to the overall importance of different streams. Input1 and Input2 are concatenated together with P L and P G Multiply them, and then add the i-th dimension and the i+C-th dimension of the result to finally get the C-dimensional output, which is the final result after the dual-stream fusion. The category corresponding to the maximum value is taken as the category of the skeleton information.

[0049] Use Pytorch to build a two-stream fusion network model, and use the SGD optimizer to update the parameters. The learning rate is set to 0.001. The validation set is used to verify the accuracy of the current model every 5 epochs. The training stops after 100 epochs. The model with the highest accuracy in the validation set is selected as the final model parameters. This model can be used for two-stream model fusion to improve the accuracy of behavior recognition.

[0050] Step 6: Figure 7As shown in the figure, the passenger skeleton is extracted and tracked for each frame of the video to be detected to obtain the human joint point sequence and the human bone sequence, and the joint point information and bone information of the human body to be predicted are respectively sent to the two corresponding trained improved MS-G3D networks (i.e., graph convolutional networks 1 and 2) to obtain two preliminary classification results, and then these two classification results are sent to the dual-stream fusion network to calculate the optimal channel and stream fusion coefficients, and finally the optimal fusion classification result is obtained.

[0051] In summary, by adopting the above scheme, the present invention improves the accuracy of the algorithm in identifying abnormal human behavior, effectively protects the life safety of personnel, has practical promotion value, and is worthy of promotion.

[0052] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A method for abnormal behavior recognition based on dual-stream attention graph convolution, characterized in that: The following steps are involved: 1) Obtain video segments captured by cameras that contain abnormal human behavior and label each person’s abnormal behavior category; 2) Use the key point extraction network to obtain the human joint point sequence in the video, combine it with the corresponding abnormal behavior category to produce a behavior classification joint point dataset, convert the human joint point sequence into a human skeleton sequence to produce a behavior classification skeleton dataset, and divide the behavior classification joint point dataset and the behavior classification skeleton dataset into training set and validation set in the same proportion; 3) Use the training sets of the behavior classification joint point dataset and the behavior classification skeleton dataset to train two improved MS-G3D networks, and use the corresponding validation sets to verify the model accuracy to select the optimal model parameters; The improved parts of the improved MS-G3D network are as follows: three modules are added on the basis of the MS-G3D network, namely, a time attention module, a space attention module, and a channel attention module; the time attention module includes a frame difference information extraction submodule, a multi-time scale feature extraction module and an output module, and the space attention module includes a node difference information extraction submodule, a multi-node scale feature extraction module and an output module; The channel attention module includes an average pooling layer, a 1×1 convolution layer, a ReLU activation function, a Sigmoid activation function, a BN layer, and two cross connections, where the average pooling layer is used to compress the original input of the channel attention module. The temporal and spatial information of the network is obtained by using a 1×1 convolution layer and a ReLU activation function to reduce the channel dimension and reduce the amount of calculation. Then, a 1×1 convolution layer and a Sigmoid activation function are used to increase the channel dimension and obtain the weight coefficient of each channel. The first cross connection is used to Multiply it by the weight factor and connect the result to the second stride Finally, the BN layer and ReLU activation function are connected to adjust the data distribution and increase the nonlinear fitting ability of the model to obtain the final channel attention module output result. ; The frame difference information extraction submodule includes two average pooling layers, a square layer and a cross connection, where the first average pooling layer is used to obtain the original input of the temporal attention module. The average feature of each node in all frames is used as the average skeleton, and then connected with The subtraction is performed, and then each element is squared through the square layer, and finally the spatial dimension is averaged through the second average pooling layer to obtain the output result of the frame difference information extraction submodule. , that is, the average value of all node features of a frame is taken as the importance of the frame, so that the greater the difference from the average skeleton, the higher the importance of the frame; The multi-time scale feature extraction module includes a hole convolution layer, a Concat layer, an average pooling layer and a Sigmoid layer, wherein multiple 13×1 hole convolution layers with expansion coefficients of 1, 5, 9, 13, 17, 21, and 24 are connected in parallel, and the input is also the output result of the frame difference information extraction submodule. , use the Concat layer to concatenate the outputs of all the dilated convolutional layers according to the channel dimension, then use the average pooling layer to compress the channel dimension, and pass the Sigmoid layer to obtain the importance coefficients of different frames ; The output module includes a cross-connection layer, a BN layer and a ReLU activation function, wherein the cross-connection layer converts the importance coefficients of different frames into Or the importance coefficients of different nodes With the original input of the temporal attention module Or the original input of the spatial attention module Multiply and then or Add, and finally connect the BN layer and ReLU activation function to obtain the final output of the temporal attention module or the spatial attention module; The node difference information extraction submodule includes two average pooling layers, a square layer and a cross connection, where the first average pooling layer is used to obtain the original input of the spatial attention module. The average feature of each node in all frames is then connected to The difference is then calculated, and each element is squared through the square layer. Finally, the second average pooling layer is used to average the time dimension to obtain the output result of the node difference information extraction submodule. , that is, the average value of all frame features of a node is taken as the importance of the node, so that the more drastic the change, the higher the importance of the node; The multi-node scale feature extraction module includes a hole convolution layer, a Concat layer, an average pooling layer and a Sigmoid activation function, wherein multiple 3×1 hole convolution layers with expansion coefficients of 1, 3, 5, 7, 9, 11, and 12 are connected in parallel, and the input is also the output result of the node difference information extraction submodule. , use the Concat layer to concatenate the outputs of all the dilated convolutional layers according to the channel dimension, then use the average pooling layer to compress the channel dimension, and use the Sigmoid activation function to obtain the importance coefficients of different nodes. ; 4) The human joint point sequence and human skeleton sequence of the same person are predicted using the trained corresponding improved MS-G3D network, and the two prediction results and the abnormal behavior category of the person are combined to produce a two-stream fusion dataset, which is divided into a training set and a validation set; 5) Use the training set of the two-stream fusion dataset to train the two-stream fusion network, and use the corresponding validation set to verify the model accuracy to select the optimal model parameters; 6) For each frame of the video to be detected, the passenger skeleton is extracted and tracked to obtain the human joint point sequence and human bone sequence. The two sequences are respectively sent to the trained corresponding improved MS-G3D network for prediction. The two prediction results are sent to the trained two-stream fusion network for prediction. The prediction result is the final abnormal behavior category of the passenger.

2. The abnormal behavior recognition method based on dual-stream attention graph convolution according to claim 1 is characterized in that: In step 5), the dual-stream fusion network includes a channel importance extraction module and a stream importance extraction module; The channel importance extraction module includes a BN layer, a Concat layer, a fully connected layer, a ReLU activation function, and a Sigmoid activation function. First, the dual-stream inputs with the number of channels C are normalized using the BN layer, and then the Concat layer is used for feature concatenation. Then, the output with the number of channels 2C is obtained through a fully connected layer, a ReLU activation function, a fully connected layer, and a Sigmoid activation function, corresponding to the importance of each channel of the dual-stream input; The stream importance extraction module includes an average pooling layer, a maximum pooling layer, a Concat layer, a fully connected layer, a ReLU activation function and a Sigmoid activation function. First, the dual-stream inputs with the number of channels C are connected to the average pooling layer and the maximum pooling layer respectively, and then the Concat layer is used to perform feature concatenation to obtain a feature vector with a channel number of 4, and then the fully connected layer, the ReLU activation function, the fully connected layer, the ReLU activation function, the fully connected layer, and the Sigmoid activation function are passed to obtain an output with a channel number of 2, corresponding to the overall importance of different streams.

Citation Information

Patent Citations

  • Yellow River ice semantic segmentation method based on multi-attention mechanism double-flow fusion network

    CN111160311A

  • End-to-end behavior recognition method and system based on self-adaptive space-time attention mechanism

    CN111401177A