A Behavior Recognition Method and System Based on a Multimodal Attention Fusion Network
By adopting a multimodal attention fusion network in the activity recognition of elderly people and fusing RGB and skeletal features, the identification problem of individual actions and human-client interactions in the activity recognition of elderly people is solved, and higher behavior recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111458271.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-02
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-02
AI Technical Summary
The identification of activities for elderly people is a challenging task because there are individual actions and human-client interactions in activities of elderly people. Many people-client interactions are local, the amplitude of activities of elderly people is not obvious, and some activities of elderly people are similar. How to capture and fuse discriminant information in RGB and skeletal modality is crucial for the modeling of activities for elderly people.
The behavior recognition method based on the multimodal attention fusion network is adopted, and the RGB features and Shift-GCN network are extracted through the ResNeXt101 network, and the bone features are extracted, and their input mode fusion network and channel fusion network are fused. Finally, the multimodal loss optimization network is used to output behavior recognition results.
It improves the accuracy of behavior recognition, effectively detects human activity categories, and improves detection accuracy.
Smart Images

Figure CN114170683B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of action recognition, and in particular to a behavior recognition method and system based on a multimodal attention fusion network. Background Art
[0002] With the continuous increase of the social population, the problem of population aging is becoming more and more serious. Coupled with the long-term absence of young people, the proportion of empty-nest elderly in many countries has risen sharply. As the physical fitness and mobility of the elderly begin to decline, some dangers are prone to occur in life. In this context, in recent years, some institutions are trying to provide intelligent monitoring of the daily activities of the elderly by applying some advanced technologies in the fields of artificial intelligence and computer vision. Compared with normal action recognition, activity recognition for the elderly is a more challenging task, because there are individual actions and human-guest interactions in the activities of the elderly, many of which are local; the amplitude of many elderly activities is not obvious; and the actions of some elderly activities are particularly similar. How to capture and fuse the discriminative information in RGB and skeletal modalities is crucial for modeling activities of the elderly.
[0003] As the key to intelligent monitoring, elderly activity recognition has attracted more and more attention. Researchers have also proposed skeleton-based and RGB-based action recognition methods, which have good modeling capabilities. However, the RGB-based action recognition method mainly obtains the motion information of the action from the RGB video, which is interfered by the background information to a certain extent. The skeleton-based action recognition method faces the challenge of identifying actions with similar postures. Summary of the invention
[0004] The purpose of the present invention is to provide a behavior recognition method and system based on a multimodal attention fusion network, which can improve the accuracy of behavior recognition.
[0005] To achieve the above object, the present invention provides the following solutions:
[0006] A behavior recognition method based on a multimodal attention fusion network, comprising:
[0007] The ResNeXt101 network in the multimodal attention fusion network is used to extract RGB features in the video to be identified; and the Shift-GCN network in the multimodal attention fusion network is used to extract bone features; the multimodal attention fusion network takes the video to be identified as input and takes the behavior recognition result as output;
[0008] The modal fusion network in the multimodal attention fusion network is used to perform modal fusion on the RGB features and the skeleton features to determine the features after modal fusion; the modal fusion network includes: a modal dimension expansion network, a modal dimension compression network and a modal dimension activation network;
[0009] The channel fusion network in the multi-modal attention fusion network is used to perform channel fusion on the features after modal fusion to determine the features after channel fusion; the channel fusion network includes: a channel dimension expansion network, a channel dimension compression network, and a channel dimension activation network;
[0010] Integrate the losses of RGB features, skeleton features, and the features after channel fusion to determine the multi-modal loss, and use the multi-modal loss to optimize the multi-modal attention fusion network to output the behavior recognition result.
[0011] Optionally, before using the ResNeXt101 network in the multi-modal attention fusion network to extract RGB features in the video to be recognized; and using the Shift-GCN network in the multi-modal attention fusion network to extract skeleton features, it further includes:
[0012] Divide the video to be recognized into multiple segments;
[0013] Randomly sample each segment to determine video frames.
[0014] Optionally, using the ResNeXt101 network in the multi-modal attention fusion network to extract RGB features in the video to be recognized; and using the Shift-GCN network in the multi-modal attention fusion network to extract skeleton features specifically includes:
[0015] Determine the RGB frame sequence and the sequence of skeletons according to the video frames;
[0016] Use the ResNeXt101 network to extract the RGB frame sequence to determine RGB features;
[0017] Use the Shift-GCN network to extract the sequence of skeletons to determine skeleton features.
[0018] Optionally, before using the modal fusion network in the multi-modal attention fusion network to perform modal fusion on the RGB features and skeleton features to determine the features after modal fusion, it further includes:
[0019] Use two multi-layer perceptrons to unify the RGB features and skeleton features;
[0020] Transpose the unified features.
[0021] Optionally, using the modal fusion network in the multi-modal attention fusion network to perform modal fusion on the RGB features and skeleton features to determine the features after modal fusion specifically includes:
[0022] Expand the features of the transposed features using a modal dimension expansion network;
[0023] Compress the features of the features after modal dimension expansion using a modal dimension compression network;
[0024] Activate the features of the features after modal dimension compression using a modal dimension activation network.
[0025] Optionally, the channel fusion network in the multi-modal attention fusion network is used to perform channel fusion on the features after modal fusion to determine the features after channel fusion. Specifically, it includes:
[0026] Expand the features of the features after modal fusion using a channel dimension expansion network;
[0027] Compress the features of the features after channel dimension expansion using a channel dimension compression network;
[0028] Activate the features of the features after channel dimension compression using a channel dimension activation network.
[0029] Optionally, the losses of the RGB features, the skeletal features, and the features after channel fusion are integrated to determine the multi-modal loss, and the multi-modal loss is used to optimize the multi-modal attention fusion network to output the behavior recognition result. Specifically, it includes:
[0030] Use the formula Determine the multi-modal loss;
[0031] Among them, is the multi-modal loss, is the RGB modal loss, is the skeletal modal loss, is the fusion modal loss, and both α and β are weights.
[0032] A behavior recognition system based on a multi-modal attention fusion network, including:
[0033] A feature extraction module, which is used to extract RGB features in the video to be recognized using the ResNeXt101 network in the multi-modal attention fusion network; and extract skeletal features using the Shift-GCN network in the multi-modal attention fusion network; the multi-modal attention fusion network takes the video to be recognized as input and the behavior recognition result as output;
[0034] A modal fusion module, which is used to perform modal fusion on the RGB features and skeletal features using the modal fusion network in the multi-modal attention fusion network to determine the features after modal fusion; the modal fusion network includes: a modal dimension expansion network, a modal dimension compression network, and a modal dimension activation network;
[0035] A channel fusion module, configured to perform channel fusion on the features after modality fusion by using a channel fusion network in a multi-modal attention fusion network to determine the features after channel fusion; the channel fusion network includes: a channel dimension expansion network, a channel dimension compression network, and a channel dimension activation network;
[0036] A multi-modal loss optimization module, configured to integrate the losses of RGB features, skeleton features, and the features after channel fusion to determine a multi-modal loss, and use the multi-modal loss to optimize the multi-modal attention fusion network to output a behavior recognition result.
[0037] Optionally, the multi-modal loss optimization module specifically includes:
[0038] A multi-modal loss determination unit, configured to use the formula to determine the multi-modal loss;
[0039] where is the multi-modal loss, is the RGB modality loss, is the skeleton modality loss, is the fused modality loss, and both α and β are weights.
[0040] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0041] A behavior recognition method and system based on a multi-modal attention fusion network provided by the present invention inputs the extracted RGB features and skeleton features into a modality fusion network, expands, compresses, and activates the features respectively from the modality dimension; inputs the features after modality fusion into a channel fusion network to fuse the features respectively from the channel dimension; constructs three sub-losses by using three types of features, namely RGB features, skeleton features, and fused multi-modal features, integrates them into a new multi-modal loss, and optimizes the entire network. The present invention can effectively detect human activity categories and has a high detection accuracy. Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0043] Figure 1 It is a schematic flowchart of a behavior recognition method based on a multi-modal attention fusion network provided by the present invention;
[0044] Figure 2Schematic diagram of a behavior recognition system based on a multi-modal attention fusion network provided by the present invention. Specific embodiments
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0046] The purpose of the present invention is to provide a behavior recognition method and system based on a multi-modal attention fusion network, which can improve the accuracy of behavior recognition.
[0047] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0048] Figure 1 Schematic diagram of the process of a behavior recognition method based on a multi-modal attention fusion network provided by the present invention. As Figure 1 shown, a behavior recognition method based on a multi-modal attention fusion network provided by the present invention includes:
[0049] S101, using the ResNeXt101 network in the multi-modal attention fusion network to extract the RGB features in the video to be recognized; and using the Shift-GCN network in the multi-modal attention fusion network to extract the skeletal features; the multi-modal attention fusion network takes the video to be recognized as the input and the behavior recognition result as the output;
[0050] Before S101, it further includes:
[0051] Dividing the video to be recognized into multiple segments;
[0052] Randomly sampling each segment to determine video frames.
[0053] S101 specifically includes:
[0054] Determining the RGB frame sequence and the skeleton sequence according to the video frames;
[0055] Using the ResNeXt101 network to extract the RGB frame sequence to determine the RGB features;
[0056] Using the Shift-GCN network to extract the skeleton sequence to determine the skeletal features.
[0057] As a specific embodiment, the video is divided into T segments, and one frame is randomly selected from each segment to obtain the RGB frame sequence {r i | i = 1, 2, 3,..., T} of the video r. The sequence s i | i = 1, 2, 3,..., T of the skeleton s is obtained using the same method.
[0058] The RGB frames are input into the ResNeXt101 network, the model parameters are fine-tuned, and then the parameters are fixed to extract the RGB feature f r = ResNeXt101(r), where
[0059] The skeleton frames are input into the Shift-GCN network, the model parameters are fine-tuned, and then the parameters are fixed to extract the skeleton feature f s = Shift-GCN(s) where
[0060] S102. The modality fusion network in the multi-modal attention fusion network is used to perform modality fusion on the RGB feature and the skeleton feature to determine the feature after modality fusion; the modality fusion network includes: a modality dimension expansion network, a modality dimension compression network, and a modality dimension activation network;
[0061] To connect features with different feature scales, before S102, it further includes:
[0062] Using two Multi-Layer Perceptrons (MLPs) to unify the RGB feature and the skeleton feature;
[0063] Transposing the unified feature to determine the feature f.
[0064] f = Concat(H r (f r ), H s (f s ));
[0065] h m = M-Net(Trans(f));
[0066] where the feature f ∈ R d×n , H r and H s represent MLPs, and Concat(·) represents the concatenation operation of modality dimensions.
[0067] S102 specifically includes:
[0068] Using the modality dimension expansion network to expand the transposed feature; that is, using the formula f ml=Conv1(Conv2(Conv3(f))) expands the spatial information from the local view to achieve the interaction of feature modal information;
[0069] Among them d m is the spatial dimension after the convolution transformation. f ml can also be regarded as a single-modal feature set
[0070] Use the modal dimension compression network to compress the features after the modal dimension expansion;
[0071] The feature f ml is averaged from a global perspective to obtain the compressed feature f of the modal dimension mlg ∈R m×l (m < n×d), that is:
[0072]
[0073] Use the modal dimension activation network to activate the features after the modal dimension compression.
[0074] The feature f mlg is input into several fully connected networks and then the attention W between modalities is learned through the Sigmoid function m . That is:
[0075] W m =σ(F3(Relu(F4(f mlg ))))
[0076] where σ is the activation function.
[0077] Apply the attention to the original features to obtain the features after modal fusion:
[0078]
[0079] S103. Use the channel fusion network in the multi-modal attention fusion network to perform channel fusion on the features after the modal fusion, and determine the features after the channel fusion; the channel fusion network includes: a channel dimension expansion network, a channel dimension compression network, and a channel dimension activation network;
[0080] S103 specifically includes:
[0081] Use the channel dimension expansion network to expand the features after the modal fusion; that is, use the formula h cg =Conv4(h m ) to achieve the interaction of feature channel information and obtain the channel dimension expansion features;
[0082] Among them
[0083] Use the channel dimension compression network to compress the features after channel dimension expansion;
[0084] Feature h cg Perform average pooling from a global perspective to obtain the compressed features of the channel dimension. That is:
[0085]
[0086] where f cg ∈R d×1 , and can also be regarded as a single-channel feature set
[0087] Use the channel dimension activation network to activate the features after channel dimension compression.
[0088] The obtained feature f cg is input into several fully connected layers, and finally through the Sigmoid function, the channel attention W c is obtained. That is:
[0089] W c =σ(F5(Relu(F6(f cg )))));
[0090] Apply the attention to the feature h mc , and obtain the feature h m after channel fusion. That is:
[0091]
[0092] S104, integrate the losses of RGB features, skeletal features, and the features after channel fusion to determine the multi-modal loss, and use the multi-modal loss to optimize the multi-modal attention fusion network and output the behavior recognition result. That is, maintain the consistency between single-modal features and fused multi-modal features.
[0093] S104 specifically includes:
[0094] Use the formula to determine the multi-modal loss; use the multi-modal loss to additionally measure the difference between the single-modal prediction loss and the fused multi-modal prediction loss.
[0095] where, is the multi-modal loss, is the RGB modal loss, is the skeletal modal loss, is the fused modal loss, and both α and β are weights.
[0096] Figure 2The following is a schematic structural diagram of a behavior recognition system based on a multi-modal attention fusion network provided by the present invention. As Figure 2 shown, a behavior recognition system based on a multi-modal attention fusion network provided by the present invention includes:
[0097] A feature extraction module 201, configured to extract RGB features in a video to be recognized by using a ResNeXt101 network in the multi-modal attention fusion network; and extract skeleton features by using a Shift-GCN network in the multi-modal attention fusion network; the multi-modal attention fusion network takes the video to be recognized as an input and outputs a behavior recognition result;
[0098] A modality fusion module 202, configured to perform modality fusion on the RGB features and the skeleton features by using a modality fusion network in the multi-modal attention fusion network to determine the features after modality fusion; the modality fusion network includes: a modality dimension expansion network, a modality dimension compression network, and a modality dimension activation network;
[0099] A channel fusion module 203, configured to perform channel fusion on the features after modality fusion by using a channel fusion network in the multi-modal attention fusion network to determine the features after channel fusion; the channel fusion network includes: a channel dimension expansion network, a channel dimension compression network, and a channel dimension activation network;
[0100] A multi-modal loss optimization module 204, configured to integrate the losses of the RGB features, the skeleton features, and the features after channel fusion to determine a multi-modal loss, and use the multi-modal loss to optimize the multi-modal attention fusion network and output a behavior recognition result.
[0101] The multi-modal loss optimization module 204 specifically includes:
[0102] A multi-modal loss determination unit, configured to use the formula to determine the multi-modal loss;
[0103] where is the multi-modal loss, is the RGB modality loss, is the skeleton modality loss, is the fused modality loss, and both α and β are weights..
[0104] In the present specification, each embodiment is described in a progressive manner. The key point of each embodiment is to describe the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0105] In this article, specific examples are used to elaborate on the principles and implementation modes of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation modes and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A behavior recognition method based on a multi-modal attention fusion network, characterized in that Including: Extracting RGB features in the video to be recognized by using the ResNeXt101 network in the multi-modal attention fusion network; And extracting skeletal features by using the Shift-GCN network in the multi-modal attention fusion network; the multi-modal attention fusion network takes the video to be recognized as the input and outputs the behavior recognition result; Performing modal fusion on the RGB features and the skeletal features by using the modal fusion network in the multi-modal attention fusion network to determine the features after modal fusion; the modal fusion network includes: a modal dimension expansion network, a modal dimension compression network, and a modal dimension activation network; Performing channel fusion on the features after modal fusion by using the channel fusion network in the multi-modal attention fusion network to determine the features after channel fusion; the channel fusion network includes: a channel dimension expansion network, a channel dimension compression network, and a channel dimension activation network; After modal fusion, integrating the losses of the RGB features, the losses of the skeletal features, and the losses of the features after channel fusion to determine the multi-modal loss, and optimizing the multi-modal attention fusion network by using the multi-modal loss to output the behavior recognition result.
2. The behavioral recognition method based on a multi-modal attention fusion network according to claim 1, wherein, The step of extracting RGB features in the video to be recognized by using the ResNeXt101 network in the multi-modal attention fusion network; And the step of extracting skeletal features by using the Shift-GCN network in the multi-modal attention fusion network are preceded by: Dividing the video to be recognized into multiple segments; Performing random sampling on each segment to determine video frames.
3. The behavior recognition method based on a multi-modal attention fusion network according to claim 2, wherein, The step of extracting RGB features in the video to be recognized by using the ResNeXt101 network in the multi-modal attention fusion network; And the step of extracting skeletal features by using the Shift-GCN network in the multi-modal attention fusion network specifically includes: Determining an RGB frame sequence and a skeleton sequence according to the video frames; Using the ResNeXt101 network to extract the RGB frame sequence to determine RGB features; Using the Shift-GCN network to extract the skeleton sequence to determine skeletal features.
4. The behavior recognition method based on a multi-modal attention fusion network according to claim 1, characterized in that, Before the step of performing modal fusion on the RGB features and the skeletal features by using the modal fusion network in the multi-modal attention fusion network to determine the features after modal fusion, it further includes: Using two multi-layer perceptrons to unify the RGB features and the skeletal features; Transposing the unified features.
5. The behavior recognition method based on a multi-modal attention fusion network according to claim 4, wherein The step of performing modal fusion on the RGB features and the skeletal features by using the modal fusion network in the multi-modal attention fusion network to determine the features after modal fusion specifically includes: Using the modal dimension expansion network to expand the transposed features; Using the modal dimension compression network to compress the features after modal dimension expansion; Using the modal dimension activation network to activate the features after modal dimension compression.
6. The behavioral recognition method based on a multi-modal attention fusion network according to claim 1, characterized in that The step of performing channel fusion on the features after modal fusion by using the channel fusion network in the multi-modal attention fusion network to determine the features after channel fusion specifically includes: Using the channel dimension expansion network to expand the features after modal fusion; Use the channel dimension compression network to compress the features after channel dimension expansion; Use the channel dimension activation network to activate the features after channel dimension compression.
7. A behavior recognition method based on a multi-modal attention fusion network according to claim 1, characterized in that, Integrate the losses of the RGB features, the skeletal features, and the features after channel fusion to determine the multimodal loss, and use the multimodal loss to optimize the multimodal attention fusion network to output the behavior recognition result, specifically including: Use the formula to determine the multi-modal loss; Among them, is the loss of the multi-modal, is the loss of the RGB modality, is the loss of the skeletal modality, is the loss of the fusion modality, where α and β are both weights.
8. A behavior recognition system based on a multi-modal attention fusion network, characterized in that, Including: A feature extraction module for extracting RGB features in the video to be recognized using the ResNeXt101 network in the multimodal attention fusion network; And use the Shift-GCN network in the multimodal attention fusion network to extract skeletal features; the multimodal attention fusion network takes the video to be recognized as input and outputs the behavior recognition result; A modality fusion module for using the modality fusion network in the multimodal attention fusion network to perform modality fusion on the RGB features and the skeletal features to determine the features after modality fusion; The modality fusion network includes: a modality dimension expansion network, a modality dimension compression network, and a modality dimension activation network; A channel fusion module for using the channel fusion network in the multimodal attention fusion network to perform channel fusion on the features after modality fusion to determine the features after channel fusion; the channel fusion network includes: a channel dimension expansion network, a channel dimension compression network, and a channel dimension activation network; A multimodal loss optimization module for integrating the losses of the RGB features, the skeletal features, and the features after channel fusion after modality fusion to determine the multimodal loss, and using the multimodal loss to optimize the multimodal attention fusion network to output the behavior recognition result.
9. The behavior recognition system based on a multi-modal attention fusion network according to claim 8, wherein, The multimodal loss optimization module specifically includes: A multi-modal loss determination unit for determining a multi-modal loss using the formula to determine the multi-modal loss; Among them, is the loss of the multi-modal, is the loss of the RGB modality, is the loss of the skeleton modality, is the loss of the fusion modality, and both α and β are weights.
Citation Information
Patent Citations
An image gesture action online detection and recognition method based on deep learning
CN109886225A
Behavior recognition method, device and system based on skeleton and RGB frame fusion
CN112906604A