A behavior recognition method and system based on multi-modal and multi-level information fusion

By adopting a multimodal multi-level information fusion network model in behavior recognition, combining video and skeletal features, the problem of low accuracy in fall behavior recognition in the existing technology is solved, efficient and accurate fall behavior monitoring is achieved, and public safety is ensured.

CN116682172BActive Publication Date: 2025-05-27CHINA RAILWAY ERYUAN ENGINEERING GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310557432.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-17
Publication Date
2025-05-27
Estimated Expiration
2043-05-17

AI Technical Summary

Technical Problem

The prior art is difficult to identify fall behavior efficiently and with high accuracy, especially when the elderly fall, and there is a problem of low recognition robustness and low accuracy.

Method used

A behavior recognition method based on multimodal multi-level information fusion is adopted. By inputting video clips and human skeleton point sequences into multimodal multi-level information fusion network model, combining video feature extraction network, bone feature extraction network and intermediate layer feature fusion module, behavior categories are generated to determine whether there is a fall behavior.

Benefits of technology

Through the use of the intermediate layer feature fusion module, the monitoring of target behavior can be efficiently completed, improving the recognition accuracy and robustness of fall behavior, thereby ensuring public safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116682172B_ABST
    Figure CN116682172B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer technology, and specifically relates to a behavior recognition method and system based on multi-modal and multi-level information fusion. The method includes the following steps: slicing a target video to generate video segments, and generating an input behavior human body bone point sequence according to the video segments; inputting the video segments and the behavior human body bone point sequence into a trained multi-modal and multi-level information fusion network model to obtain a behavior category, and the behavior category is used to determine whether there is a fall behavior; the multi-modal and multi-level information fusion network model includes a video feature extraction network, a bone feature extraction network, and an intermediate layer feature fusion module. The present invention combines video and bone information through the intermediate layer feature fusion module, can efficiently complete the monitoring of target behaviors, and further enables relevant personnel to take corresponding measures to ensure public safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and relates to a behavior recognition method and system based on multi-modal and multi-level information fusion. Background Art

[0002] If a person's fall, especially an elderly person's fall, is not discovered in time, it may pose a threat to life and health. Currently, the discovery of fall behaviors mainly relies on manual detection. On the one hand, due to people's avoidance of harm, even if a fall behavior is discovered, people may turn a blind eye to these behaviors in order to avoid trouble. On the other hand, during specific time periods (such as late at night), there are few people, resulting in these behaviors not being discovered.

[0003] Behavior recognition is an important research issue in the field of computer vision research and is widely applied in many fields such as video surveillance, behavior analysis, and human-computer interaction. Traditional human body recognition mainly relies on RGB video sequences. However, due to the limitations of traditional algorithms in feature extraction, problems such as low algorithm robustness and low accuracy often occur. In recent years, with the development of artificial intelligence technology, the accuracy of behavior recognition based on deep learning of RGB video sequences has been greatly improved. However, there are often complex backgrounds in video scenes, which pose certain challenges to behavior recognition. With the development of high-precision depth sensors and human pose estimation algorithms, the skeletal joint points of the human body are easily obtained. The skeleton-based behavior recognition method based on deep learning can effectively avoid the influence of video backgrounds on behavior recognition. However, some background information is beneficial to behavior recognition, and skeleton-based behavior recognition completely ignores video background information, limiting the expressiveness of behavior recognition.

[0004] Aiming at the problem that it is difficult to efficiently and accurately recognize fall behaviors in the prior art, no effective solution has been proposed yet. Summary of the Invention

[0005] In order to overcome the defects existing in the above-mentioned prior art, the present invention provides a behavior recognition method and system based on multi-modal and multi-level information fusion.

[0006] In order to achieve the above object, the following technical solutions are provided:

[0007] A behavior recognition method based on multi-modal and multi-level information fusion includes the following steps:

[0008] Slice the target video to generate video segments, and generate an input sequence of skeletal points of the behavior human body according to the video segments; input the video segments and the sequence of skeletal points of the behavior human body into a trained multi-modal and multi-level information fusion network model to obtain a behavior category, and the behavior category is used to determine whether there is a fall behavior;

[0009] The multi-modal multi-level information fusion network model includes a video feature extraction network, a skeleton feature extraction network, and an intermediate layer feature fusion module. The video clip is input into the video feature extraction network, and the human body skeleton point sequence of the behavior is input into the skeleton feature extraction network. The video feature extraction network and the skeleton feature extraction network each include a number of intermediate layers, and the number of intermediate layers are sequentially linked. And after the video features output by the output layer of the video feature extraction network and the skeleton features output by the output layer of the skeleton feature extraction network are concatenated, they are sequentially input into a fully connected layer and a softmax layer to obtain the behavior category.

[0010] The intermediate layer feature fusion module is used to fuse the video features output by the current intermediate layer of the video feature extraction network and the skeleton features output by the current intermediate layer of the skeleton feature extraction network to obtain a video feature fusion weight and a skeleton feature fusion weight. The video feature fusion weight is multiplied by the video features output by the current intermediate layer of the video feature extraction network in the channel dimension and then input into the next layer, and the skeleton feature fusion weight is multiplied by the skeleton features output by the current intermediate layer in the skeleton feature extraction network in the channel dimension and then input into the next layer.

[0011] As a preferred solution of the present invention, the intermediate layer feature fusion module includes a compression layer, a fusion layer, a separation layer, and an excitation layer that are sequentially linked.

[0012] As a preferred solution of the present invention, the compression layer is represented by the formula:

[0013]

[0014]

[0015] Among them, h and w correspond to the height and width of specific pixels in the video features, t represents the specific number of frames in the video, c represents the specific number of channels in video feature A, c' represents the specific number of channels in skeleton feature B, v represents the specific skeleton point, S A , S B respectively represent the compressed video features and skeleton features, and R C refers to a c-order matrix.

[0016] As a preferred solution of the present invention, the fusion layer is represented by the following formula:

[0017]

[0018] C z =(C + C') / 4, where Z represents the fused feature, W represents the learning weight, S A , S B respectively represent the compressed video features and skeleton features, b represents the bias, C ZRepresents the number of feature channels after fusion, refers to matrix C Z of order

[0019] As a preferred embodiment of the present invention, the separation layer is represented by the following formula:

[0020]

[0021]

[0022] wherein, E A , E B respectively represent the separated video features and skeletal features, W A , W B respectively represent the deep learning weights of video features and skeletal features, b A , b B respectively represent the depth biases of video features and skeletal features, c represents the specific number of channels in video feature A, c' represents the specific number of channels in skeletal feature B, R C refers to a matrix of order c, R C refers to a matrix of order c;

[0023] The excitation layer is represented by the following formula:

[0024]

[0025]

[0026] wherein, A and B respectively represent the features of video feature A and skeletal feature B after intermediate feature fusion, σ(.) represents the sigmoid activation function, ⊙ represents multiplication in the channel dimension, R HWTC represents the video feature, R TVC ' represents the skeletal feature, H and W correspond to the height and width of specific pixels in the video feature, T represents the specific number of frames in the video, C represents the specific number of channels in video feature A, C' represents the specific number of channels in skeletal feature B, and V represents the specific skeletal points.

[0027] As a preferred embodiment of the present invention, the video feature extraction network adopts a 3D ResNet-18 network, and the skeletal feature extraction network adopts an ST-GCN network.

[0028] As a preferred embodiment of the present invention, the sequence of human body skeletal points of the behavior is enhanced and then input into the skeletal feature extraction network. The enhancement process includes rotation processing and shear processing. The rotation processing is used to imitate the transformation of the viewing angle of the acquisition device of the target video, and the shear processing is used to imitate the change of the body inclination when a person performs an action.

[0029] As a preferred embodiment of the present invention, the rotation process is represented by the following matrix:

[0030]

[0031]

[0032]

[0033] R = R z (γ)R y (β)R x (α);

[0034] wherein, R x , R y , R z respectively represent the rotation of the bone point around the x, y, and z axes, and α, β, and γ represent the rotation angles.

[0035] As a preferred embodiment of the present invention, the shear process is represented by the following matrix:

[0036] wherein represents the tilt factor from the subscript axis to the superscript axis.

[0037] Based on the same concept, a behavior recognition system based on multi-modal and multi-level information fusion is also proposed, including at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a behavior recognition method based on multi-modal and multi-level information fusion according to any one of the above.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] The present invention combines video and bone information through an intermediate layer feature fusion module, can efficiently complete the monitoring of target behaviors, and further enables relevant personnel to take corresponding measures to ensure public safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a flowchart of a fall behavior recognition method based on multi-modal and multi-level information fusion in Embodiment 1 of the present invention;

[0041] Figure 2 is a schematic diagram of extracting human bone joint points in Embodiment 1 of the present invention;

[0042] Figure 3 is a schematic diagram of bone data enhancement in Embodiment 1 of the present invention;

[0043] Figure 4 Schematic diagram of the intermediate layer feature fusion module in Embodiment 1 of the present invention. Detailed implementation manners

[0044] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation of the present invention.

[0045] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it may be a mechanical connection or an electrical connection, or it may be the communication inside two elements. It may be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.

[0046] Embodiment 1

[0047] As Figure 1 shown, taking the fall behavior recognition as an example, the present application provides a fall behavior recognition method based on multi-modal and multi-level information fusion, which specifically includes the following steps:

[0048] Collect a fall video dataset and preprocess the fall video dataset. In this embodiment, preferably but not limited to, collect the fall video dataset through an RGB camera, and divide the dataset into a training set and a test set after preprocessing it.

[0049] Construct a network model, which includes a video feature extraction network, a skeleton feature extraction network, and an intermediate layer feature fusion module, and train the constructed network model.

[0050] In this embodiment, a video feature extraction network is built based on 3D convolution. The feature extraction network selects the 3DResNet-18 network, and its structure includes: input layer → 3d res1 → 3d res2 → 3d res3 → 3d res4 → global pooing layer → fully connected layer. After completing the construction of the video feature extraction network, use the preprocessed fall video dataset to train the network and extract video features.

[0051] Generate a fall human skeleton point sequence according to the preprocessed fall video dataset.

[0052] In this embodiment, as Figure 2As shown, preferably but not limited to, a fall human body bone point extraction network is used to generate a corresponding bone point sequence for the video frames of the preprocessed fall video dataset. After the human body bone point sequence is generated, data augmentation processing is performed on it, such as Figure 3 As shown, the data augmentation processing techniques include Rotation and Shear, where Rotation mimics the transformation of the camera perspective and Shear mimics the change in the body inclination when a person performs an action. Among them, Rotation is represented by the following matrix:

[0053]

[0054]

[0055]

[0056] R = R z (γ)R y (β)R x (α)

[0057] Among them, R x , R y , R z respectively represent the rotation of the bone point around the x, y, and z axes, and α, β, and γ represent the rotation angles;

[0058] Shear is represented by the following matrix:

[0059] Among them represents the inclination factor from the subscript axis to the superscript axis. For example represents the inclination factor from the x-axis to the y-axis.

[0060] Then, a bone feature extraction network is built, and the fall human body bone point sequence is used to train the bone feature extraction network to extract bone features. In this embodiment, preferably but not limited to, a bone feature extraction network is built based on graph convolution. The bone feature extraction network selects the ST-GCN network, and its structure includes: input layer → spatial graph convolution layer → temporal convolution layer → spatial graph convolution layer → temporal convolution layer →... → global pooing layer → fully connected layer. After the bone feature extraction network is built, the above human body bone sequence is used to train the network.

[0061] Build an intermediate layer feature fusion module, and fuse the intermediate layer video features obtained during the training of the video feature extraction network and the intermediate layer bone features obtained during the training of the bone feature extraction network in the intermediate layer feature fusion module.

[0062] There are N intermediate feature fusion modules, where N is a positive integer. In this embodiment, N is preferably but not limited to 2. Since the network model needs to be trained through multiple iterations during training, the intermediate feature fusion module fuses video features and skeletal features to obtain fusion weights, which are respectively multiplied by the input video features and input skeletal features in the channel dimension to obtain video features and skeletal features that fuse different modality information. The video features and skeletal features that fuse different modality information are used as the intermediate layer video features and intermediate layer skeletal features to be fused in another intermediate feature fusion module and output video features and skeletal features that fuse different modality information.

[0063] As Figure 4 shown, the structure of the intermediate feature fusion module includes: compression layer → fusion layer → separation layer → excitation layer.

[0064] The intermediate feature fusion module takes the intermediate layer video feature A ∈ R HWTC and the intermediate layer skeletal feature B ∈ R TVC ' as inputs, where H and W correspond to the height and width of specific pixels in the video feature, T represents the specific number of frames in the video, C represents the specific number of channels in the video feature A, C' represents the specific number of channels in the skeletal feature B, and V represents the specific skeletal points.

[0065] The compression layer is represented by the formula:

[0066]

[0067] where S A , S B represents the compressed feature, R * refers to the *-order matrix, and R *×* represents the *×* -order matrix;

[0068] The fusion layer can be represented by the following formula:

[0069]

[0070] C Z = (C + C') / 4, where Z represents the fused feature, W' represents the learned weight, b represents the bias, and C Z represents the number of channels of the fused feature;

[0071] The separation layer can be represented by the following formula:

[0072]

[0073]

[0074] where E A , E Brespectively represent the separated features, W A , W B represents the deep learning weight, b A , b B represents the deep bias, W A , W B , b A , b B are all obtained by machine learning;

[0075] The excitation layer can be represented by the following formula:

[0076] A = 2 * σ(E A ) ⊙ A; A ∈ R HWTC ;

[0077] B = 2 * σ(E B ) ⊙ A; B ∈ R TVC ';

[0078] wherein, A and B respectively represent the features of feature A and B after passing through the intermediate feature fusion layer, σ(.) represents the sigmoid activation function, ⊙ represents multiplication in the channel dimension, R HWTC represents video features, R TVC ' represents skeleton features.

[0079] Finally, the video features finally output by the video feature extraction network and the skeleton features finally output by the skeleton feature extraction network are concatenated, and the behavior category is obtained by using fully connected and softmax operations.

[0080] In practical applications, the monitored video information is preprocessed and then input into the above network model for recognition to determine whether there is a fall behavior.

[0081] The recognition of other target behaviors can refer to the operation of this fall behavior recognition, and all belong to the protection scope of this application. Moreover, multiple target behaviors can be trained during the training stage to achieve the recognition of multiple target behaviors.

[0082] This application also proposes a behavior recognition system, including a processing module and a storage module connected to the processing module, and the two communicate with each other. The storage module is used to store at least one executable instruction, and the executable instruction enables the processing module to perform operations corresponding to the above-mentioned behavior recognition method based on multi-modal and multi-level information fusion.

[0083] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0084] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A behavior recognition method based on multi-modal and multi-level information fusion, characterized in that, it includes the following steps: Slice the target video to generate video segments, and generate an input behavioral human body bone point sequence according to the video segments; input the video segments and the behavioral human body bone point sequence into a trained multi-modal and multi-level information fusion network model to obtain a behavior category, and the behavior category is used to judge whether there is a fall behavior; The multi-modal and multi-level information fusion network model includes a video feature extraction network, a bone feature extraction network and an intermediate layer feature fusion module. The video segments are input into the video feature extraction network, and the behavioral human body bone point sequence is input into the bone feature extraction network; the video feature extraction network and the bone feature extraction network each include several intermediate layers, and the several intermediate layers are connected in sequence; and after the video features output by the output layer of the video feature extraction network and the bone features output by the output layer of the bone feature extraction network are spliced, they are sequentially input into a fully connected layer and a softmax layer to obtain a behavior category; The intermediate layer feature fusion module is used to fuse the video features output by the current intermediate layer of the video feature extraction network and the bone features output by the current intermediate layer of the bone feature extraction network to obtain a video feature fusion weight and a bone feature fusion weight. Multiply the video feature fusion weight with the video features output by the current intermediate layer of the video feature extraction network through channel multiplication and then input it into the next layer. Multiply the bone feature fusion weight with the bone features output by the current intermediate layer in the bone feature extraction network through channel multiplication and then input it into the next layer.

2. A behavior recognition method based on multi-modal and multi-level information fusion according to claim 1, characterized in that, the intermediate layer feature fusion module includes a compression layer, a fusion layer, a separation layer and an excitation layer connected in sequence.

3. A behavior recognition method based on multi-modal and multi-level information fusion according to claim 2, characterized in that, the compression layer is represented by the formula: ; ; Among them, H and W correspond to the height and width of specific pixels in the video features, T represents the specific number of frames in the video, C represents the specific number of channels in video feature A, C' represents the specific number of channels in skeleton feature B, and V represents the specific skeleton points. respectively represent the compressed video features and skeleton features. refers to a matrix of order C.

4. A behavior recognition method based on multi-modal and multi-level information fusion according to claim 3, characterized in that, the fusion layer is represented by the following formula: ; , where Z represents the fused feature, W’ represents the learned weight, respectively represent the compressed video feature and skeleton feature, b represents the bias, represents the number of channels of the fused feature, refers to an order matrix.

5. A behavior recognition method based on multi-modal and multi-level information fusion according to claim 4, characterized in that, the separation layer is represented by the following formula: ; ; Among them, respectively represent the separated video features and skeleton features, respectively represent the deep learning weights of the video features and the skeleton features, respectively represent the depth biases of the video features and the skeleton features. c represents the specific number of channels in video feature A, and c' represents the specific number of channels in skeleton feature B, refers to a matrix of order C; The excitation layer is represented by the following formula: ; ; Among them, respectively represent the features after the intermediate feature fusion of video feature A and skeleton feature B, represents the sigmoid activation function, represents multiplication in the channel dimension, represents the video feature, represents the skeleton feature. H and W correspond to the height and width of specific pixels in the video feature, T represents the specific number of frames in the video, C represents the specific number of channels in video feature A, C' represents the specific number of channels in skeleton feature B, and V represents the specific skeleton points.

6. A behavior recognition method based on multi-modal and multi-level information fusion according to claim 1, characterized in that, the video feature extraction network adopts a 3D ResNet-18 network, and the bone feature extraction network adopts an ST-GCN network.

7. A behavior recognition method based on multi-modal and multi-level information fusion according to claim 1, characterized in that, the behavioral human body bone point sequence is enhanced and then input into the bone feature extraction network. The enhancement processing includes rotation processing and shear processing. The rotation processing is used to imitate the transformation of the acquisition device's perspective of the target video, and the shear processing is used to imitate the change of the body inclination when a person performs an action.

8. A behavior recognition method based on multi-modal and multi-level information fusion according to claim 7, characterized in that, the rotation processing is represented by the following matrix: ; ; ; ; Among them, respectively represent the rotation of the bone point around the x, y, and z axes, represents the rotation angle.

9. A behavior recognition method based on multi-modal and multi-level information fusion according to claim 7, characterized in that, the shearing processing is represented by the following matrix: , where , represents the inclination factor from the subscript axis to the superscript axis.

10. A behavior recognition system based on multi-modal and multi-level information fusion, characterized in that, it includes at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a behavior recognition method based on multi-modal and multi-level information fusion according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Human body posture correction device based on Kinect sensor

    CN104157107A

  • Human body behavior recognition method based on RGB video and skeleton sequence

    CN111967379A