Human body behavior recognition method and device based on multi-modal feature fusion, terminal equipment and storage medium
By extracting key points and fusion of multimodal features of the target video, the timing and optical flow feature matrix are generated, and the inaccurate behavior recognition problem caused by ignoring environmental changes in the prior art is solved, and a higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510633270.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-22
AI Technical Summary
Existing human behavior recognition methods are difficult to accurately identify human behavior, especially when ignoring environmental spatial changes.
By extracting key points on the target video, a key point feature sequence is generated, and combining timing feature extraction and optical flow feature extraction, a timing feature matrix, spatial feature matrix and optical flow feature matrix are constructed, multimodal feature fusion is performed, and finally the behavior recognition model is input for behavior recognition.
The recognition accuracy of the behavior recognition model is improved, so that it can pay attention to the action changes and environmental spatial changes of the target person at the same time.
Smart Images

Figure CN120526480A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a method, apparatus, terminal device and storage medium for multimodal feature fusion human behavior recognition. Background Art
[0002] In recent years, with the continuous development of computer technology, computer vision has become one of the most active and popular disciplines in the computer field. Computer vision uses cameras and computer equipment to simulate human visual functions, performing operations such as identification, tracking, and measurement of target objects, forming machine vision, and further conducting image processing and analysis based on this. Human behavior recognition, as an emerging research direction in the field of computer vision, has attracted widespread attention and in-depth research. This research direction focuses on videos containing human movement, using computer technology and graphics and image processing methods to extract feature information from the videos, and then accurately determine the action categories or behavior patterns of human activities in the videos.
[0003] However, traditional behavior recognition methods based on feature engineering usually extract and identify static features by intercepting video frames. This method, which only applies to frame-level features, is difficult to capture the dynamic changes of human movements. Secondly, traditional behavior recognition methods only focus on human movements when extracting features, ignoring the dynamic changes of the environmental space, and easily miss important feature information, resulting in recognition errors. For example, if a person in a video squats to pick up an object, if the model only focuses on the human movements in the static video frame, it will only be identified as "squatting" or "standing", while combined with environmental changes (changes in the position of the object), it can be accurately identified as "picking up". Therefore, current human behavior recognition methods have difficulty in accurately identifying human action behaviors. Summary of the Invention
[0004] The present invention provides a multimodal feature fusion human behavior recognition method, device, terminal device and storage medium, which can solve the defect that current human behavior recognition methods are difficult to accurately identify human action behaviors.
[0005] An embodiment of the present invention provides a method for human behavior recognition based on multimodal feature fusion, comprising: obtaining a target video to be recognized;
[0006] Extracting key points from the target video, identifying key point features of the target person in each video frame of the target video, and generating a key point feature sequence;
[0007] Extracting time series features from the key point feature sequence, identifying dynamic change features of the key point features of the human body according to the time sequence, and generating a time series feature matrix;
[0008] Performing spatial feature extraction and optical flow feature extraction on the target video, identifying the spatial features and optical flow features of each of the video frames, and constructing a spatial feature matrix and an optical flow feature matrix;
[0009] Fusing the temporal feature matrix, the spatial feature matrix, and the optical flow feature matrix of each video frame to generate a fused feature matrix of the video frame;
[0010] The fused feature matrix is input into a preset behavior recognition model, so that the behavior recognition model recognizes and outputs the target behavior of the target person according to the fused feature matrix.
[0011] Furthermore, extracting key points from the target video, identifying key point features of the target person in each video frame of the target video, and generating a key point feature sequence includes:
[0012] Splitting the target video into several video frames;
[0013] Performing human skeleton detection on the video frame to identify facial features and body joints of the target person in the video frame as initial key points;
[0014] Normalizing the size of the target person according to the initial key points to generate a normalization factor;
[0015] Normalizing the initial key points according to the normalization factor to generate corresponding target key points;
[0016] Traversing the video frames, calculating the angles and distances between each target key point in the currently traversed video frame and other target key points as key point sub-features of the currently traversed video frame;
[0017] The target key points and key point sub-features of each video frame are used as the human body key point features of the target person in each video frame;
[0018] The key point features of the human body of several video frames are arranged according to the time sequence to generate the key point feature sequence.
[0019] Furthermore, the extracting of time series features from the key point feature sequence, identifying the dynamic change features of the key point features of the human body according to the time sequence, and generating a time series feature matrix includes:
[0020] The key point feature sequence is input into a preset time series feature extraction model, so that the forward LSTM sub-model of the time series feature extraction model processes the human key point features in the key point feature sequence in a positive order to generate forward time series information of each human key point feature, and the backward LSTM sub-model of the time series feature extraction model processes the human key point features in the key point feature sequence in a reverse order to generate backward time series information of each human key point feature, and then the forward time series information and the backward time series information of each human key point feature are fused to generate a dynamic change feature of each human key point feature;
[0021] The time series feature matrix is constructed according to the dynamic change characteristics of each of the key point features of the human body.
[0022] Furthermore, performing spatial feature extraction and optical flow feature extraction on the target video, identifying the spatial features and optical flow features of each video frame, and constructing a spatial feature matrix and an optical flow feature matrix includes:
[0023] Inputting several video frames of the target video into a spatial feature extraction model so that the spatial feature extraction model recognizes and outputs image edge features, texture features, and local features of the video frames;
[0024] Using the image edge features, texture features, and local features of each of the video frames as spatial features;
[0025] The video frames are traversed, and the moving direction and moving speed of corresponding pixels between the currently traversed video frame and the next video frame are calculated, and the moving direction and moving speed are used as the optical flow features of the currently traversed video frame.
[0026] Furthermore, the fusion feature matrix is:
[0027] X t =[H t ; F t ; V t ];
[0028] Among them, X t is the fusion feature matrix of the t-th video frame, H t is the temporal feature matrix of the t-th video frame, F t is the spatial feature matrix of the t-th video frame, V t is the optical flow feature matrix of the t-th video frame.
[0029] Furthermore, the behavior recognition model identifies and outputs the target behavior of the target person according to the fusion feature matrix, including:
[0030] Using a global attention mechanism, the temporal feature matrix, the spatial feature matrix, and the optical flow feature matrix in each of the fused feature matrices are weighted to generate a weighted feature matrix;
[0031] Using the Softmax function, based on the weighted feature matrix, the probability distribution of the target person in a plurality of preset behaviors is calculated;
[0032] The behavior corresponding to the maximum probability in the probability distribution is used as the target behavior of the target person.
[0033] Another embodiment of the present invention further provides a multimodal feature fusion human behavior recognition device, comprising: a video acquisition module, for acquiring a target video to be recognized;
[0034] A human feature extraction module is used to extract key points from the target video, identify key point features of the target person in each video frame of the target video, and generate a key point feature sequence;
[0035] A time series feature extraction module is used to extract time series features from the key point feature sequence, identify the dynamic change characteristics of the key point features of the human body according to the time sequence, and generate a time series feature matrix;
[0036] An environmental feature extraction module is used to extract spatial features and optical flow features of the target video, identify the spatial features and optical flow features of each video frame, and construct a spatial feature matrix and an optical flow feature matrix;
[0037] a feature fusion module, configured to fuse the temporal feature matrix, the spatial feature matrix, and the optical flow feature matrix of each video frame to generate a fused feature matrix of the video frame;
[0038] The behavior recognition module is used to input the fused feature matrix into a preset behavior recognition model, so that the behavior recognition model recognizes and outputs the target behavior of the target person according to the fused feature matrix.
[0039] Furthermore, the human feature extraction module extracts key points from the target video, identifies key point features of the target person in each video frame of the target video, and generates a key point feature sequence, including:
[0040] Splitting the target video into several video frames;
[0041] Performing human skeleton detection on the video frame to identify facial features and body joints of the target person in the video frame as initial key points;
[0042] Normalizing the size of the target person according to the initial key points to generate a normalization factor;
[0043] Normalizing the initial key points according to the normalization factor to generate corresponding target key points;
[0044] Traversing the video frames, calculating the angles and distances between each target key point in the currently traversed video frame and other target key points as key point sub-features of the currently traversed video frame;
[0045] The target key points and key point sub-features of each video frame are used as the human body key point features of the target person in each video frame;
[0046] The key point features of the human body of several video frames are arranged according to the time sequence to generate the key point feature sequence.
[0047] Another embodiment of the present invention also provides a terminal device, comprising: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the steps of a human behavior recognition method with multimodal feature fusion as described in any one of the above embodiments of the present invention.
[0048] Another embodiment of the present invention also provides a computer-readable storage medium item, including: a stored computer program, which, when the computer program is running, controls the device where the computer-readable storage medium is located to execute the steps of a human behavior recognition method with multimodal feature fusion as described in any one of the above embodiments of the present invention.
[0049] The following beneficial effects are achieved by implementing the present invention:
[0050] The present invention provides a multimodal feature fusion human behavior recognition method, apparatus, terminal device, and storage medium. The method performs a key point extraction operation on a target video to identify the key point features of a target person in each video frame of the target video, thereby generating a key point feature sequence. Furthermore, the method performs temporal feature extraction on the key point feature sequence to identify the dynamic change characteristics of the key point features of the human body according to the time sequence, thereby generating a temporal feature matrix that can reflect the dynamic changes of human behavior. Subsequently, the spatial features and optical flow features of each video frame in the target video are identified to construct a spatial feature matrix and an optical flow feature matrix. The temporal feature matrix, the spatial feature matrix, and the optical flow feature matrix are fused to generate a fused feature matrix that can be used to characterize the target person's action changes and environmental space changes in the video frames. Finally, the fused feature matrix is input into a behavior recognition model so that the behavior recognition model recognizes and outputs the target person's behavior based on the fused feature matrix. Therefore, the present invention captures the dynamic change characteristics of human body movements from the target video through temporal feature extraction, and captures the spatial change characteristics from the target video through spatial feature extraction and optical flow feature extraction, so that the behavior recognition model can pay attention to the movement changes of the target person and the spatial changes of the environment when performing behavior recognition, thereby effectively improving the recognition accuracy of the behavior recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the implementation. Obviously, the drawings described below are only some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0052] Figure 1 1 is a flow chart of a method for human behavior recognition using multimodal feature fusion provided by one embodiment of the present invention;
[0053] Figure 2 This is a structural diagram of a multimodal feature fusion human behavior recognition device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0054] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned figure descriptions are intended to cover non-exclusive inclusions.
[0056] In the description of the embodiments of this application, the technical terms "first" and "second" are used only to distinguish different objects and should not be understood to indicate or imply relative importance or implicitly specify the quantity, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "plurality" is more than two, unless otherwise clearly and specifically defined.
[0057] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0058] In the description of the embodiments of this application, the term "and / or" is simply a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0059] In the description of the embodiments of the present application, the term "multiple" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).
[0060] In the description of the embodiments of the present application, unless otherwise expressly specified or limited, technical terms such as "installed," "connected," "connected," and "fixed" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integration; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; internal connections between two components or interactions between two components. Those skilled in the art can understand the specific meanings of the above terms in the embodiments of the present application based on specific circumstances.
[0061] See also Figure 1To address the drawback of current human behavior recognition methods that they are unable to accurately identify human motion behaviors, an embodiment of the present invention provides a human behavior recognition method using multimodal feature fusion, comprising:
[0062] S1. Obtain the target video to be identified;
[0063] In a preferred embodiment of the present invention, taking a security monitoring scenario as an example, a remote connection is established with a camera installed in a security area, and the monitoring video captured by the camera in real time is obtained as the target video.
[0064] S2, extracting key points from the target video, identifying key point features of the target person in each video frame of the target video, and generating a key point feature sequence;
[0065] In a preferred embodiment of the present invention, when only one target person exists in a target video frame, key point extraction is performed directly on the target person. When multiple target persons exist in a target video frame, a target recognition algorithm is used to define different labels for different target persons. Based on the labels, key point extraction is performed on the target video multiple times, generating a key point feature sequence for each target person each time.
[0066] Preferably, extracting key points from the target video, identifying key point features of the target person in each video frame of the target video, and generating a key point feature sequence includes:
[0067] S21, splitting the target video into several video frames;
[0068] S22, performing human skeleton detection on the video frame, identifying facial features and body joints of the target person in the video frame as initial key points;
[0069] S23, normalizing the size of the target person according to the initial key points to generate a normalization factor;
[0070] S24, normalizing the initial key points according to the normalization factor to generate corresponding target key points;
[0071] S25, traversing the video frames, calculating the angle and distance between each target key point and other target key points in the currently traversed video frame as the key point sub-feature of the currently traversed video frame;
[0072] S26, using the target key points and key point sub-features of each video frame as the human body key point features of the target person in each video frame;
[0073] S27. Arrange the key point features of the human body of several video frames according to the time sequence to generate the key point feature sequence.
[0074] In a preferred embodiment of the present invention, the target video is split into a video frame sequence:
[0075] I={I1,I2,……,I t},t∈T,
[0076] Where I is the video frame sequence, T is the number of frames that the target video is split into, and I t is the tth video frame.
[0077] Use open source algorithms such as OpenPose and HRNet to detect the human skeleton and obtain initial key points:
[0078] P t ={(x1,y1),(x2,y2),...,(x K ,y K )}, t = 1, 2, ..., T;
[0079] Among them, P t is the initial key point set of the t-th video frame, K is the number of key points (set K = 17), (x i ,y i ) is the coordinate of the i-th initial key point.
[0080] In this embodiment, the initial key points extracted include: nose, right eye, left eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, right hip, left knee, right knee, left ankle, right ankle, left wrist, right wrist, and neck. In addition to the initial key points, the coordinates of the hip center point are also extracted as key coordinates for correcting the character's proportions.
[0081] Furthermore, since the problem of shooting angle may cause abnormal human body proportions, directly using the original coordinates is not conducive to subsequent feature extraction. Therefore, it is necessary to standardize the initial key points through the following steps:
[0082] First, normalize the size of the target person and calculate the shoulder width of the person according to the coordinates of the right shoulder and left shoulder in the initial key points as the normalization factor:
[0083]
[0084] Among them, s t is the normalization factor of the t-th video frame, is the horizontal coordinate of the left shoulder of the t-th video frame, is the horizontal coordinate of the right shoulder of the t-th video frame, is the ordinate of the left shoulder of the t-th video frame, is the vertical coordinate of the right shoulder of the t-th video frame.
[0085] The initial key points are normalized using the following formula to generate the corresponding target key points;
[0086]
[0087] in, is the key coordinate of the t-th video frame, is the i-th initial key point of the t-th video frame, is the target key point corresponding to the i-th initial key point.
[0088] Furthermore, the angle and distance between each target key point and other target key points are calculated using the following formula:
[0089]
[0090] in, is the distance between the i-th target key point and the j-th target key point in the t-th video frame, is the angle between the i-th target keypoint and the j-th target keypoint in the t-th video frame.
[0091] Finally, the following key features of the human body that can be used to identify human posture and behavior patterns are constructed for each video frame:
[0092]
[0093] Among them, F t is the key point feature of the human body in the t-th video frame.
[0094] S3, extracting time series features from the key point feature sequence, identifying dynamic change features of the key point features of the human body according to the time sequence, and generating a time series feature matrix;
[0095] Preferably, extracting the time series features from the key point feature sequence, identifying the dynamic change features of the key point features of the human body according to the time sequence, and generating a time series feature matrix includes:
[0096] S31, inputting the key point feature sequence into a preset temporal feature extraction model, so that the forward LSTM sub-model of the temporal feature extraction model processes the human key point features in the key point feature sequence in a positive order to generate forward temporal information of each human key point feature, and the backward LSTM sub-model of the temporal feature extraction model processes the human key point features in the key point feature sequence in a reverse order to generate backward temporal information of each human key point feature, and then fusing the forward temporal information and the backward temporal information of each human key point feature to generate a dynamic change feature of each human key point feature;
[0097] S32. Construct the temporal feature matrix based on the dynamic change characteristics of each of the key point features of the human body. In a preferred embodiment of the present invention, the temporal feature extraction model is a neural network model based on the Bi-LSTM (Bidirectional Long Short-Term Memory) architecture. Bi-LSTM is an improved recurrent neural network (RNN) model that can simultaneously capture the contextual information of sequence data by combining forward and backward LSTM layers.
[0098] In this embodiment, the key point feature sequence P1, P2, ..., P T Input into the time series feature extraction model so that the forward LSTM sub-model processes the human key point features in the key point feature sequence in positive order to generate the hidden state of each human key point feature (forward time series information):
[0099]
[0100] The backward LSTM sub-model of the temporal feature extraction model processes the human key point features in the key point feature sequence in reverse order to generate the hidden state of each human key point feature (backward temporal information):
[0101]
[0102] The forward and backward time series information of each key point feature of the human body are fused to generate the dynamic change features of each key point feature of the human body:
[0103]
[0104] in, is the forward temporal information of the key point features of the human body in the t-th video frame, is the backward temporal information of the key point features of the human body in the t-th video frame, σ is the sigmoid function, W fis the first weight matrix of the forward LSTM sub-model, which is used to control the positive sequence of key point feature sequence P T Hidden state The impact of U f is the second weight matrix of the forward LSTM sub-model, which is used to control the previous hidden state in the positive sequence of key point features. Hide state for the current moment The impact of b f is the first bias matrix of the forward LSTM sub-model, W b is the third weight matrix of the backward LSTM sub-model, which is used to control the reverse order of the key point feature sequence P T Hidden state The impact of U b is the fourth weight matrix of the forward LSTM sub-model, which is used to control the previous hidden state in the reverse key point feature sequence Hide state for the current moment The impact of b b is the second bias matrix of the backward LSTM sub-model.
[0105] Finally, the time series feature matrix H is obtained:
[0106] H={h1, h2, ..., h T}.
[0107] S4, performing spatial feature extraction and optical flow feature extraction on the target video, identifying the spatial features and optical flow features of each video frame, and constructing a spatial feature matrix and an optical flow feature matrix;
[0108] Preferably, performing spatial feature extraction and optical flow feature extraction on the target video, identifying the spatial features and optical flow features of each video frame, and constructing a spatial feature matrix and an optical flow feature matrix includes:
[0109] S41, inputting several video frames of the target video into a spatial feature extraction model, so that the spatial feature extraction model recognizes and outputs image edge features, texture features, and local features of the video frames;
[0110] S42, using the image edge features, texture features, and local features of each of the video frames as spatial features;
[0111] In a preferred embodiment of the present invention, in order to extract the spatial features of each video frame, a deep convolutional neural network such as ResNet is used to extract features from the image. Assume that the tth frame of the target video is I t , after ResNet processing, the extracted spatial features are expressed as:
[0112] Ft =ResNet(I t );
[0113] Among them, F t is the spatial feature of the t-th frame.
[0114] ResNet extracts deep spatial features such as edges, textures, local features, etc. through multi-layer convolution and residual connections. Finally, the visual feature matrix of the entire video is obtained:
[0115] F=[F1,F2,...,F T ]∈R T×dv ;
[0116] Among them, dv is the dimension of spatial features.
[0117] S43, traversing the video frames, calculating the moving direction and moving speed of corresponding pixels between the currently traversed video frame and the next video frame, and using the moving direction and moving speed as the optical flow features of the currently traversed video frame.
[0118] In a preferred embodiment of the present invention, optical flow is a computer vision technique that estimates the motion of an object in an image by analyzing the changes in pixel intensity between two images. In this case, the motion vector V is calculated using the OpticalFlow(·) function. t , the vector represents the video frame I t The displacement and speed of each pixel in . Specifically, the optical flow calculation formula is:
[0119] V t =OpticalFlow(I t , I t+1 );
[0120] Among them, V t is the optical flow feature of the t-th video frame, which contains the motion direction and speed information of each pixel; I t , I t+1 are two adjacent video frames;
[0121] Finally, the motion feature matrix of all frames is expressed as:
[0122] V=[V1,V2,...,V T ]∈R T×dm ;
[0123] Among them, dm is the dimension of optical flow features.
[0124] S5, fusing the temporal feature matrix, the spatial feature matrix, and the optical flow feature matrix of each video frame to generate a fused feature matrix of the video frame;
[0125] Preferably, the fusion feature matrix is:
[0126] X t =[H t ; F t ; V t ];
[0127] Among them, X t is the fusion feature matrix of the t-th video frame, H t is the temporal feature matrix of the t-th video frame, F t is the spatial feature matrix of the t-th video frame, V t is the optical flow feature matrix of the t-th video frame.
[0128] S6. Input the fused feature matrix into a preset behavior recognition model, so that the behavior recognition model recognizes and outputs the target behavior of the target person according to the fused feature matrix.
[0129] Preferably, the behavior recognition model identifies and outputs the target behavior of the target person according to the fusion feature matrix, including:
[0130] S61, using a global attention mechanism to weight the temporal feature matrix, the spatial feature matrix, and the optical flow feature matrix in each of the fused feature matrices to generate a weighted feature matrix;
[0131] S62, using a Softmax function to calculate the probability distribution of the target person in a plurality of preset behaviors according to the weighted feature matrix;
[0132] S63. The behavior corresponding to the maximum probability in the probability distribution is used as the target behavior of the target person.
[0133] In a preferred embodiment of the present invention, Transformer is used to perform global attention modeling on the fused features:
[0134] Q=W Q Z, K = W K Z, V = W V Z;
[0135]
[0136] Among them, Q, K, V are query, key, and value matrices, and W Q ,W K ,W V is the learnable parameter of Transformer, d k is the scaling factor of the feature dimension, and A is the attention-weighted fusion feature matrix.
[0137] Finally, Softmax is used for behavior classification:
[0138] y=softmax(W o Attention(X)+b o );
[0139] Among them, W o is the weight matrix of the fully connected layer, and Attention(X) is the fused feature matrix, i.e. A. y is the probability distribution of behavior categories, and the behavior category corresponding to the maximum probability is selected as the final recognition result.
[0140] In summary, this embodiment provides a method for human behavior recognition based on multimodal feature fusion. By performing a key point extraction operation on the target video, the key point features of the target person in each video frame of the target video are identified to generate a key point feature sequence. Further, by performing temporal feature extraction on the key point feature sequence, the dynamic change features of the key point features of the human body are identified according to the time sequence, and a temporal feature matrix that can reflect the dynamic changes of human behavior is generated. Then, by identifying the spatial features and optical flow features of each video frame in the target video, a spatial feature matrix and an optical flow feature matrix are constructed, and the temporal feature matrix, the spatial feature matrix and the optical flow feature matrix are fused to generate a fused feature matrix that can be used to characterize the target person's action changes and environmental space changes in the video frame. Finally, the fused feature matrix is input into the behavior recognition model so that the behavior recognition model recognizes and outputs the target person's behavior based on the fused feature matrix. Therefore, the present invention captures the dynamic change characteristics of human body movements from the target video through temporal feature extraction, and captures the spatial change characteristics from the target video through spatial feature extraction and optical flow feature extraction, so that the behavior recognition model can pay attention to the movement changes of the target person and the spatial changes of the environment when performing behavior recognition, thereby effectively improving the recognition accuracy of the behavior recognition model.
[0141] like Figure 2 As shown, based on the above method embodiment, a corresponding device embodiment is provided;
[0142] An embodiment of the present invention provides a human behavior recognition device using multimodal feature fusion, comprising:
[0143] A video acquisition module is used to acquire the target video to be identified;
[0144] A human feature extraction module is used to extract key points from the target video, identify key point features of the target person in each video frame of the target video, and generate a key point feature sequence;
[0145] A time series feature extraction module is used to extract time series features from the key point feature sequence, identify the dynamic change characteristics of the key point features of the human body according to the time sequence, and generate a time series feature matrix;
[0146] An environmental feature extraction module is used to extract spatial features and optical flow features of the target video, identify the spatial features and optical flow features of each video frame, and construct a spatial feature matrix and an optical flow feature matrix;
[0147] a feature fusion module, configured to fuse the temporal feature matrix, the spatial feature matrix, and the optical flow feature matrix of each video frame to generate a fused feature matrix of the video frame;
[0148] The behavior recognition module is used to input the fused feature matrix into a preset behavior recognition model, so that the behavior recognition model recognizes and outputs the target behavior of the target person according to the fused feature matrix.
[0149] Preferably, the human feature extraction module extracts key points from the target video, identifies key point features of the target person in each video frame of the target video, and generates a key point feature sequence, including:
[0150] Splitting the target video into several video frames;
[0151] Performing human skeleton detection on the video frame to identify facial features and body joints of the target person in the video frame as initial key points;
[0152] Normalizing the size of the target person according to the initial key points to generate a normalization factor;
[0153] Normalizing the initial key points according to the normalization factor to generate corresponding target key points;
[0154] Traversing the video frames, calculating the angles and distances between each target key point in the currently traversed video frame and other target key points as key point sub-features of the currently traversed video frame;
[0155] The target key points and key point sub-features of each video frame are used as the human body key point features of the target person in each video frame;
[0156] The key point features of the human body of several video frames are arranged according to the time sequence to generate the key point feature sequence.
[0157] It can be understood that the above-mentioned device embodiment corresponds to the method embodiment of the present invention, which can implement any of the above-mentioned method embodiments of the present invention to provide a human behavior recognition method with multimodal feature fusion.
[0158] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Furthermore, in the drawings of the device embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which may be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement the present invention without inventive effort.
[0159] Based on the above-mentioned embodiment of a method for recognizing human behavior by fusion of multimodal features, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, a method for recognizing human behavior by fusion of multimodal features according to any embodiment of the present invention is implemented.
[0160] For example, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to implement the present invention. The one or more module elements may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.
[0161] The terminal device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0162] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, connecting various parts of the entire terminal device using various interfaces and lines.
[0163] Based on the above-mentioned method embodiments, another embodiment of the present invention provides a computer-readable storage medium, including a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a human behavior recognition method with multimodal feature fusion as described in any one of the above-mentioned method embodiments of the present invention.
[0164] Wherein, the module / unit integrated in the device / terminal equipment, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0165] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A human behavior recognition method based on multimodal feature fusion, characterized in that: include: Obtain the target video to be identified; Extracting key points from the target video, identifying key point features of the target person in each video frame of the target video, and generating a key point feature sequence; Extracting time series features from the key point feature sequence, identifying dynamic change features of the key point features of the human body according to the time sequence, and generating a time series feature matrix; Performing spatial feature extraction and optical flow feature extraction on the target video, identifying the spatial features and optical flow features of each of the video frames, and constructing a spatial feature matrix and an optical flow feature matrix; Fusing the temporal feature matrix, the spatial feature matrix, and the optical flow feature matrix of each video frame to generate a fused feature matrix of the video frame; The fused feature matrix is input into a preset behavior recognition model, so that the behavior recognition model recognizes and outputs the target behavior of the target person according to the fused feature matrix.
2. The method for human behavior recognition based on multimodal feature fusion according to claim 1, wherein: The step of extracting key points from the target video, identifying key point features of the target person in each video frame of the target video, and generating a key point feature sequence includes: Splitting the target video into several video frames; Performing human skeleton detection on the video frame to identify facial features and body joints of the target person in the video frame as initial key points; Normalizing the size of the target person according to the initial key points to generate a normalization factor; Normalizing the initial key points according to the normalization factor to generate corresponding target key points; Traversing the video frames, calculating the angles and distances between each target key point in the currently traversed video frame and other target key points as key point sub-features of the currently traversed video frame; The target key points and key point sub-features of each video frame are used as the human body key point features of the target person in each video frame; The key point features of the human body of several video frames are arranged according to the time sequence to generate the key point feature sequence.
3. The method for human behavior recognition based on multimodal feature fusion according to claim 2, wherein: The step of extracting time series features from the key point feature sequence, identifying dynamic change features of key point features of the human body according to the time sequence, and generating a time series feature matrix includes: The key point feature sequence is input into a preset time series feature extraction model, so that the forward LSTM sub-model of the time series feature extraction model processes the human key point features in the key point feature sequence in a positive order to generate forward time series information of each human key point feature, and the backward LSTM sub-model of the time series feature extraction model processes the human key point features in the key point feature sequence in a reverse order to generate backward time series information of each human key point feature, and then the forward time series information and the backward time series information of each human key point feature are fused to generate a dynamic change feature of each human key point feature; The time series feature matrix is constructed according to the dynamic change characteristics of each of the key point features of the human body.
4. The method for human behavior recognition based on multimodal feature fusion according to claim 3, wherein: The step of extracting spatial features and optical flow features from the target video, identifying spatial features and optical flow features of each video frame, and constructing a spatial feature matrix and an optical flow feature matrix includes: Inputting several video frames of the target video into a spatial feature extraction model so that the spatial feature extraction model recognizes and outputs image edge features, texture features, and local features of the video frames; Using the image edge features, texture features, and local features of each of the video frames as spatial features; The video frames are traversed, and the moving direction and moving speed of corresponding pixels between the currently traversed video frame and the next video frame are calculated, and the moving direction and moving speed are used as the optical flow features of the currently traversed video frame.
5. The method for human behavior recognition based on multimodal feature fusion according to claim 4, wherein: The fusion feature matrix is: X t =[H t ;F t ;V t ]; Among them, X t is the fusion feature matrix of the t-th video frame, H t is the temporal feature matrix of the t-th video frame, F t is the spatial feature matrix of the t-th video frame, V t is the optical flow feature matrix of the t-th video frame.
6. The method for human behavior recognition based on multimodal feature fusion according to claim 5, wherein: The behavior recognition model identifies and outputs the target behavior of the target person according to the fusion feature matrix, including: Using a global attention mechanism, the temporal feature matrix, the spatial feature matrix, and the optical flow feature matrix in each of the fused feature matrices are weighted to generate a weighted feature matrix; Using the Softmax function, based on the weighted feature matrix, the probability distribution of the target person in a plurality of preset behaviors is calculated; The behavior corresponding to the maximum probability in the probability distribution is used as the target behavior of the target person.
7. A multimodal feature fusion human behavior recognition device, characterized in that: include: A video acquisition module is used to acquire the target video to be identified; A human feature extraction module is used to extract key points from the target video, identify key point features of the target person in each video frame of the target video, and generate a key point feature sequence; A time series feature extraction module is used to extract time series features from the key point feature sequence, identify the dynamic change characteristics of the key point features of the human body according to the time sequence, and generate a time series feature matrix; An environmental feature extraction module is used to extract spatial features and optical flow features of the target video, identify the spatial features and optical flow features of each video frame, and construct a spatial feature matrix and an optical flow feature matrix; a feature fusion module, configured to fuse the temporal feature matrix, the spatial feature matrix, and the optical flow feature matrix of each video frame to generate a fused feature matrix of the video frame; The behavior recognition module is used to input the fused feature matrix into a preset behavior recognition model, so that the behavior recognition model recognizes and outputs the target behavior of the target person according to the fused feature matrix.
8. The multimodal feature fusion human behavior recognition device according to claim 7, characterized in that: The human feature extraction module extracts key points from the target video, identifies key point features of the target person in each video frame of the target video, and generates a key point feature sequence, including: Splitting the target video into a plurality of video frames; Performing human skeleton detection on the video frame to identify facial features and body joints of the target person in the video frame as initial key points; Normalizing the size of the target person according to the initial key points to generate a normalization factor; Normalizing the initial key points according to the normalization factor to generate corresponding target key points; Traversing the video frames, calculating the angles and distances between each target key point in the currently traversed video frame and other target key points as key point sub-features of the currently traversed video frame; The target key points and key point sub-features of each video frame are used as the human body key point features of the target person in each video frame; The key point features of the human body of several video frames are arranged according to the time sequence to generate the key point feature sequence.
9. A terminal device, characterized in that: The invention comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for human behavior recognition by multimodal feature fusion according to any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium, characterized in that include: A stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute a human behavior recognition method based on multimodal feature fusion as described in any one of claims 1 to 6.
Citation Information
Cited By
Gait recognition method, device and equipment and storage medium
CN121214552A