A super joint and multimodal network and its application in behavior recognition method

By introducing skeleton dependency and multimodal spatiotemporal feature representation in human behavior recognition, combined with depth maps and skeleton sequences, the problems of insensitive joint posture changes and neglected dependencies in the prior art are solved, and more efficient action recognition effects are achieved.

CN114782992BActive Publication Date: 2025-05-06CHANGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210464698.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-05-06
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

The existing human behavior recognition algorithm is not sensitive enough when dealing with local posture changes in human joints, and ignores the dependence between joints, resulting in limited recognition effect.

Method used

A human behavior recognition method based on skeleton dependency and multimodal spatiotemporal feature representation is proposed. Through super joints and multimodal networks, combining depth maps and skeleton sequences, rich texture information and spatiotemporal features are learned, and action recognition performance is improved by fusing scores of different modes.

Benefits of technology

By establishing joint dependency and multimodal feature fusion, the accuracy and robustness of human behavior recognition are significantly improved, especially when processing multimodal data, it shows a high recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782992B_ABST
    Figure CN114782992B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of neural network technology, and in particular to a super-joint and multimodal network and a behavior recognition method thereof, comprising: collecting a human body depth map, extracting features from the depth map using a DMMS stream, and calculating a depth data prediction score; collecting a human body skeleton sequence, extracting original joint and super-joint data respectively, combining super-joints and common joints to construct skeleton information, and sending it into a structured spatiotemporal feature learning model to obtain static and dynamic joint data streams, static and dynamic super-joint data streams, and adaptively weighting the original joint and super-joint data streams to obtain joint data prediction scores and super-joint prediction scores respectively; adding the classification prediction scores of the DMMS stream and the skeleton stream to generate a final prediction score. The present invention learns the rich texture information of human body parts in space from the depth map, and learns the rich spatiotemporal features in the changes of motion postures from the skeleton sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of neural network technology, and in particular to a super-joint and multimodal network and a behavior recognition method thereof. Background Art

[0002] Human behavior recognition has become a hot topic in the field of machine learning and computer vision because of its broad application prospects. It has very important theoretical research value. Unlike the recognition of static images, behavior recognition needs to comprehensively consider continuously changing images and construct the spatiotemporal feature representation of human behavior from a series of action sequences composed of static human posture images to achieve classification and recognition of actions. At present, human behavior recognition has been widely used in the fields of assisted human-computer interaction, sports analysis, intelligent monitoring and virtual reality.

[0003] At present, behavior recognition algorithms are mainly divided into two categories:

[0004] (1) Action recognition based on traditional methods. In traditional methods, scholars extract manual features as effective spatiotemporal feature descriptors. Donglu Li et al. combined the characteristics of depth data with point cloud data and used depth information to reconstruct the three-dimensional point cloud space to improve the spatial information expression ability of depth data. Then they evenly divided the point cloud space into many grid spaces, calculated the ratio of the number of point clouds in the grid space as the spatial distribution feature for action recognition. Liu et al. established distribution sectors on the two-dimensional projection plane of the three-dimensional skeleton joint trajectory, and then used histogram statistics of the distribution sectors of specific projections as the skeleton-based action descriptor HDS-SP to characterize spatial and temporal information.

[0005] (2) Action recognition based on deep learning. The method of action recognition based on extracting manual features has a high complexity in implementation. Traditional methods require feature engineering, and the analysis and design of motion features greatly affect the classification accuracy. In the action recognition method based on deep learning, the feature learning network is first used to learn the feature representation of the action sequence, and then the features are sent to the classification network for classification. Shengquan Wang et al. proposed a multi-level deep feature fusion enhancement network (MDFFEN), which first extracts spatiotemporal features from the sub-modules of the two-stream network to form multi-depth features, and then fuses these spatiotemporal features for classification. Haoran Wang et al. proposed a skeleton edge motion network (SEMN), which mines the motion information of the human body based on the skeleton edge motion, and combines the angle change of the skeleton edge with the corresponding joint motion to determine the motion of the skeleton edge.

[0006] Although depth images carry dense texture information, the rich texture information also brings high information redundancy, making them insensitive to local posture changes of human joints. Human skeleton data has become a hot research object in the field of human action recognition due to its high sensitivity to posture changes. Compared with depth image data, the three-dimensional spatial structure information provided by skeleton data can greatly reduce the influence of factors such as height, body shape and clothing of the data collection object. Although skeleton joints are specially designed to capture the spatial structure information of the human body, the information carried by the joint points can only represent the position but not the dependency relationship with adjacent points. Previous methods usually use the position of the joints as the representation of the spatial structure information of the joints, while ignoring the dependency relationship between the joints. Summary of the invention

[0007] In view of the shortcomings of existing algorithms, the present invention proposes a human behavior recognition method based on skeleton dependency and multimodal spatiotemporal feature representation, which learns the rich texture information of human body parts in space from depth maps and the rich spatiotemporal features in motion posture changes from skeleton sequences.

[0008] The technical solution adopted by the present invention is: a super joint and multimodal network and a behavior recognition method thereof include the following steps:

[0009] S1, collect human body depth data, input DMMS stream to extract features of the depth map, and calculate the DMMs stream prediction score;

[0010] S2. Collect human skeleton sequence, extract original joint and super-joint data respectively, construct skeleton information by combining super-joints and ordinary joints, calculate static and motion data of super-joints and ordinary joints respectively and send them into structured spatiotemporal feature learning model to obtain static and dynamic joint data streams, static and dynamic super-joint data streams, perform feature-level fusion on static and motion data from ordinary joints and super-joints respectively, and calculate prediction scores of ordinary joints and super-joints.

[0011] S3, firstly, adaptively weight the classification scores of the original joints and super joints to obtain the prediction scores of the skeleton flow;

[0012] S21, constructing super joints according to the dependency relationship of common joints, and constructing skeleton information by combining super joints and common joints;

[0013] Further, the super joint construction includes:

[0014] First, calculate RWE and RWH, which are the direction vectors of RW pointing to RE and RH respectively. The calculation formula is as follows:

[0015]

[0016] Among them, RW, RE and RH represent the right wrist, right elbow and right hand in the human skeleton respectively;

[0017] Then, calculate the Cartesian product of the two intersecting vectors to get the normal vector n of the plane where the vectors are located. The calculation formula is as follows:

[0018]

[0019] Then, taking the direction vectors of the two bones as the starting point RE, calculate the angle between the two direction vectors. The calculation formula is as follows:

[0020]

[0021] Finally, we get the hyperjoint data vector HyperJoint, the formula is:

[0022]

[0023] Furthermore, combining super joints and common joints to construct skeleton information includes the following steps:

[0024] The skeleton diagram of the human body is represented by the variable G(V,H), and the formula is:

[0025]

[0026]

[0027] Among them, V is the set of spatial nodes that constitute the human skeleton graph, and H is the dependency relationship established on the joints of the human skeleton;

[0028] The dynamic information of a skeleton is defined as the difference between the skeleton elements on adjacent skeletons in the skeleton sequence, which is used to establish a relationship between the skeleton postures in the time domain. The formula is as follows:

[0029]

[0030]

[0031] S22, sending the skeleton information into the structured spatiotemporal feature learning model, including the following steps:

[0032] The spatiotemporal feature learning module of convolutional neural network is used to extract human motion features based on skeleton information;

[0033] Furthermore, the spatiotemporal feature learning module includes a node temporal feature learning block and a spatial global feature learning block. By superimposing the spatiotemporal blocks, a feature learning module is constructed to learn effective spatiotemporal features of the original skeleton nodes and the dependencies between nodes from the skeleton data;

[0034] Furthermore, the node temporal feature learning block construction includes the following steps:

[0035] First, the convolutional block attention module is used to increase attention to the tensor of the input network, which means improving the feature expression ability of the input data;

[0036] Then, a convolutional layer with a convolution kernel size of (1×1) is used in the network to automatically learn the position features of the skeleton joints; the formulation is defined as:

[0037] f T (X) = σ(φ(Attn(X))) (9)

[0038] Where X represents a third-order tensor φ represents the function of the convolutional layer structure, σ represents the ReLU activation function, and Attn represents an attention block;

[0039] Finally, a convolution kernel of size (3×1) is used to aggregate the feature information learned at the skeleton joint positions in time, and the output is a third-order tensor It is defined by the following formula:

[0040] A=φ(f T (X)) (10).

[0041] Furthermore, the spatial global feature learning block construction includes:

[0042] The convolution kernel of size (3×3) is used to automatically learn the semantic features of the coordinated motion of skeleton nodes in the spatial domain on the network. The convolution operation can aggregate the spatial structural information of the elements that constitute the human body, which is formally defined as:

[0043] f S (A) = φ(A') (11)

[0044] Among them, A' represents the transformation of the matrix, which transforms a third-order tensor Transformed into

[0045] S3, adding the classification prediction scores of the DMMS stream and the skeleton stream to generate the final prediction score;

[0046] Beneficial effects of the present invention:

[0047] 1. Skeleton node dependency is proposed to remodel the spatial structure information representation of the human skeleton sequence, and the isolated nodes that constitute the human skeleton are established with the adjacent points. The human skeleton sequence and the skeleton node dependency sequence are used as the input of the network for spatiotemporal feature learning.

[0048] 2. A new spatiotemporal feature learning network is designed to learn the spatiotemporal features of skeleton sequences and skeleton node dependencies. In the network, a node temporal feature learning block (JTFLB) is constructed to explore the temporal relationship between elements on the human skeleton. Due to the integrity and coordination of human movement, a spatial global feature learning block (SGFLB) is constructed to aggregate the semantic features learned in the temporal domain of the elements that constitute the human skeleton.

[0049] 3. Collaboratively extract rich spatiotemporal features from the depth sequence and skeleton sequence of the human body, and improve the action recognition performance by fusing the scores of different modalities; In order to verify the effectiveness of the present invention, an evaluation was carried out on public datasets of different scales. Compared with a single modality, the recognition effect was significantly improved after multimodal fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a structural block diagram of the super joint and multimodal network of the present invention and the behavior recognition method thereof;

[0051] Figure 2 It is a schematic diagram of calculating joint posture direction of the present invention;

[0052] Figure 3 is a schematic diagram of calculating joint angles of the present invention;

[0053] Figure 4 It is a schematic diagram of establishing a super joint of the present invention;

[0054] Figure 5 It is a block diagram of joint timing feature learning of the present invention;

[0055] Figure 6 is a spatial global feature learning block diagram of the present invention;

[0056] Figure 7 is the frame statistics of the UTD-MHAD and NTU RGB+D datasets of the present invention;

[0057] Figure 8 is the difference between the basic model and the fusion model of the present invention on UTD-MHAD;

[0058] Fig. 9 It is the difference between the basic model and the fusion model of the present invention on NTU RGB+D CS;

[0059] Fig.10 It is the difference between the basic model and the fusion model of the present invention on NTU RGB+D CV. DETAILED DESCRIPTION

[0060] The present invention is further described below in conjunction with the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner, and therefore it only shows the components related to the present invention.

[0061] In order to evaluate the effectiveness of the method of the present invention, experiments were conducted on public datasets based on depth maps and skeleton information; large public datasets can provide a wider range of training data for the model, making the model stronger; in order to verify the robustness of the method of the present invention, a classic small dataset was used in the selection of the dataset, therefore, experiments were conducted on several datasets with completely different scales: UTD-MHAD and NTU-RGB+D.

[0062] The present invention is based on the PyTorch framework, wherein the Python version is 3.7.0 and the pytorch version is 1.10.1; the hardware platform of this experiment is a desktop computer, wherein the motherboard is MSI B460M MORTAR, the CPU is Intel i710700, the main frequency is 2.9GHz, the memory is 16GB, the operating system is Windows 10 Professional Edition, the GPU resource is NVIDIA TeslaV100, and the video memory is 32GB; the software tools used in the experiment are PyCharm and Anaconda3.

[0063] The UTD-MHAD dataset includes a total of 861 action samples, and uses a fixed-position acquisition device to provide each action sample with four types of data captured from one perspective: RGB, depth map, skeleton sequence, and inertial sensor sequence; it includes 27 action categories, and each action category is repeated 3-4 times by 8 subjects; in order to facilitate comparison with the existing technology, the action sequences of subjects numbered 1, 3, 5, and 7 are used as training sets, and the rest are used as test sets.

[0064] NTU RGB+D is a large dataset and provides more challenging action samples and more modal information. The NTU RGB+D dataset includes a total of 56,880 action samples, and uses fixed acquisition devices at different positions to provide each sample with four types of data: RGB, depth map, skeleton sequence, and infrared radiation video captured from three perspectives. It includes 60 action categories, each of which is completed 1-2 times by 40 subjects. Unlike the UTD-MHAD acquisition method, NTU RGB+D provides 17 types of multi-view and multi-modal data collected by acquisition devices at different levels and distances. Two standard experimental strategies are used on the dataset, which are cross-subject and cross-perspective. In the cross-subject strategy, all subjects are first divided into two groups, where the action sequences of subjects numbered 1, 2, 4, 5, 8, 9, 13, 14, 15, 16, 17, 18, 19, 25, 27, 28, 31, 34, 35, and 38 are used as training sets, and the rest are used as test sets. In the cross-view strategy, the cameras with different viewpoints in the acquisition device are first divided into two groups, of which the 37,920 action sequences acquired by cameras numbered 2 and 3 are used as the training set, and the 18,960 action sequences acquired by camera numbered 1 are used as the test set.

[0065] like Figure 1 As shown, a super joint and multimodal network and a behavior recognition method thereof include the following steps:

[0066] Figure 1 There are two flows, namely Figure 1 Upper half: DMMs flow and Figure 1 Lower half: Skeleton flow, Figure 1 (a) Calculate the static and motion data of the original joint and super joint respectively; Figure 1 (b) A structured spatiotemporal feature learning module is established by stacking JTFLB and SGFLB to learn spatiotemporal feature representations from skeleton sequences; Figure 1 (c) Feature-level fusion of static and motion data from joints and super-joints respectively; Figure 1 (d) First, the prediction scores of the original joint and the super-joint are adaptively weighted and fused, and then the prediction scores of the two streams are added together to generate the final prediction score;

[0067] S1 collects the human body depth map, uses the DMMS stream to extract features from the depth map, and calculates the depth data prediction score;

[0068] S2. Collect human skeleton sequence, extract original joint and super-joint data respectively, and send them into structured spatiotemporal feature learning model to obtain static and dynamic joint data streams, static and dynamic super-joint data streams respectively, and perform adaptive weight fusion on the original joint and super-joint data streams to obtain joint data prediction scores and super-joint prediction scores respectively;

[0069] Furthermore, given a human skeleton sequence {S1,S2,...,S T}, calculate an action descriptor that can better describe the change of action; the human skeleton sequence is regarded as a function It represents the changes in the spatial structure of the human skeleton in the time dimension t=1,2,...,T; however, the joints that make up the human skeleton only carry position information, so dependencies are established between naturally connected joints in the human body to obtain high-level information other than joint positions.

[0070] In the human behavior represented by skeleton sequence, joints are the main moving parts, such as Figure 2 As shown, RW, RE and RH represent the right wrist, right elbow and right hand in the human skeleton respectively. They are naturally connected on the human body. Taking the movement of the right wrist joint as an example, the coordinates of RW, RE and RH determine a plane in space, and the normal vector of the plane determines the direction. The normal vector of the plane is calculated as the direction descriptor of the local posture centered on the right wrist joint;

[0071] First, calculate RWE and RWH, which are the direction vectors of RW pointing to RE and RH respectively. The calculation formula is as follows:

[0072]

[0073] Then, calculate the Cartesian product of the two intersecting vectors to get the normal vector n of the plane where the vectors are located. The calculation formula is as follows:

[0074]

[0075] For different actions, the amplitude of bone movement is different, such as Figure 3 As shown, when the right wrist joint moves, the angle between the two bones connected by the joint will also change. Therefore, the angle between the two bones connected by the joint is calculated as the descriptor of the movement amplitude;

[0076] First, we need to calculate the direction vectors of the two bones with RE as the starting point, as shown in formula (2); then, calculate the angle between the two vectors, as shown in the following formula:

[0077]

[0078] Since the human body local posture descriptor and the human body local limb motion descriptor are obtained by aggregating the information of adjacent joints, they are combined together to form a joint dependency relationship; the joint dependency relationship corresponds to a specific joint. In order to distinguish it from the joint position, the vector representing the joint dependency relationship is called HyperJoint. The calculation of HyperJoint is as follows:

[0079]

[0080] The constructed HyperJoint is significantly different from the joints in the skeleton data. The skeleton sequence constructed by HyperJoint is regarded as a function The original skeleton sequence only represents the changes in the positions of the human skeleton joints in the time dimension t=1,2,...,T, while HyperJoint aggregates the information from adjacent points to represent the local posture direction and local limb movement amplitude.

[0081] like Figure 4 As shown in the figure, the information carried by ordinary joint points is used to describe the position of the joint in three-dimensional space, that is, the joint information does not include the multi-dimensional association relationship between the joint points; the relationship between objects in reality is often a complex multi-dimensional association relationship; therefore, if the unary relationship carried by ordinary joint points is converted into a multi-dimensional relationship, a lot of useful information will be generated; the difference between HyperJoint and ordinary joint points is not only the different dimensions of the joint points. Compared with ordinary joint points, HyperJoint more accurately describes the relationship between multi-dimensional objects that are associated.

[0082] S21, constructing super joints according to the dependency relationship of common joints, and constructing skeleton information according to the super joints and common joints;

[0083] The skeleton graph of the human body is represented by the variable G(V,H), where V is the set of spatial nodes (ordinary joints) that constitute the human skeleton graph, and H is the dependency relationship (super joint) established on the human skeleton joints; the node information V and the dependency relationship H are represented by the following formula:

[0084]

[0085]

[0086] The dynamic information of a skeleton is defined as the difference between the skeleton elements on adjacent skeletons in the skeleton sequence, which is used to establish a relationship between the skeleton postures in the time domain. The formula is as follows:

[0087]

[0088]

[0089] S22, structured spatiotemporal feature learning model for skeleton data;

[0090] In the natural environment, human behavior is dynamic and continuous; however, in the acquisition device, the complete human behavior is divided into a series of human postures on the time axis, and the human postures can be further decomposed into human skeleton points in three-dimensional space. Since a continuous action shows a hierarchical structure from local to global in the process of formation, the spatiotemporal feature learning module of the convolutional neural network is used to extract human action features on the skeleton data;

[0091] The feature learning module consists of two main functional blocks, namely the node temporal feature learning block and the spatial global feature learning block. The node temporal feature learning block processes the temporal features of the skeleton nodes, while the spatial global feature learning block processes the spatial information distribution features of the skeleton nodes. By superimposing the spatiotemporal blocks, the feature learning module is constructed to learn the effective spatiotemporal features of the dependency relationship between the original skeleton nodes and the super-joint nodes from the skeleton data.

[0092] Furthermore, the temporal relationship of the elements on the human skeleton is learned through the node temporal feature learning block (JTFLB); the node temporal feature learning block is as follows Figure 5 As shown, the node temporal feature learning block decomposes the human skeleton into a series of basic components. Each component has its own motion trajectory when moving. These trajectories have semantic features in the time domain. For the node temporal feature learning block, the input data is a third-order tensor Among them, N represents the number of elements that constitute the human skeleton, C represents the number of channels composed of element dimensions, and T represents the number of frames of the input action.

[0093] First, the convolutional block attention module is used to increase the attention representation of the tensor of the input network to improve the feature expression ability of the input data; the convolutional attention module consists of a channel attention module and a spatial attention module; then a convolutional layer with a convolution kernel size of (1×1) is used to automatically learn the position features of the skeleton joints in the network; the formula is defined as:

[0094] f T (X) = σ(φ(Attn(X))) (9)

[0095] Where X represents a third-order tensor φ represents the function of the convolutional layer structure, σ represents the ReLU activation function, and Attn represents an attention block.

[0096] Finally, a convolution kernel of size (3×1) is used to aggregate the feature information learned from the skeleton joint positions in time, and the output is a third-order tensor It is defined by the following formula:

[0097] A=φ(fT(X)) (10)

[0098] Furthermore, the spatial global feature learning block (SGFLB) is used to learn the relationship between the elements on the human skeleton in motion, such as Figure 6 As shown in the figure, the spatial global feature learning block involves different parts of human body movement in spatial learning, because the integrity and coordination of human body movement have spatial semantic characteristics; for the spatial global feature learning block, the input data is a third-order tensor Among them, N is the number of elements that make up the human skeleton, C represents the number of channels, and T represents the number of frames of the input action;

[0099] The convolution kernel of size (3×3) is used to automatically learn the semantic features of the coordinated motion of skeleton nodes in the spatial domain on the network. The convolution operation can aggregate the spatial structural information of the elements that constitute the human body, which is formally defined as:

[0100] f S (A) = φ(A') (11)

[0101] Among them, A' represents the transformation of the matrix, which transforms a third-order tensor Transformed into

[0102] S3, adding the classification prediction scores of the DMMS stream and the skeleton stream to generate the final prediction score;

[0103] The effect of the length of the skeleton sequence on the recognition rate is studied. Figure 7 The middle one is the frame number statistical histogram using two scale data sets; Figure 7 (a) is the frame number statistical histogram of the UTD-MHAD dataset. Figure 7 It can be seen from (a) that the frame number distribution interval of the UTD-MHAD dataset is [40,125]; Figure 7 (b) is the main frame statistics histogram of the NTU RGB+D dataset. Figure 7 (b) shows that the main frame number distribution range of the NTU RGB+D dataset is [25,200]. In order to effectively explore the impact of frame rate on the experimental results, the frame number is divided into three levels: 32, 64 and 128. Since UTD-MHAD is a small-scale dataset, an additional experiment with a frame number of 16 is added. Table 1 shows the experimental results. In this table, we can clearly see the changes in recognition rate corresponding to the selection of skeleton sequences of different lengths.

[0104] Table 1 Effect of frame number on classification performance

[0105]

[0106] As the number of frames increases, the recognition rates on the two datasets change significantly. In the UTD-MHAD dataset, when the number of frames is 16, the recognition rate on CS is only 43.35%. As the number of frames increases to 32, the recognition rate increases to 77.67%. When the number of frames is less than 32, the information provided to the model in the frame interval with a central frame number of 16 is not enough to express the entire action. When the number of frames increases from 32 to 64, the recognition rate drops to 76.74%. The action frames added in this interval fail to provide the model with discriminative information. On the contrary, the redundancy caused by the increase in action frames leads to a decrease in accuracy. As the number of frames increases from 64 to 128, the recognition rate rises to 91.16%. Therefore, in the UTD-MHAD dataset, due to the large redundancy in the middle part of the action sequence, obtaining enough frames will make the model more discriminative.

[0107] In the NTU RGB+D dataset, the recognition rates of the two experimental strategies also increase with the increase of the number of frames, so a higher number of frames does bring more discriminative information to the model; to ensure the consistency of the model parameters, the number of frames sampled from the two datasets is 128 frames to verify the robustness of the method of the present invention on large-scale data.

[0108] The effect of a single modality in classification is limited, so the deep sequence and the skeleton sequence are fused to improve the recognition accuracy. Table 2 shows the recognition accuracy before and after fusion. The recognition accuracy of the method of the present invention is significantly improved in both data sets. In order to analyze the improvement effect of fusion deep features, Figure 8-10 The changes in accuracy of the method of the present invention on specific actions of two data sets are listed in FIG.

[0109] Figure 8 In the experiment, the recognition accuracy of each action category in the UTD-MHAD dataset was improved by the method of the present invention. Among them, the action numbered 6 was improved the most after the fusion of deep features, and its recognition accuracy was improved by nearly 30%. Among the actions with improved recognition accuracy, the recognition accuracy of most actions was improved by about 5%; Figure 8 It was found that the method of the present invention has very good recognition accuracy for actions numbered 7, 12, 13, 17, 18, 22, 24, 25 and 26.

[0110] The method of the present invention also has a good effect on the NTU RGB+D dataset. Fig. 9 The middle is the result of the experimental strategy CS. The recognition accuracy of the model of the present invention for actions numbered 1, 10, 11, 12, 13, 16, 17, and 20 is significantly improved; Fig.10The middle is the result of the experimental strategy CV. The fused model has significantly improved the recognition accuracy of actions numbered 4, 11, 12, 28, 29, and 30. Overall, the accuracy of the fused model on CS is improved by 6.76%, and the accuracy on CV is improved by 3.28%.

[0111] Table 2 Multimodal performance

[0112]

[0113] The method of the present invention is compared with the prior art method on the UTD-MHAD data set, and the results are shown in Table 3:

[0114] Table 3 Comparison with existing technical methods on the UTD-MHAD dataset

[0115]

[0116]

[0117] The results of the method of the present invention and other methods on UTD-MHAD are shown in Table 3. The recognition rate of the method of the present invention is the highest. The other methods in Table 3 are divided into methods based on manual features and methods based on deep learning. Compared with the BayesianGC-LSTM method, the accuracy is improved by 4.64%. Compared with the network based on skeleton edges, the accuracy is improved by 1.15%.

[0118] The present invention establishes explicit dependency relationships between independent joints and adjacent points. This feature can be used to describe the local posture and motion amplitude during human behavior. In order to better learn the semantic features of skeleton sequences in space and time, we design two functional modules, JTFLB and SGFLB. JTFLB learns the temporal features of skeleton nodes, and SGFLB aggregates the temporal features learned from skeleton nodes in the spatial domain.

[0119] Based on the above ideal embodiments of the present invention, the relevant staff can make various changes and modifications without departing from the technical concept of the present invention through the above description. The technical scope of the present invention is not limited to the contents of the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A super joint and multimodal network and a behavior recognition method thereof, characterized in that: The following steps are involved: S1 collects the human body depth map, uses the DMMS stream to extract features from the depth map, and calculates the depth data prediction score; S2. Collect human skeleton sequence, extract original joint and super joint data respectively, construct skeleton information by combining super joint and common joint, send skeleton information into structured spatiotemporal feature learning model, obtain static and dynamic joint data stream, static and dynamic super joint data stream respectively, and adaptively weight the original joint and super joint data stream to obtain joint data prediction score and super joint prediction score; The step S2 comprises: S21, constructing super joints according to the dependency relationship of common joints, and constructing skeleton information according to the super joints and common joints; The super joint construction includes: First, calculate RWE and RWH, which are the direction vectors of RW pointing to RE and RH respectively. The calculation formula is as follows: Among them, RW, RE and RH represent the right wrist, right elbow and right hand in the human skeleton respectively; Then, calculate the Cartesian product of the two intersecting vectors to get the normal vector n of the plane where the vectors are located. The calculation formula is as follows: Then, taking the direction vectors of the two bones as the starting point RE, calculate the angle between the two direction vectors. The calculation formula is as follows: Finally, we get the hyperjoint data vector HyperJoint, and the calculation formula is: The skeleton information constructed by combining super joints and ordinary joints includes: The skeleton diagram of the human body is represented by the variable G(V,H), and the formula is: Among them, V is the set of spatial nodes that constitute the human skeleton graph, and H is the dependency relationship established on the joints of the human skeleton; The dynamic information of a skeleton is defined as the difference between the skeleton elements on adjacent skeletons in the skeleton sequence, which is used to establish a relationship between the skeleton postures in the time domain. The formula is as follows: S22, feeding the skeleton information into the structured spatiotemporal feature learning model; S3. Add the classification prediction scores of the DMMS stream and the skeleton stream to generate the final prediction score.

2. The hyper-joint and multimodal network and the behavior recognition method thereof according to claim 1, characterized in that: The sending of skeleton information into a structured spatiotemporal feature learning model includes: utilizing a spatiotemporal feature learning module of a convolutional neural network to extract human motion features based on skeleton information.

3. The hyper-joint and multimodal network and the behavior recognition method thereof according to claim 2, characterized in that: The spatiotemporal feature learning module includes a node temporal feature learning block and a spatial global feature learning block. By superimposing the spatiotemporal blocks, a feature learning module is constructed to learn effective spatiotemporal features of the dependency relationship between original skeleton nodes from skeleton information.

4. The hyper-joint and multimodal network and the behavior recognition method thereof according to claim 3, characterized in that: The construction of the node temporal feature learning block includes the following steps: First, the convolutional block attention module is used to increase attention to the tensor of the input network, which means improving the feature expression ability of the input data; Then, a convolutional layer with a convolution kernel size of (1×1) is used in the network to automatically learn the position features of the skeleton joints; the formulation is defined as: f T (X)=σ(φ(Attn(X))) (9) Where X represents a third-order tensor φ represents the function of the convolutional layer structure, σ represents the ReLU activation function, and Attn represents an attention block; Finally, a convolution kernel of size (3×1) is used to aggregate the feature information learned at the skeleton joint positions in time, and the output is a third-order tensor It is defined by the following formula: A=φ(f T (X)) (10).

5. The hyper-joint and multimodal network and the behavior recognition method thereof according to claim 1, characterized in that: The construction of the spatial global feature learning block includes: using a convolution kernel of size (3×3) to automatically learn the semantic features of the coordinated motion of skeleton nodes in the spatial domain on the network. The convolution operation can aggregate the spatial structural information of the elements that constitute the human body, which is formally defined as: f S (A)=φ(A') (11) Among them, A' represents the transformation of the matrix, which transforms a third-order tensor Transformed into

Citation Information

Patent Citations

  • User identity recognition method and system in combination with user gait information

    CN112101176A

  • Human body behavior recognition method based on multi-scale attention map convolutional network

    CN113343901A