A Human Behavior Recognition Method Based on the Motion Coordination Space
By introducing the fusion of motion state measurement coefficients and deep motion maps in human behavior recognition, the problem of difficult to reflect the integrity and coordination of human movements in the prior art is solved, and a more efficient behavior recognition effect is achieved.
Patent Information
- Application Number
- CN202210224741.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-03-09
AI Technical Summary
When the existing human behavior recognition methods use bone data, it is difficult to effectively reflect the integrity and coordination of human movements, and the recognition rate is low.
By introducing motion state measurement coefficients to measure the degree of motion contribution of each joint node, the characteristics of motion co-spatial are extracted, and the deep motion map are fused, and feature extraction and fusion are used using a small convolutional neural network VGG-16.
It improves the accuracy and efficiency of human behavior recognition, and eliminates redundant information through multimodal fusion, providing a more complete behavior recognition method system.
Smart Images

Figure CN114677621B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for human behavior recognition, and particularly to a method for human behavior recognition based on a motion coordination space. Background Art
[0002] Human behavior recognition is one of the research hotspots in the field of computer vision, and many research results have been widely applied in the fields of image analysis, human-computer interaction, intelligent monitoring, video retrieval, motion sensing games, and health detection.
[0003] The related research on behavior recognition can be traced back to an experiment by Johansson in 1973 (GUNNAR JOHANSSON. (1973) Visual perception of biological motion and a model for its analysis In Perception & Psychophysics 1973. Vol. 14. No. 2. 201·211), which described human motion by using the movement of 10-12 key human body nodes to recognize human behavior. In the early stage, behavior recognition research was mainly based on RGB video sequences, but restricted by factors such as illumination, perspective, and background, the behavior recognition based on image videos had certain limitations. With the development of imaging technology, especially the introduction of depth cameras, the research object of human behavior recognition has also started to develop from the initial RGB images to depth images. Compared with the previous RGB images, the depth map sequences collected by structured light depth sensors are insensitive to illumination changes and provide depth data of human behavior.
[0004] Since bone data overcomes the influence of factors such as lighting and background, it can accurately provide the coordinates of human joint points and more directly describe human behavior. Sensors such as Microsoft Kinect and some advanced human pose estimation algorithms make it easier for us to obtain accurate 3D bone data. Using skeleton sequences can well overcome the influence of appearance factors, and has the advantages of clear and simple features and strong spatial information correlation. Therefore, bone data has received more and more attention from researchers in the field of human behavior recognition and detection, and the application of bone data is becoming more and more extensive: Shotton et al. (Shotton Jamie, Sharp Toby, Kipman Alex, et al. Real-time human pose recognition in parts from single depth images[J]. Communications of the Association for Computing Machinery, 2013, 56(1): 116-124) proposed a new method to predict the positions of human joint points from depth images. Lv et al. (Lv, F. and Nevatia, R. (2006) Recognition and segmentation of 3-d human action using hmm and multi-class adaboost. In Euro-pean Conference on Computer Vision, pp. 359–372. Springer.) introduced an action recognition system based on bone joints by decomposing the high-dimensional three-dimensional joint space into a set of feature spaces. Each feature is associated with the combined movement of a single joint or corresponding multiple joints. In addition, Xia et al. (Xia, L., Chen, C.-C. and Aggarwal, J.K. (2012) View invariant human action recognition using histograms of 3d joints. In 2012 IEEE Computer Society Conf. Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 20–27. IEEE.) also proposed HOJ3D features to characterize various human behaviors.In addition, Yang et al. (Yang, X. and Tian, Y. L. (2012) Eigenjoints-based action recognition using naive-bayes-nearest-neighbor. In 2012 IEEE Computer Society Conf. Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 14–19. IEEE) represented human actions using the positions of human skeletal joints, the temporal displacements of the joints, and the offsets of the joints relative to the initial frame of the human skeleton. However, these recognition methods are still relatively single, and for the application of skeletal data, they cannot well reflect the integrity and coordination of human actions. Summary of the Invention
[0005] Object of the Invention: The object of the present invention is to provide a human behavior recognition method that can improve the recognition rate by using a motion state measurement coefficient to judge the contribution degree of each bone point to human motion, and then extracting motion coordination spatial features from the processed key frame sequence.
[0006] Technical Solution: The human behavior recognition method of the present invention includes the following steps:
[0007] S1, respectively perform key frame extraction based on the motion state measurement coefficient for the initial bone sequence and the depth map sequence;
[0008] S2, for the bone sequence processed in step S1, extract the motion coordination spatial vector and splice it into the motion coordination spatial feature; for the depth map sequence processed in step S1, extract the DMM feature to obtain the depth motion map;
[0009] S3, simultaneously input the depth motion map and the motion coordination spatial feature into the deep network for score fusion.
[0010] Further, in step S2, when extracting the motion coordination spatial vector, the vector calculation principle for each joint is as follows: according to the motion change amplitude of each joint point, each joint point is multiplied by the corresponding motion state measurement coefficient of each joint point and then added together.
[0011] Further, the implementation steps of the corresponding motion state measurement coefficient for each joint point are as follows:
[0012] S21, taking the Spine point as the origin coordinate, the area of the triangle formed by two spatial vectors pointing to the same bone point between two adjacent frames of images is S, and the included angle is θ;
[0013] When θ ∈ (0, 90°]:
[0014]
[0015] When the angle change is greater than 90°, it indicates that the movement range of this part is larger. To ensure that S continues to be positively correlated with the movement range, when θ ∈ (90°, 180°]:
[0016]
[0017] Among them, joint_1 and joint_2 respectively represent the spatial vectors describing the joint movement states in two consecutive frames of images;
[0018] S22, project the two spatial vectors onto the XOY plane, YOZ plane, and XOZ plane respectively, and then calculate the areas S XOY 、S YOZ 、S XOZ ; Process the n-frame images in the LeftArm, RightArm, LeftLeg, and RightLeg regions of a complete skeleton graph action sequence in sequence:
[0019]
[0020] S23, then normalize the values within the same region to obtain the motion state measurement coefficient W joint :
[0021]
[0022] Furthermore, when extracting the key frames of the motion state measurement coefficients corresponding to each joint point, according to the difference in the action changes between two adjacent frames of images, one frame is extracted from the two frames of images, and the remaining other frame of image is used to represent the two adjacent frames of images;
[0023] The principle for judging the difference in action changes between two adjacent frames is as follows: Add up the motion state measurement coefficients of all parts between two adjacent frames of images. The larger the obtained value, the greater the difference in actions between the two adjacent frames, and vice versa, it indicates that the difference is smaller; The implementation steps are as follows:
[0024] S031: Calculate the sum of the motion state measurement coefficients of all regions between two adjacent frames of images in sequence to obtain a set of n - 1 motion state measurement coefficients:
[0025] {W1, W2, W3…W n-1};
[0026] S032: Add two adjacent motion state measurement coefficients to obtain the action change parameter C for measuring the key frame i :
[0027] C i =Wi +W i+1 where \(i\in(1,n - 2)\);
[0028] S033: Sort \(C\) i in ascending order. If the \(i\)-th item is the minimum value, then delete the \((i + 1)\)-th frame image:
[0029] C delete =\(\min\{C_1,C_2,C_3,\cdots,C\) n-2 \}\);
[0030] Repeat steps S031 - S033 until a key frame sequence that satisfies the description of human behavior is obtained.
[0031] Furthermore, in step S3, the deep network selects the small convolutional neural network VGG - 16. A total of four convolutional neural networks are trained. One is used to extract skeletal features for motion collaborative spatial features, and the other three are respectively used to extract depth features from the three views of DMM. Finally, the weighted fusion method and the product method are used to fuse the scores obtained from training.
[0032] Compared with the prior art, the present invention has the following remarkable effects:
[0033] 1. Aiming at the characteristics of the integrity and coordination of human motion, based on human skeletal data, a motion collaborative spatial feature model that can synthesize information of each joint point is proposed;
[0034] 2. A new quantization standard, the motion state measurement coefficient, is proposed for the contribution degree of each joint point to the motion; based on the motion state measurement coefficient, a new key frame extraction method is proposed, which removes redundant data and improves the calculation efficiency;
[0035] 3. Based on the idea of multi - modal fusion, skeletal data and depth data are combined to make the data more complete, realize the complementarity of multiple heterogeneous information, eliminate the redundancy between modalities, and establish a new behavior recognition method system, providing new ideas and theoretical basis for the research and application of human behavior recognition methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the overall framework schematic diagram of the present invention;
[0037] Figure 2 is the human body area diagram of the present invention;
[0038] Figure 3 is the left upper limb area diagram of the present invention;
[0039] Figure 4 is the schematic diagram of the motion state measurement coefficient of the present invention;
[0040] Figure 5 This is the VGG-16 network structure diagram of the present invention;
[0041] Figure 6 This is the effect diagram of extracting key frames of the present invention. Detailed implementation manners
[0042] The present invention will be further described in detail below in conjunction with the accompanying drawings of the specification and the detailed implementation manners.
[0043] The present invention uses a motion state measurement coefficient to describe the motion changes of each part of the human body in each action, combines the individual bone data into a comprehensive vector, removes redundant data through key frames, reduces the calculation amount, and then fuses with depth features to achieve a better recognition effect.
[0044] The present invention extracts key frames based on the motion state measurement coefficient for the initial bone sequence and depth map sequence, then extracts the motion collaborative space vector from the processed bone sequence and splices it into the motion collaborative space feature; extracts the DMM feature from the processed depth map sequence. Finally, fuses the DMM (Depth Motion Map) feature based on depth data and the SFMI feature based on bone data, providing a method for recognizing human behaviors using multi-modalities. The overall framework is as Figure 1 shown.
[0045] (I) Motion collaborative space feature
[0046] When a human body performs various actions, the motion conditions of each part of the body are not the same. Most of the existing behavior recognitions based on bone data separately use the information of each bone node, while the joint parts in the human body movement process are coordinated and interact with each other, and the human body movement has the characteristics of integrity and coordination. In order to better reflect these characteristics, the present invention proposes a motion collaborative space vector.
[0047] (11) Determine the motion collaborative space vector
[0048] During the movement process of the human body, the torso part is often relatively stable, while the limbs part is relatively active. That is, between two adjacent frames of images, the changes in the limbs part of a person are often greater. Therefore, compared with the torso part that is prone to generating duplicate information, the spatio-temporal information generated by the movement of the limbs part is more helpful for behavior recognition.
[0049] Therefore, the present invention divides the human body into four regions: LeftArm, RightArm, LeftLeg, and RightLeg, as Figure 2 shown. Calculate the motion collaborative space vector representing the limb within each region to describe the limb movement and reflect the integrity and coordination.
[0050] Taking the LeftArm area as an example, the specific description of the motion coordination space vector is as follows Figure 3 As shown, there are four skeletal points in the left upper limb: LeftShoulder (left shoulder), LeftElbow (left elbow), LeftWrist (left wrist), and LeftHand (left hand). Taking the Spine point as the center, the states of each joint point can be represented by the left shoulder vector left elbow vector left wrist vector left hand vector to describe. A certain movement of the left upper limb is jointly determined by the overall action and coordination of these four vectors.
[0051] Simply splicing the vectors directly can reflect the movements of each skeletal point, but it cannot accurately reflect the contribution degree of each joint point to the overall movement state. For example, in the left hand waving action, only the movement of the left upper limb changes significantly, and among the four skeletal points of the left upper limb, the movement amplitude of the left hand (LeftHand) is also larger than that of the left shoulder (LeftShoulder). To better identify human movements, the present invention combines the movement states of each part and proposes a motion coordination space vector in order to more accurately reflect the integrity and coordination of the movement.
[0052] Still taking the left upper limb as an example, according to the movement change amplitude of each joint point, corresponding weights are given, that is, multiply by the corresponding W (movement state measurement coefficient) of each joint point and then add them together. The description of the movement state of the left upper limb is more accurate, and the expression is:
[0053]
[0054] Among them, W LS 、W LE 、W LW 、W LH are the movement state measurement coefficients of the left shoulder, left elbow, left wrist, and left hand parts respectively.
[0055] Obtained according to formula (1), respectively describing the (left upper limb area motion coordination space vector), (right upper limb area motion coordination space vector), (left lower limb area motion coordination space vector), (right lower limb area motion coordination space vector), and then according to the corresponding movement state measurement coefficients of each motion coordination space vector, they are spliced into the motion coordination space feature (Spatial Features of MotionCoordination), and the expression is:
[0056]
[0057] (12) Determine the motion state measurement coefficient
[0058] Since the change amplitudes of different parts of the human body are different in different actions, in terms of the vectors of each joint point, the motion states of each part can be evaluated from two aspects: the vector length and the rotation angle. The greater the change amplitudes of the vector length and the rotation angle, the more intense the movement of that part and the higher the contribution to the motion state.
[0059] The area enclosed between two vectors formed by the same joint point at different times can take into account both the length and the angle attributes and can well describe the motion change amplitude of this node. Based on this idea, the present invention proposes a motion state measurement coefficient W.
[0060] Taking two adjacent frames of images in the LeftArm area as an example, as Figure 4 shown, respectively represent the spatial vectors describing the motion state of the left hand in the front and back frames of images. Taking the Spine point as the origin coordinate, the area of the triangle formed by two spatial vectors pointing to the same bone point between two adjacent frames of images is S, and the included angle is θ.
[0061] When θ ∈ (0, 90°]:
[0062]
[0063] When the angle change is greater than 90°, it indicates that the motion amplitude of this part is larger. In order to keep S positively correlated with the motion amplitude, when θ ∈ (90°, 180°]:
[0064]
[0065] To better retain the spatial information of the motion state measurement coefficient, the two spatial vectors can be projected onto the XOY plane, YOZ plane, and XOZ plane respectively and then S can be calculated respectively. Process the n frames of images in the LeftArm area of a complete skeleton map action sequence in sequence:
[0066]
[0067]
[0068]
[0069]
[0070] Among them, S LS 、S LE 、S LW 、S LHThey are respectively: the spatial area enclosed by the spatial vectors of two adjacent frames of images of the left shoulder, left elbow, left wrist, and left hand; S XOY 、S YOZ 、S XOZ respectively represent the areas after the spatial vectors are projected onto the XOY plane, YOZ plane, and XOZ plane.
[0071] After normalizing the values in the same region, the motion state measurement coefficient W can be obtained:
[0072]
[0073]
[0074]
[0075]
[0076] After obtaining the motion coordination spatial vectors in sequence according to the time series, the motion state measurement coefficients W representing each region can be obtained in sequence. Among them, W LS 、W LE 、W LW 、W LH are respectively the motion state measurement coefficients of the left shoulder, left elbow, left wrist, and left hand.
[0077] (2) Key frame extraction for the motion state measurement coefficient
[0078] A person's actions cannot be static. Moreover, for a complete action sequence, the distribution of a person's actions on each frame of the image is also uneven, and there will be a situation where the action changes less between two adjacent frames of images. At this time, one frame can be extracted from these two frames, and the remaining frame can be used to represent these two frames. In this way, redundant information can be simplified, the calculation amount can be reduced, and the recognition efficiency can be improved.
[0079] When extracting key frames, the motion state measurement coefficient can be used as a good measurement standard. Adding up the motion state measurement coefficients of all parts between two adjacent frames of images, the obtained value can accurately reflect the action change between the two frames of images. The larger the value, the greater the difference in actions between two adjacent frames, and vice versa.
[0080] The specific method steps are as follows:
[0081] Input: The joint point coordinates of the skeleton map sequence, and this skeleton map sequence has n frames of images.
[0082] Output: The key frame sequence based on the motion state measurement coefficient.
[0083]
[0084] Step 1: Calculate the sum of the motion state measurement coefficients for all regions between two adjacent frames in sequence, and a set of n - 1 motion state measurement coefficients can be obtained:
[0085] {W1, W2, W3…W n-1} (13)
[0086] W1 represents the sum of the motion state measurement coefficients for all regions between the first and second frames, W2 represents the sum of the motion state measurement coefficients for all regions between the second and third frames, and so on, until W n-1
[0087] Step 2: Add two adjacent motion state measurement coefficients to obtain the action change parameter C for measuring the key frame:
[0088] C i =W i +W i+1 , i ∈ (1, n - 2) (14)
[0089] Step 3: Sort C by size. If the i-th item is the minimum value, then delete the (i + 1)-th frame image:
[0090] C delete =min{C1, C2, C3…C n-2} (15)
[0091] Repeat the above steps Step 1 - Step 3 until a key frame sequence sufficient to describe human behavior is obtained.
[0092] (III) VGG - 16 Convolutional Neural Network
[0093] To address the problem of the small number of samples in the UTD - MHAD dataset, a small convolutional neural network VGG - 16 (refer to Krizhevsky A, Sutskever I, Hinton G E. Imagenet classification with deep convolutional neural networks[C] / / International Conference on Neural Information Processing Systems, 2012: 1106 - 1114.) is used to extract the motion collaborative spatial features and the feature information of DMM (Depth Motion Maps). And because the number of dataset samples is small, overfitting is likely to occur when retraining from scratch. Therefore, training is carried out based on the pre - trained model of ILSVRC - 2012, and the network parameters are fine - tuned.
[0094] The VGG-16 network structure is as follows Figure 5 shown. By repeatedly stacking small 3*3 convolutional kernels and 2*2 max-pooling kernels, the network structure is continuously deepened to improve performance.
[0095] In the present invention, a total of four convolutional neural networks are trained. One is used to extract skeletal features of motion collaborative spatial features, and the other three are respectively used to extract depth features from the three views of the DMM. Finally, the weighted fusion method and the product method are used to fuse the scores obtained from the training.
[0096] (4) Experimental verification
[0097] The experiment was run on a Lenovo Y7000P laptop with Windows 10 system, an i7-10750H CPU with 2.60 GHz, 16.00 GB of installed memory, Matlab R2018b version, and Python 3.7.0 version.
[0098] (41) Experimental data
[0099] The present invention is tested on the multi-modal action dataset UTD-MHAD (see Chen, C.; Jafari, R.; Kehtarnavaz, N. Utd-mhad: A multimodal dataset for human action recognition utilizing a depth camera and a wearable inertial sensor. In Proceedings of the 2015 IEEE International Conference on Image Processing (ICIP), Quebec City, QC, Canada, 27–30 September 2015.). This dataset consists of 27 different actions: (1) Slide right arm left, (2) Slide right arm right, (3) Wave right hand, (4) Clap hands in front, (5) Throw with right arm, (6) Cross arms at chest, (7) Basketball shoot, (8) Draw an x with right hand, (9) Draw a circle with right hand (clockwise), (10) Draw a circle with right hand (counterclockwise), (11) Draw a triangle, (12) Bowling (right hand), (13) Front punch, (14) Baseball right swing, (15) Tennis right hand forehand swing, (16) Spin arms (both arms), (17) Tennis serve, (18) Push with both hands, (19) Knock on door with right hand, (20) Grab object with right hand, (21) Pick up and throw with right hand, (22) Jog in place, (23) Walk in place, (24) Sit to stand, (25) Stand to sit, (26) Forward lunge (left foot forward), (27) Squat (arms straight). The inertial sensor is worn on the right wrist or right thigh of the subject, depending on whether the action is mainly of the arm or leg type. Specifically, for actions 1 to 21, the inertial sensor is placed on the right wrist of the subject; for actions 22 to 27, the inertial sensor is placed on the right thigh of the subject.
[0100] The UTD-MHAD dataset was collected in an indoor environment using a Microsoft Kinect sensor and a wearable inertial sensor. The dataset contains 27 actions performed by 8 subjects (4 females and 4 males). Each subject repeated each action 4 times. After deleting three corrupted sequences, the dataset includes 861 data sequences. Four data modalities of RGB video, depth video, skeletal joint positions, and inertial sensor signals are recorded in three channels or threads. One channel is used to capture the depth video and skeletal positions simultaneously, one channel is for the RGB video, and one channel is for the inertial sensor signals (triaxial acceleration and triaxial rotation signals). To synchronize the data, the timestamp of each sample is recorded.
[0101] (42) Experimental setup
[0102] In the present invention, two different settings are adopted for the UTD-MHAD dataset.
[0103] Setting 1:
[0104] In order to evaluate the performance characteristics of the model according to the change of the training dataset size, three experiments are conducted on the samples in the UTD-MHAD dataset:
[0105] In Test1, 1 / 4 of the data is used as the training set and 3 / 4 of the data is used as the test set;
[0106] In Test2, 1 / 2 of the data is used as the training set and 1 / 2 of the data is used as the test set;
[0107] In Test3, 3 / 4 of the data is used as the training set and 1 / 4 of the data is used as the test set.
[0108] Setting 2:
[0109] All samples in the UTD-MHAD are classified simultaneously. The samples of subjects 1, 3, 5, 7 are used for training, and the samples of subjects 2, 4, 6, 8 are used for testing.
[0110] In the present invention, the VGG-16 network architecture CNN is used for classification recognition, and the pre-trained model ILSVRC-2012 is used. The network parameter settings refer to the settings in the literature (Krizhevsky A, Sutskever I, Hinton G E. Imagenet classification with deep convolutional neural networks[C] / / International Conference on Neural Information Processing Systems, 2012: 1106 - 1114.): the learning rate is set to 0.001, the batch size is set to 32; the maximum number of training iterations is set to 20,000 times, and the learning rate is decreased once every 5000 iterations.
[0111] (43) Experimental Results and Analysis
[0112] (431) Experimental Results of the Motion Coordination Spatial Feature Model
[0113] To evaluate the effects of different methods for extracting the spatial features of motion coordination, three different models were established and compared according to the parameters of Setting 1: Model 1 is the spatial features of motion coordination after projection onto three Cartesian planes without using the motion state measurement coefficient; Model 2 is the spatial features of motion coordination that directly calculates the motion state measurement coefficient without projecting the motion coordination spatial vector onto the three Cartesian planes; Model 3 is the spatial features of motion coordination after being processed with the motion state measurement coefficient and projected onto the three Cartesian planes. Model 1 and Model 3, Model 2 and Model 3 are two groups of control groups.
[0114] Table 1 Comparison of Models for Spatial Features of Motion Coordination
[0115]
[0116] As can be seen from Table 1, the average recognition rate of Model 3 increased by 3.3% compared to Model 1. The introduction of the motion state measurement coefficient enables the model to have a clear discrimination criterion for the contributions of various parts to the motion, provides a weight ratio discrimination basis for the extraction and splicing of the motion coordination spatial vectors in each region, and improves the recognition effect.
[0117] The average recognition rate of Model 3 increased by 31.1% compared to Model 2. The processing of projecting onto the three Cartesian planes enhances the model's extraction of spatial information and significantly improves the recognition effect.
[0118] It can be seen from this experiment that both the introduction of the motion state measurement coefficient and the projection onto the three Cartesian planes effectively improve the recognition effect of the spatial features of motion coordination. Therefore, Model 3 will be used to extract the spatial features of motion coordination in the subsequent experiments of the present invention.
[0119] (432) Extraction of Key Frames
[0120] The present invention uses a key frame extraction algorithm based on the motion state measurement coefficient to extract key frames from the sequence of skeleton diagrams. The number of key frames directly affects the performance of the experiment. Therefore, it is necessary to set an appropriate percentage of the remaining frames, which can not only remove redundant information but also completely retain key information.
[0121] The present invention conducted experiments on the key frame extraction ratio according to Setting 1. When the percentage of the remaining frames is set differently, the influence on the recognition rate is as Figure 6 shown.
[0122] From Figure 6 it can be seen that when the key frame ratio is around 70%, the recognition effect reaches the peak, with an average increase of 1.5% compared to the original recognition rate, and the computational amount is reduced and the running time is decreased. Therefore, the key frame ratio for the subsequent experiments in this paper is set to 70%.
[0123] (433) Multi-modal Fusion Experiment Results
[0124] To evaluate the impact of modal fusion on the model performance, according to Setting 2, the results of using skeletal data alone, depth data alone, and after modal fusion were compared, as shown in Table 2:
[0125] Table 2 Comparison of Single-modal and Multi-modal Methods
[0126]
[0127] As can be seen from Table 2, the effect of fusing scores using the product method is better than that using the weighted average method. And the results after modal fusion all reach the best effect. The recognition rate of SFMC-DMM after fusion is 19.6% higher than that of SFMC using skeletal data alone and 1% higher than that of DMM using depth data alone. It proves that the information provided by skeletal data and depth data is complementary, and multi-modal data can describe human behaviors more accurately after fusion.
[0128] (434) Method Comparison
[0129] The present invention was compared with other methods according to Setting 2, and the results are shown in Table 3:
[0130] Table 3 Precision Comparison of Each Method on UTD-MHAD
[0131]
[0132] Reference for MSTM-HOG: Chao Xin, Hou Zhenjie, Liang Jiuzhen, Yang Tianjin. (2020). Integrally Cooperative SpatioTemporal Feature Representation of MotionJoints for Action Recognition. Sensors (Basel, Switzerland). 20.10.3390 / s20185180.
[0133] Kinect, Kinect + Inertial reference: Chen, C.; Jafari, R.; Kehtarnavaz, N. Utd-mhad: A multimodal dataset for human action recognition utilizing a depth camera and a wearable inertial sensor. In Proceedings of the 2015 IEEE International Conference on Image Processing (ICIP), Quebec City, QC, Canada, 27–30 September 2015。
[0134] 3DHoT-MBC reference: Zhang, B. C.; Yang, Y.; Chen, C.; Yang, L. L.; Han, J. G.; Shao, L. Action recognition using 3d histograms of texture and a multi-class boosting classifier. IEEE Trans. Image Process. 2017, 26, 4648–4660。
[0135] SDSR reference: Annadani Y, Rakshith D, Biswas S. Sliding dictionary based sparse representation for action recognition[C] / / Computer Vision and Pattern Recognition, 2016:1-7。
[0136] DMM-CTHOG-LBP-EOH reference: Bulbul M F, Jiang Y, Ma J. Dmms-based multiple features fusion for human action recognition[J]. International Journal of Multimedia Data Engineering and Management, 2015, 6(4): 23-39。
[0137] As can be seen from Table 3, the recognition rate of the method of the present invention on the UTD-MHAD dataset reaches 89.8%, which is higher than other methods listed, and the recognition rate is increased by at least 1.4%. The evaluation results prove the superiority of this method.
[0138] In summary, the present invention provides a behavior recognition algorithm for the integrity and coordination of human motion. By referring to the motion relationship between each bone node and using the motion state measurement coefficient to judge the contribution degree of each bone point to human motion, the motion coordination spatial features are extracted from the processed key frame sequence. After multi-modal fusion with the depth motion map, good recognition results are obtained on the UTD-MHAD dataset.
Claims
1. A human behavior recognition method based on a motion coordination space, characterized in that Including the steps: S1. Extract key frames for the initial bone sequence and depth map sequence respectively based on the motion state measurement coefficient; S2. For the bone sequence processed in step S1, extract the motion coordination space vectors and splice them into motion coordination space features; Extract DMM features from the depth map sequence processed in step S1 to obtain a depth motion map; S3. Input the depth motion map and the motion coordination space features into the depth network simultaneously for score fusion; The implementation steps of the motion state measurement coefficient corresponding to each joint point are as follows: S21. Evaluate the motion state of each part according to the vector length and rotation angle. Taking the Spine point as the origin coordinate, the area of the triangle formed by two spatial vectors pointing to the same bone point between two adjacent frames of images is S, and the included angle is θ; When θ ∈ (0, 90°]: When the angle change is greater than 90°, it means that the motion amplitude of this part is larger. To ensure that S is still positively correlated with the motion amplitude, when θ ∈ (90°, 180°]: Among them, joint_1 and joint_2 respectively represent the spatial vectors describing the joint motion state in the previous and subsequent frames of images; S22. Project the two spatial vectors onto the XOY plane, YOZ plane, and XOZ plane respectively, and then calculate the areas S XOY , S YOZ , S XOZ ; Process the n-frame images in the LeftArm, RightArm, LeftLeg, and RightLeg regions of a complete skeleton map action sequence in sequence: S23. After normalizing the values in the same area, the motion state measurement coefficient W is obtained. joint :
2. The human behavior recognition method based on a motion collaboration space according to claim 1, wherein In step S2, when extracting the motion coordination space vectors, the vector calculation principle for each joint is as follows: According to the motion change amplitude of each joint point, each joint point is multiplied by the corresponding motion state measurement coefficient of each joint point and then added together.
3. The human behavior recognition method based on the motion cooperation space according to claim 2, wherein When extracting the key frames of the motion state measurement coefficient corresponding to each joint point, according to the difference in action changes between two adjacent frames of images, one frame is extracted from the two frames of images, and the remaining other frame of image is used to represent the two adjacent frames of images; The difference judgment principle for the action changes between two adjacent frames is as follows: Add up the motion state measurement coefficients of all parts between two adjacent frames of images. The larger the obtained value, the greater the difference in actions between two adjacent frames, and vice versa, it means the smaller the difference. The implementation steps are as follows: S031: Calculate the sum of the motion state measurement coefficients of all regions between two adjacent frames of images in sequence to obtain a set of n - 1 motion state measurement coefficients: {W1, W2, W3…W n-1}; S032: Add the motion state measurement coefficients of two adjacent items to obtain the action change parameter C for measuring the key frame i : C i = W i + W i+1 , i ∈ (1, n - 2); S033: Sort C i in ascending order. If the i-th item is the minimum value, delete the (i + 1)-th frame image: C delete = min{C1, C2, C3…C n-2}; Repeat steps S031 - S033 until a key frame sequence that satisfies the description of human behavior is obtained.
4. The human behavior recognition method based on the motion cooperation space according to claim 1, characterized in that In step S3, the depth network selects the small convolutional neural network VGG - 16. A total of four convolutional neural networks are trained. One is used to extract bone features from the motion coordination space features, and the other three are respectively used to extract depth features from the three views of the DMM. Finally, the weighted fusion method and the product method are used to fuse the scores obtained from the training.
Citation Information
Patent Citations
Gait recognition method for multi-gait feature combined collaborative dictionary
CN108921062A
Action recognition method based on depth image and skeleton information
CN110263720A