A method, device and equipment for skeleton-based human behavior recognition
By extracting joint and bone features of the skeleton data and stitching and fusion of features, the problem of insufficient use of skeleton data in the prior art is solved, and a higher accuracy of human behavior recognition is achieved.
Patent Information
- Application Number
- CN202111616700.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-27
AI Technical Summary
The existing skeleton human behavior recognition methods do not use the information of skeleton data sufficiently, resulting in the impact of recognition accuracy.
By obtaining the skeleton data of the target object, calculate the joint difference data and bone difference data, and perform feature extraction to obtain the skeleton data characteristics and skeleton difference data characteristics. Then, feature data splicing and fusion are performed to obtain joint splicing features and bone splicing features. Finally, key position features of different dimensions are strengthened and fused with these features to obtain the action classification prediction results.
By extracting the joints and bones in the human skeleton separately, the skeleton data information can be made more fully, noise interference can be reduced, and the accuracy of human behavior recognition can be improved.
Smart Images

Figure CN114582012B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human behavior recognition, and in particular to a skeleton human behavior recognition method, device and equipment. Background Art
[0002] Human behavior recognition is a popular research topic in the field of machine vision, aiming to capture and extract the feature information existing in the space and time of human actions, and further determine the types of human actions based on these feature information. The skeleton human behavior recognition method is a human behavior recognition method that uses the skeleton data directly or indirectly extracted from human actions as input data. Compared with RGB video data, skeleton data has the advantages of being insensitive to environmental interference, high effective information density, and small data storage space.
[0003] The current skeleton human behavior recognition method uses the Microsoft Kinect sensor to directly collect labeled skeleton data or uses the OpenPose algorithm to extract skeleton data from RGB human action videos as input data, and uses recurrent neural network methods, convolutional neural network methods and graph convolutional neural network methods in deep learning to perform human behavior recognition on the skeleton data. However, the existing human behavior recognition methods are the transformation of a single data feature extraction network, or the fusion method of adding the prediction scores obtained based on a single data to fuse the skeleton data and its derivative data. These methods lack sufficient attention to the derivative data and the fusion data, which results in insufficient use of the information of the skeleton data, and further affects the final recognition accuracy. Summary of the Invention
[0004] Therefore, the technical problem to be solved by the present invention is to overcome the defect of insufficient use of the information of the skeleton data in the prior art, so as to provide a skeleton human behavior recognition method, device and equipment.
[0005] According to a first aspect, an embodiment of the present invention provides a skeleton human behavior recognition method, including: obtaining skeleton data of a target object, where the skeleton data includes joint data and bone data; calculating joint difference data and bone difference data based on key points of the skeleton data; performing feature extraction based on the joint difference data and the bone difference data to obtain skeleton data features and skeleton difference data features, and obtaining joint data features and joint difference data features based on the skeleton data features; respectively performing feature data splicing and fusion based on the joint data features and the joint difference data features, and the skeleton data features and the skeleton difference data features to obtain joint splicing features and bone splicing features; respectively performing enhanced fusion of the key position features of different dimension branches with the joint splicing features and the bone splicing features to obtain an action classification prediction result.
[0006] Optionally, the key point calculation based on the skeleton data to obtain joint difference data includes: extracting joint data based on the skeleton data; establishing a joint coordinate system based on the joint data; extracting joint change data within a preset time based on the joint coordinate system, and calculating the difference within the preset time based on the joint change data and the joint data, so as to obtain the joint difference data.
[0007] Optionally, the key point calculation based on the skeleton data to obtain bone difference data includes: extracting bone data based on the skeleton data; establishing a bone coordinate system based on the bone data; extracting bone position change data within a preset time based on the bone coordinate system, and calculating the difference within the preset time based on the bone change data and the bone data, so as to obtain the bone difference data.
[0008] Optionally, the feature extraction based on the bone difference data to obtain skeleton data features and skeleton difference data features includes: constructing a skeleton difference data coordinate system based on the bone difference data and a preset time; extracting a skeleton difference change image within a preset time based on the skeleton difference data coordinate system; obtaining the skeleton data features and the skeleton difference data features based on the skeleton difference change image.
[0009] Optionally, the feature data splicing and fusion based on the joint data features and the joint difference data features to obtain joint splicing features includes: constructing a network layer based on the joint data features and the joint difference data features; sorting the data based on the first network layer to obtain a first sorting result; performing feature data splicing and fusion based on the first sorting result to obtain the joint splicing features.
[0010] Optionally, the feature data splicing and fusion based on the skeleton data features and the skeleton difference data features to obtain bone splicing features includes: constructing a network layer based on the skeleton data features and the skeleton difference data features; sorting the data based on the second network layer to obtain a second sorting result; performing feature data splicing and fusion based on the second sorting result to obtain the bone splicing features.
[0011] Optionally, the key position features of different dimensional branches are respectively and intensively fused with the joint splicing features and the bone splicing features to obtain an action classification prediction result, including: establishing a fusion layer based on the joint splicing features and the bone splicing features; extracting key position feature information based on the fusion layer, where the key position is obtained based on the skeleton data features and the skeleton difference data features; obtaining a prediction value of the skeleton data based on the key position feature information; obtaining the action classification prediction result based on the prediction value.
[0012] According to a second aspect, an embodiment of the present invention provides a skeletal human behavior recognition device, including: an acquisition module, configured to acquire skeletal data of a target object, where the skeletal data includes joint data and bone data; a calculation module, configured to calculate joint difference data and bone difference data based on key points of the skeletal data; a feature extraction module, configured to perform feature extraction based on the joint difference data and the bone difference data to obtain skeletal data features and skeletal difference data features, and obtain joint data features and joint difference data features based on the skeletal data features; a fusion module, configured to perform feature data splicing and fusion based on the joint data features and joint difference data features, and the skeletal data features and skeletal difference data features respectively, to obtain joint splicing features and bone splicing features; and a prediction module, configured to perform enhanced fusion of key position features of different-dimensional branches with the joint splicing features and bone splicing features respectively to obtain an action classification prediction result.
[0013] According to a third aspect, a skeletal human behavior recognition device includes: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the skeletal human behavior recognition method according to any one of the first aspect or any optional manner.
[0014] According to a fourth aspect, a computer-readable storage medium is characterized in that the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the skeletal human behavior recognition method according to the first aspect or any optional implementation manner.
[0015] The technical solution of the present invention has the following advantages:
[0016] An embodiment of the present invention provides a skeletal human behavior recognition method, device and equipment. The method includes the following steps: acquiring skeletal data of a target object, calculating joint difference data and bone difference data based on key points of the skeletal data, extracting feature information based on the joint difference data and the bone difference data to obtain skeletal data features and skeletal difference features, obtaining joint data features and joint difference data features based on the skeletal data features, performing feature data splicing and fusion on the obtained joint data features and joint difference data features, and the skeletal data features and skeletal difference data features respectively to obtain joint splicing features and bone splicing features, and performing enhanced fusion of key position features of different dimensions with the joint splicing features and bone splicing features respectively to obtain an action classification prediction result. By separately extracting features of joints and bones in the human skeleton, the present invention can identify detailed information of the human body, and at the same time eliminate more interference of noise information, making the recognition of human behavior more accurate. Description of the Drawings
[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a specific example flowchart of a skeleton human behavior recognition method according to an embodiment of the present invention;
[0019] Figure 2 It is a schematic diagram of an example of a joint coordinate system of a skeleton human behavior recognition method according to an embodiment of the present invention;
[0020] Figure 3 It is a schematic diagram of an example of joint difference data of a skeleton human behavior recognition method according to an embodiment of the present invention;
[0021] Figure 4 It is a schematic diagram of an example of skeleton difference data of a skeleton human behavior recognition method according to an embodiment of the present invention;
[0022] Figure 5 It is a schematic diagram of an example of a network layer of a skeleton human behavior recognition method according to an embodiment of the present invention;
[0023] Figure 6 It is a schematic diagram of an example of the overall network layer of a skeleton human behavior recognition method according to an embodiment of the present invention;
[0024] Figure 7 It is a schematic diagram of an example of enhanced fusion of a skeleton human behavior recognition method according to an embodiment of the present invention;
[0025] Figure 8 It is a schematic diagram of an example of the structure of a skeleton human behavior recognition device according to an embodiment of the present invention;
[0026] Figure 9 It is a schematic diagram of an example of the connection of a skeleton human behavior recognition device according to an embodiment of the present invention. Specific Embodiments
[0027] The following will clearly and completely describe the technical solutions of the present invention with reference to the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0028] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0029] In the embodiments of the present invention, relevant data of joints and bones are extracted based on the skeleton data of the target object, and then the relevant data of each bone of the joints are fused to obtain the action classification prediction result. In the following embodiments, the human skeleton is taken as an example, and the present invention can also be applied to the action recognition of other skeletons, and the present application is not limited thereto.
[0030] Figure 1 The flowchart of a skeleton human behavior recognition method according to an embodiment of the present invention is shown. The method specifically includes the following steps:
[0031] S100: Obtain the skeleton data of the target object, where the skeleton data includes joint data and bone data.
[0032] Specifically, the action information of the target object is collected by a collection device, and the skeleton data of the target object is extracted based on the collected action information. The skeleton data includes the joint data and bone data of the target object. In practical applications, the collection device can be, for example, a Microsoft Kinect sensor or a video collection device.
[0033] S200: Calculate joint difference data and bone difference data based on the key points of the skeleton data.
[0034] Specifically, by preprocessing the skeleton data, the key points of the skeleton data are obtained. Based on the key points, joint data in the channel dimension, time dimension, and spatial dimension are obtained. Based on the joint data D, data frames are calculated in the time dimension, and the data form is where t 0 is the time frame of the joint data, and C, T, and S respectively represent the channel dimension, time dimension, and spatial dimension of the skeleton data. Based on the joint data, joint difference data generated by the action transformation of the target object within a preset time is obtained. Based on the spatial dimension, bone data is obtained. Based on the bone data and the time dimension, bone difference data generated by the action transformation of the target object within a preset time is obtained.
[0035] S300: Extract features based on the joint difference data and bone difference data to obtain skeleton data features and skeleton difference data features, and obtain joint data features and joint difference data features based on the skeleton data features.
[0036] Specifically, feature extraction is performed based on the joint difference data and the bone difference data. The skeleton data features and the skeleton difference data features are obtained according to the joint change difference features and the bone change difference features of the target object during the movement process. The joint data features and the joint difference data features are extracted from the skeleton data features.
[0037] S400: Feature data splicing and fusion are respectively performed based on the joint data features and the joint difference data features, the skeleton data features and the skeleton difference data features, to obtain joint splicing features and bone splicing features.
[0038] Specifically, feature data splicing and fusion are performed based on the joint data features and the joint difference data features, the skeleton data features and the skeleton difference data features. The joint splicing features and the bone splicing features are obtained according to the fused overall skeleton and joints.
[0039] S500: The key position features of different dimensional branches are respectively and intensively fused with the joint splicing features and the bone splicing features to obtain an action classification prediction result.
[0040] Specifically, the key position features of the skeletons in the channel dimension, the time dimension, and the space dimension are intensively fused again with the joint splicing features and the bone splicing features to obtain a complete skeleton after multi - perspective attention enhancement fusion. Based on the complete skeleton, the action of the target object is analyzed to obtain an action classification prediction result.
[0041] In the embodiment of the present invention, the action information of the target object is collected by a collection device, the skeleton data of the target object is extracted based on the collected action information, the key points of the skeleton data are obtained by pre - processing the skeleton data, different - dimensional joint data are obtained based on the key points, the joint difference data and the bone difference data generated by the action transformation of the target object within a preset time are obtained based on the joint data, feature data splicing and fusion are performed based on the joint data features and the joint difference data features, the skeleton data features and the skeleton difference data features, the joint splicing features and the bone splicing features are obtained according to the fused overall skeleton and joints, the key position features of the skeletons in different dimensions are intensively fused again with the joint splicing features and the bone splicing features to obtain a complete skeleton, and based on the complete skeleton, the action of the target object is analyzed to obtain an action classification prediction result. In the embodiment of the present invention, by collecting the skeleton data of the target object during the action, different - dimensional data feature information is obtained, and the action and behavior of the target object are analyzed thereby, so as to improve the accuracy of the recognition task and reduce the influence of image noise.
[0042] In an optional embodiment of the present invention, the above - mentioned step S200 calculates the joint difference data based on the key points of the skeleton data, including the following steps:
[0043] (1) Extract joint data based on the skeleton data;
[0044] (2) Establish a joint coordinate system based on the joint data;
[0045] (3) Extract joint change data within a preset time based on the joint coordinate system, and calculate the difference within the preset time based on the joint change data and the joint data, so as to obtain joint difference data.
[0046] Specifically, extract joint data according to the skeleton data, and establish a joint coordinate system according to the joint data. The joint coordinate system is as Figure 2 shown. In the joint coordinate system, take the joint points of the target object as the origin, and naturally connect each joint point following the human skeleton. Obtain joint change data within a preset time based on the joint coordinate system, and calculate the difference between the joint change data and the joint data within the preset time to obtain joint difference data.
[0047] Exemplarily, the data of the joint data at time t within a preset time is extracted and denoted as (x t1 , y t2 , z t3 ). Calculate the difference of the same joint node in adjacent time dimensions based on the data, expressed as (x t1 - x t1-1 , u t1 - y t1-1 , z t1 - z t1-1 ). In practical applications, in order to align the data for subsequent calculations, for example, the average value of all data can be set at time 0 of the preset time. The process is as Figure 3 shown, that is
[0048] In the embodiment of the present invention, extract joint data according to the skeleton data, and establish a joint coordinate system according to the joint data. In the joint coordinate system, take the joint points of the target object as the origin, and naturally connect each joint point following the human skeleton. Obtain joint change data within a preset time based on the joint coordinate system, and calculate the difference between the joint change data and the joint data within the preset time to obtain joint difference data. By establishing a joint coordinate system in the embodiment of the present invention, the difference data of each joint within a preset time can be calculated more accurately, further improving the accuracy of identifying the action behavior of the target object.
[0049] In an alternative embodiment of the present invention, the above step S200 calculates the bone difference data based on the key points of the skeleton data, including the following steps:
[0050] (1) Extract bone data based on the skeleton data;
[0051] (2) Establish a bone coordinate system based on the bone data;
[0052] (3) Extract the bone position change data within a preset time based on the bone coordinate system, and calculate the difference within the preset time based on the bone change data and the bone data, so as to obtain the bone difference data.
[0053] Specifically, extract the bone data from the spatial dimension according to the skeleton data, and establish a bone coordinate system according to the bone data. The bone coordinate system is as Figure 2 shown. In the bone coordinate system, take the joint points of the target object as the origin, and naturally connect the joint points following the human skeleton. Each connecting line is the bone coordinate of the target object. Obtain the bone change data within a preset time based on the bone coordinate system, and calculate the difference between the bone change data and the bone data within the preset time based on the time dimension of the skeleton data to obtain the bone difference data.
[0054] Exemplarily, record the data of the bone data at time t within a preset time as (x t2 , y t2 , z t2 ). Calculate the difference in the corresponding bone change under the adjacent time dimension based on the data, expressed as (x t2 - x t2-1 , y t2 - y t2-1 , z t2 - z t2- 1)). In practical applications, in order to align the data and facilitate subsequent calculations, for example, the average value of all data can be set at time 0 of the preset time. The process is as Figure 3 shown, that is
[0055] In an alternative embodiment of the present invention, the above step S300 performs feature extraction based on the bone difference data to obtain the skeleton data features and the skeleton difference data features, including the following steps:
[0056] (1) Construct a skeleton difference data coordinate system based on the bone difference data and the preset time;
[0057] (2) Extract the bone difference change image within a preset time based on the skeleton difference data coordinate system;
[0058] (3) Obtain the skeleton data features and the skeleton difference data features based on the bone difference change image.
[0059] In the embodiment of the present invention, construct a skeleton difference data coordinate system according to the bone difference data and the preset time. The skeleton difference data coordinate system is asFigure 4 As shown, based on the skeletal difference data coordinate system, the change of the skeletal position within a preset time is obtained, the skeletal difference change image is extracted according to the change, and the entire skeletal change image is constructed based on the skeletal difference change image. The skeletal data features and the skeletal difference data features are extracted from the skeletal change image. In the embodiment of the present invention, the skeletal data image and the skeletal difference data are obtained through the skeletal data, so that the entire action change of the target object can be analyzed more comprehensively, the loss of lower-level feature information is avoided, and the recognition accuracy of the action behavior of the target object is further improved.
[0060] In an alternative embodiment of the present invention, in the above step S400, the feature data splicing and fusion are performed based on the joint data features and the joint difference data features to obtain the joint splicing features, including the following steps:
[0061] (1) Construct a first network layer based on the joint data features and the joint difference data features;
[0062] (2) Perform data sorting based on the first network layer to obtain a first sorting result;
[0063] (3) Perform feature data splicing and fusion based on the first sorting result to obtain the joint splicing features.
[0064] Specifically, a network layer is constructed based on the joint data features, the joint difference data features, the skeletal data, and the skeletal difference data features. The network layer is as Figure 5 shown. Among them, a first network layer is constructed based on the joint data features and the joint difference data features. The joint data features and the joint difference data features of different layers are comprehensively sorted according to the first network layer, and the joint feature data are spliced based on the first sorting result to obtain the joint splicing features, and the joint behavior prediction scores are obtained based on the joint splicing features.
[0065] Exemplarily, the joint data features and the joint difference data features of the different layers are respectively denoted as and where L is the maximum number of feature usage layers, and the corresponding numbers can be incremented from deep to shallow, for example. The joint data features and the joint difference data features are spliced by the splicing method of the channel dimension, that is, where the cat function concatenates the features and the feature in the channel dimension into
[0066] In an embodiment of the present invention, a first network layer is constructed based on the joint data features and joint difference data features. According to the first network layer, the joint data features and joint difference data features of different layers are comprehensively sorted, and the joint feature data is spliced based on the first sorting result to obtain joint splicing features. By constructing the first network layer in the embodiment of the present invention, the joint data features and joint difference data features are spliced and combined, so that the change of joints caused by the movement of the target object within a preset time can be accurately obtained, and the behavior characteristics of the target object can be analyzed more comprehensively.
[0067] In an optional embodiment of the present invention, in the above step S400, feature data splicing and fusion are performed based on the bone data and bone difference data features to obtain bone splicing features, including the following steps:
[0068] (1) Construct a second network layer based on the bone data and bone difference data features;
[0069] (2) Perform data sorting based on the second network layer to obtain a second sorting result;
[0070] (3) Perform feature data splicing and fusion based on the second sorting result to obtain bone splicing features.
[0071] Specifically, a network layer is constructed based on the joint data features, joint difference data features, bone data, and bone difference data features. The network layer is as Figure 5 shown. Among them, a second network layer is constructed based on the bone data and bone difference data features. According to the second network layer, the bone data and bone difference data features of different layers are comprehensively sorted, and the bone feature data is spliced based on the second sorting result to obtain bone splicing features, and a bone behavior prediction score is obtained based on the bone splicing features.
[0072] Exemplarily, the bone data and bone difference data features of different layers are respectively denoted as and where L is the maximum number of feature usage layers, and the corresponding numbers can be incremented from deep to shallow, for example. The bone data and bone difference data features are spliced by the splicing method in the channel dimension, that is, where the cat function splices the feature and the feature in the channel dimension into
[0073] In an embodiment of the present invention, a second network layer is constructed based on the bone data and the bone difference data features. The bone data and the bone difference data features of different layers are comprehensively sorted according to the second network layer, and the bone feature data are spliced based on the second sorting result to obtain the bone splicing feature. In the embodiment of the present invention, by constructing the second network layer, the bone data and the bone difference data features are spliced and combined, so that the change of the bone position caused by the movement of the target object within a preset time can be accurately obtained, and the behavior characteristics of the target object can be analyzed more comprehensively.
[0074] In an optional embodiment of the present invention, the above step S500 strengthens and fuses the key position features of different-dimensional branches with the joint splicing feature and the bone splicing feature respectively to obtain the action classification prediction result, including the following steps:
[0075] (1) Establish a fusion layer based on the joint splicing feature and the bone splicing feature;
[0076] (2) Extract key position feature information based on the fusion layer, and the key position is obtained based on the skeleton data feature and the skeleton difference data feature;
[0077] (3) Obtain the prediction value of the skeleton data based on the key position feature information;
[0078] (4) Obtain the action classification prediction result based on the prediction value.
[0079] Specifically, as Figure 6 shown, an average fusion layer is established based on the joint splicing feature and the bone splicing feature. Key position feature information is extracted in the fusion layer. The key position is obtained by multi-view attention focusing on the skeleton data feature and the skeleton difference data feature. According to the joint splicing feature data and the bone splicing feature data of the average fusion layer, the weight ratios of the spatial dimension branch, the time dimension branch, and the parameter channel branch are calculated. Based on the weight ratios and the key position feature information, the prediction value of the fused skeleton data is obtained, and the action classification prediction result of the target object is obtained according to the prediction value.
[0080] Exemplarily, as Figure 7 shown, the fusion of the joint data and the joint difference data can be expressed as where the joint data is f I1 ∈R C×T×S , the joint difference data is f I2 ∈R C×T×S , W com1 ∈R C×T×S is the attention weight calculated by the multi-view attention of the input feature f I1 for the multi-view attention calculation, and Wcom2 ∈R C×T×S is the input feature f I2 The attention weights for multi - perspective attention calculation, is the element - by - element multiplication of matrices. The W com ∈R C×T×S The calculation method is: where Sig is the Sigmoid function. The attention weight W s ∈R 1×1×S The calculation method is: W s = reshape(reshape(W as +W ms ))、W as = FC w2 (ReLU(FC w1 (GAP(reshape(f I )))))、W as = FC w2 (ReLU(FC w1 (GMP(reshape(f I ))))) where reshape is the matrix dimension transformation operation, which transforms the input f I ∈R C×T×S to f I ∈R S×C×T , GAP is the global average pooling operation, GMP is the global max pooling operation, FC w1 is the fully - connected layer with weight w1, and r is the channel reduction factor, ReLU is the ReLU function. The attention weight W c ∈R C×1×1 The calculation method is: W c = W ac +W mc 、W as = FC w4 (ReLU(FC w3 (GAP(f I ))))、W as = FC w4 (ReLU(FC w3 (GAP(f I )))) where, The attention weight W t ∈R 1×T×1 The calculation method is: W t = AP s (Sig(Conv9(AP c (f I )))) where APc is the average sampling operation on the parameter channel branch, AP s is the average sampling operation on the spatial dimension branch, and Conv9 is a one-dimensional convolution operation with a convolution kernel size of 9.
[0081] Exemplarily, such as Figure 7 shown, the bone data and the bone difference data are fused and spliced to obtain the weight ratios of the spatial dimension branch, the temporal dimension branch, and the parameter channel branch. The calculation process of obtaining the weight ratios of the spatial dimension branch, the temporal dimension branch, and the parameter channel branch by the fusion and splicing is the same as the process of obtaining the weight ratios of the spatial dimension branch, the temporal dimension branch, and the parameter channel branch by the fusion and splicing of the above joint data and joint difference data, and will not be elaborated here.
[0082] Exemplarily, based on the weight ratios and the key position feature information, the predicted values of the fused skeleton data are obtained. According to the joint data and the bone data prediction scores, the skeleton data prediction scores are obtained. According to the joint difference data prediction scores and the bone difference data prediction scores, the skeleton difference data prediction scores are obtained, and are respectively denoted as The predicted values of the network layer are calculated through the prediction scores where α, β, and γ are the weight parameters of the above spatial dimension branch, temporal dimension branch, and parameter channel branch, and Mod ∈ {j, b}.
[0083] In the embodiments of the present invention, an average fusion layer is established based on the joint splicing features and the bone splicing features. The key position feature information is extracted in the fusion layer. The key position is obtained by multi-view attention focusing on the skeleton data features and the skeleton difference data features. According to the joint splicing feature data and the bone splicing feature data of the average fusion layer, the weight ratios of the spatial dimension branch, the temporal dimension branch, and the parameter channel branch are calculated. Based on the weight ratios and the key position feature information, the predicted values of the fused skeleton data are obtained. According to the predicted values, the action classification prediction results of the target object are obtained. In the embodiments of the present invention, by calculating the joint prediction scores, the bone prediction scores, and the predicted scores of the fused skeleton respectively, the predicted values of the overall network layer are comprehensively obtained, so that the behavioral characteristics of the target object can be recognized more comprehensively and accurately, and the prediction is more accurate.
[0084] Such as Figure 8 shown, the embodiments of the present invention provide a skeleton human behavior recognition device, including an acquisition module 1, a calculation module 2, a feature extraction module 3, a fusion module 4, and a prediction module 5, where
[0085] An acquisition module 1 for acquiring the skeleton data of a target object, where the skeleton data includes joint data and bone data. For detailed content, reference can be made to the relevant description of step S100 in any of the above method embodiments;
[0086] A calculation module 2 for calculating joint difference data and bone difference data based on the key points of the skeleton data. For detailed content, reference can be made to the relevant description of step S200 in any of the above method embodiments;
[0087] A feature extraction module 3 for extracting features based on the joint difference data and bone difference data to obtain skeleton data features and skeleton difference data features, and obtaining joint data features and joint difference data features based on the skeleton data features. For detailed content, reference can be made to the relevant description of step S300 in any of the above method embodiments;
[0088] A fusion module 4 for respectively performing feature data splicing and fusion based on the joint data features and joint difference data features, and the skeleton data features and skeleton difference data features to obtain joint splicing features and bone splicing features. For detailed content, reference can be made to the relevant description of step S400 in any of the above method embodiments;
[0089] A prediction module 5 for respectively performing enhanced fusion of the key position features of different dimensional branches with the joint splicing features and bone splicing features to obtain an action classification prediction result. For detailed content, reference can be made to the relevant description of step S500 in any of the above method embodiments.
[0090] An embodiment of the present invention provides a skeleton human behavior recognition device. By using an acquisition device to acquire the action information of a target object, extracting the skeleton data of the target object based on the acquired action information, preprocessing the skeleton data to obtain the key points of the skeleton data, obtaining joint data of different dimensions based on the key points, obtaining joint difference data and bone difference data generated by the action transformation of the target object within a preset time based on the joint data, performing feature data splicing and fusion based on the joint data features and joint difference data features, and the skeleton data features and skeleton difference data features, obtaining joint splicing features and bone splicing features according to the fused overall skeleton and joints, performing secondary enhanced fusion of the key position features of the skeletons of different dimensions with the joint splicing features and bone splicing features to obtain a complete skeleton, and analyzing the action of the target object based on the complete skeleton to obtain an action classification prediction result. The embodiment of the present invention acquires data feature information of different dimensions by collecting the skeleton data of the target object performing an action, analyzes the action and behavior of the target object thereby, and can thus improve the accuracy of the recognition task and reduce the influence of image noise.
[0091] For the specific limitations and beneficial effects of the skeletal human behavior recognition device, reference may be made to the limitations of the skeletal human behavior recognition method in the foregoing text, which will not be elaborated herein. Each module of the above-mentioned skeletal human behavior recognition device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the electronic device in hardware form or independent thereof, or stored in the memory of the electronic device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.
[0092] An embodiment of the present invention further provides a skeletal human behavior recognition device, such as Figure 9 shown Figure 9 FIG. is a schematic structural diagram of a skeletal human behavior recognition device provided by an optional embodiment of the present invention. The skeletal human behavior recognition device may include at least one processor 41, at least one communication interface 42, at least one communication bus 43, and at least one memory 44. Among them, the communication interface 42 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the communication interface 42 may further include a standard wired interface and a wireless interface. The memory 44 may be a high-speed RAM memory (Random Access Memory, volatile random access memory), or a non-volatile memory, such as at least one disk memory. Optionally, the memory 44 may further be at least one storage device located far from the aforementioned processor 41. Among them, the processor 41 may be combined with Figure 8 the device described, and an application program is stored in the memory 44, and the processor 41 calls the program code stored in the memory 44 to execute the steps of the skeletal human behavior recognition method in any of the above method embodiments.
[0093] Among them, the communication bus 43 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The communication bus 43 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 9 only a thick line is shown in, but it does not mean that there is only one bus or one type of bus.
[0094] Among them, the memory 44 may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory 44 may further include a combination of the above types of memories.
[0095] Among them, the processor 41 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP.
[0096] Among them, the processor 41 may further include a hardware chip. The above hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0097] Optionally, the memory 44 is further configured to store program instructions. The processor 41 may call the program instructions to implement the skeleton human behavior recognition method as shown in the embodiments of the present invention. Figure 1 The skeleton human behavior recognition method shown in the embodiments.
[0098] An embodiment of the present invention also provides a non-transitory computer storage medium, which stores computer-executable instructions that can execute the skeleton human behavior recognition method in any of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memories.
[0099] Obviously, the above embodiments are merely examples for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A method for skeleton-based human behavior recognition, characterized in that, it includes: Obtain the skeleton data of the target object, where the skeleton data includes joint data and bone data; Calculate joint difference data and bone difference data based on the key points of the skeleton data; Extract features based on the joint difference data and bone difference data to obtain skeleton data features and skeleton difference data features, and obtain joint data features and joint difference data features based on the skeleton data features; Perform feature data splicing and fusion based on the joint data features and joint difference data features, and the skeleton data and skeleton difference data features respectively to obtain joint splicing features and bone splicing features; Reinforce and fuse the key position features of different dimensional branches with the joint splicing features and bone splicing features respectively to obtain the action classification prediction result; Reinforce and fuse the key position features of different dimensional branches with the joint splicing features and bone splicing features respectively to obtain the action classification prediction result, including: Establish a fusion layer based on the joint splicing features and bone splicing features; Extract key position feature information based on the fusion layer, and the key position is obtained based on the skeleton data features and skeleton difference data features; Obtain the prediction value of the skeleton data based on the key position feature information; Obtain the action classification prediction result based on the prediction value.
2. The method for skeleton-based human behavior recognition according to claim 1, characterized in that, Calculating the joint difference data based on the key points of the skeleton data includes: Extract joint data based on the skeleton data; Establish a joint coordinate system based on the joint data; Extract the joint change data within a preset time based on the joint coordinate system, and calculate the difference within the preset time based on the joint change data and joint data to obtain the joint difference data.
3. The method for skeleton-based human behavior recognition according to claim 1, characterized in that, Calculating the bone difference data based on the key points of the skeleton data includes: Extract bone data based on the skeleton data; Establish a bone coordinate system based on the bone data; Extract the bone position change data within a preset time based on the bone coordinate system, and calculate the difference within the preset time based on the bone change data and bone data to obtain the bone difference data.
4. The method for skeleton-based human behavior recognition according to claim 1, characterized in that, Extracting features based on the bone difference data to obtain skeleton data features and skeleton difference data features includes: Construct a skeleton difference data coordinate system based on the bone difference data and a preset time; Extract the bone difference change image within a preset time based on the skeleton difference data coordinate system; Obtain skeleton data features and skeleton difference data features based on the bone difference change image.
5. The method for skeleton-based human behavior recognition according to claim 1, characterized in that, Performing feature data splicing and fusion based on the joint data features and joint difference data features to obtain joint splicing features includes: Construct a first network layer based on the joint data features and joint difference data features; Perform data sorting based on the first network layer to obtain a first sorting result; Perform feature data splicing and fusion based on the first sorting result to obtain joint splicing features.
6. The skeleton human behavior recognition method according to claim 1, wherein, Perform feature data splicing and fusion based on the skeleton data and the skeleton difference data features to obtain bone splicing features, including: Construct a second network layer based on the skeleton data and the skeleton difference data features; Perform data sorting based on the second network layer to obtain a second sorting result; Perform feature data splicing and fusion based on the second sorting result to obtain bone splicing features.
7. A skeleton human behavior recognition device, wherein, comprising: An acquisition module for acquiring the skeleton data of a target object, where the skeleton data includes joint data and bone data; A calculation module for calculating joint difference data and bone difference data based on the key points of the skeleton data; A feature extraction module for performing feature extraction based on the joint difference data and the bone difference data to obtain skeleton data features and skeleton difference data features, and obtaining joint data features and joint difference data features based on the skeleton data features; A fusion module for performing feature data splicing and fusion based on the joint data features and the joint difference data features, and the skeleton data features and the skeleton difference data features respectively to obtain joint splicing features and bone splicing features; A prediction module for performing enhanced fusion of the key position features of different dimensional branches with the joint splicing features and the bone splicing features respectively to obtain an action classification prediction result; The prediction module is further configured to: Establish a fusion layer based on the joint splicing features and the bone splicing features; Extract key position feature information based on the fusion layer, where the key position is obtained based on the skeleton data features and the skeleton difference data features; Obtain the prediction value of the skeleton data based on the key position feature information; Obtain an action classification prediction result based on the prediction value.
8. A skeleton human behavior recognition device, wherein, comprising: A communication unit, a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the steps of the method according to any one of claims 1-6.
9. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Behavior identification method based on skeleton video
CN110309732A
Student classroom behavior analysis monitoring method and system based on behavior recognition
CN113361352A