A Human Action Recognition Method and Device Based on Multimodal Fusion
Through the multimodal fusion method of depth cameras and wearable acceleration sensors, the feature-level and decision-level fusion is used to use deep learning models to solve the problems of environmental impact and data drift in the prior art, and high-precision human motion recognition is achieved.
Patent Information
- Application Number
- CN202310615495.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-05-29
AI Technical Summary
Among the existing human body motion recognition methods, the image-based method is highly restricted by the experimental environment and is easily affected by occlusion, light changes, camera angle, etc. The method based on the wearable acceleration sensor is highly flexible but the data position is sensitive and may drift. Single-modal sensors are difficult to cope with various situations under real conditions, and different modal data structures are difficult to effectively integrate.
Using a multimodal fusion method based on depth cameras and wearable acceleration sensors, through the feature level and decision-making level fusion of acceleration data and bone data, deep learning models are used to identify human movements, including data cleaning, bone joint heatmap generation, time aggregation and decision-making fusion algorithms, to achieve accurate identification of bone data and acceleration data.
It improves the accuracy and robustness of human body movement recognition, can accurately identify movements in complex environments, reduces the influence of external factors, and enhances the adaptability and accuracy of the method.
Smart Images

Figure CN116758628B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of action recognition, and in particular, to a human action recognition method and device based on multimodal fusion. Background Art
[0002] Human action recognition is a method of recognizing and classifying human actions through computer vision technology. It can be applied to various fields, such as sports, medical, security, etc.
[0003] Among the two existing methods of human action recognition, the image-based method has a low per capita deployment cost and is more common, but it is greatly restricted by the experimental environment and is easily affected by occlusion, illumination changes, camera angles, etc.; while the method based on wearable acceleration sensors has higher flexibility, but the acceleration data is sensitive to the position of the sensor on the body, and the wearable acceleration sensor may drift during long-term recording. The existing technology usually uses single-modal sensors, making it difficult to use skeletal data and acceleration data simultaneously and also difficult to handle various situations that may occur under real conditions. In addition, the data structures of the existing two modalities are different, so how to fuse them is a problem. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a human action recognition method based on multimodal fusion to eliminate or improve one or more defects existing in the prior art.
[0005] One aspect of the present invention provides a human action recognition method based on multimodal fusion, and the steps of the method include:
[0006] Based on the acceleration parameters in the acceleration data collected by an acceleration sensor, an acceleration data segment is intercepted in the acceleration data, and a joint point data frame corresponding to the corresponding time period is intercepted in the joint point data frame collected by a depth camera;
[0007] The intercepted acceleration data segment is constructed into an acceleration parameter matrix, and the acceleration parameter matrix is input into a pre-trained first model, and the first model outputs a first feature vector;
[0008] Generate a joint point heat map tensor for each joint point in the joint point data frame, and use a time aggregation strategy to process each joint point heat map tensor into an aggregation tensor;
[0009] All the aggregation tensors are input into a pre-trained second model, and the second model outputs a second feature vector;
[0010] Based on the first feature vector and the second feature vector, a classification algorithm is used to fuse the first feature vector and the second feature vector to calculate an action classification result.
[0011] Adopting the above solution, this solution uses a human action recognition method jointly recognized based on a depth camera and a wearable acceleration sensor. After the participants wear the acceleration sensors at the body key points, the background collects the acceleration data of different joint points in real time. At the same time, 25 skeletal joints and their three-dimensional spatial positions are tracked through the depth camera. Since there will be some noises in the tracking process, the Gaussian distribution is used to obtain the three-dimensional spatial positions of the skeletal joints. By using a classification algorithm to fuse the first feature vector and the second feature vector, the skeletal data and the acceleration data can be used simultaneously, so as to perform accurate human action recognition.
[0012] In some embodiments of the present invention, in the step of using a classification algorithm to fuse the first feature vector and the second feature vector and calculating the action classification result, the first feature vector and the second feature vector are combined and input into a preset third model, and the classifier of the third model outputs the parameters corresponding to each action, and a first classification result is output based on the parameters corresponding to each action output by the classifier of the third model.
[0013] In some embodiments of the present invention, in the step of using a classification algorithm to fuse the first feature vector and the second feature vector and calculating the action classification result, the first feature vector and the second feature vector are respectively input into a preset first classifier and a second classifier. The first classifier and the second classifier respectively output the parameters corresponding to each action. Based on the parameters corresponding to each action respectively output by the first classifier and the second classifier, a decision fusion algorithm is used to output the fusion parameters of each action, and a second classification result is output based on the fusion parameters of each action.
[0014] In some embodiments of the present invention, in the step of using a decision fusion algorithm to output a second classification result based on the parameters corresponding to each action respectively output by the first classifier and the second classifier, the decision fusion algorithm of the following formula is used to output the second classification result:
[0015]
[0016] where p I (c) represents the parameter corresponding to action c output by the first classifier, and p S (c) represents the parameter corresponding to action c output by the second classifier, and P(c) represents the fusion parameter of action c.
[0017] In some embodiments of the present invention, the step of using a classification algorithm to fuse the first feature vector and the second feature vector and calculating the action classification result further includes:
[0018] Determine whether the first classification result and the second classification result are consistent;
[0019] If they are consistent, output the actions corresponding to the first classification result and the second classification result as the final action;
[0020] If they are inconsistent, obtain the parameters output by the classifier of the third model when outputting the first classification result and the fusion parameters when outputting the second classification result, compare the magnitudes of the two parameters, and use the first classification result or the second classification result corresponding to the larger parameter as the final action.
[0021] In some embodiments of the present invention, both the first model and the second model include a plurality of first convolutional units, and each first convolutional unit includes a convolutional layer, an activation layer, and a pooling layer connected in sequence.
[0022] In some embodiments of the present invention, the third model includes a plurality of second convolutional units and a classifier connected in sequence, and the second convolutional unit includes a fully connected layer, an activation layer, and a dropout layer connected in sequence.
[0023] In some embodiments of the present invention, the acceleration data segment includes the acceleration parameters of each joint point, and the acceleration parameters of each joint point include the acceleration signals of three axes, the angular velocity signals of three axes, the overall acceleration signal, and the overall angular velocity signal at each time point. In the step of constructing the intercepted acceleration data segment into an acceleration parameter matrix,
[0024] Calculate the overall acceleration signal based on the acceleration signals of three axes, and calculate the overall angular velocity signal based on the angular velocity signals of three axes;
[0025] Average the acceleration data segment into the first preset number of segments based on the time sequence, and calculate the matrix parameters of each acceleration parameter of each joint point using a data cleaning algorithm;
[0026] Construct the matrix parameters into an acceleration parameter matrix, and construct an acceleration parameter matrix with the first preset number of rows and the number of columns being the number of joint points * the number of acceleration parameters.
[0027] In some embodiments of the present invention, in the step of calculating the matrix parameters of each acceleration parameter of each joint point using a data cleaning algorithm, calculate the average value of each acceleration parameter of each joint point in each time period as the matrix parameter.
[0028] In some embodiments of the present invention, in the step of processing each joint point heatmap tensor into an aggregated tensor using a time aggregation strategy, use a maximum time aggregation strategy or a sum time aggregation strategy to process each joint point heatmap tensor into an aggregated tensor.
[0029] The second aspect of the present invention further provides a human action recognition device based on multi-modal fusion. The device includes a computer device, the computer device includes a processor and a memory, computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps implemented by the method described above.
[0030] The third aspect of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps implemented by the aforementioned human action recognition method based on multi-modal fusion.
[0031] The additional advantages, objectives, and features of the present invention will be partially elaborated in the following description, and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be pointed out and obtained specifically in the description and the accompanying drawings.
[0032] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other objectives that the present invention can achieve will be more clearly understood according to the following detailed description. Description of the Drawings
[0033] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0034] Figure 1 It is a schematic diagram of an implementation manner of the human action recognition method based on multi-modal fusion of the present invention;
[0035] Figure 2 It is a schematic diagram of another implementation manner of the human action recognition method based on multi-modal fusion of the present invention;
[0036] Figure 3 It is a schematic diagram of the acquisition process of the first classification result;
[0037] Figure 4 It is a schematic diagram of the acquisition process of the second classification result;
[0038] Figure 5 It is a schematic diagram of the process of solving the final action based on the first classification result and the second classification result;
[0039] Figure 6 It is a schematic diagram of the processing process of the time aggregation strategy. Detailed Embodiments
[0040] To make the objectives, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0041] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0042] To solve the above problems, as Figure 1 shown, the present invention proposes a human action recognition method based on multi-modal fusion. The method steps include:
[0043] Step S100, based on the acceleration parameters in the acceleration data collected by the acceleration sensor, intercept an acceleration data segment in the acceleration data, and intercept the joint point data frame corresponding to the corresponding time period in the joint point data frame collected by the depth camera;
[0044] In the specific implementation process, the acceleration data of the time period from one 0 value to another adjacent 0 value of the acceleration parameter is used as an acceleration data segment; and the joint point data frame corresponding to the corresponding time period is intercepted from the joint point data frame collected by the depth camera in the same time period.
[0045] Step S200, construct the intercepted acceleration data segment into an acceleration parameter matrix, input the acceleration parameter matrix into a pre-trained first model, and the first model outputs a first feature vector;
[0046] In the specific implementation process, the acceleration parameters in the intercepted acceleration data segment are constructed into the acceleration parameter matrix.
[0047] Step S300, generate a joint point heat map tensor for each joint point in the joint point data frame, and process each joint point heat map tensor into an aggregated tensor by using a time aggregation strategy;
[0048] In the specific implementation process, the depth camera can be a Kinect depth camera or a Vicon depth camera: The Kinect depth camera is a commonly used depth camera that can obtain the bone structure and depth information of a human or an animal, and the data can be converted into an editable format and a heat map effect can be added by using software (such as OpenNI); The Vicon depth camera is a high-precision depth camera that can obtain the bone structure and motion information of a human or an animal, and the data can be converted into an editable format and a heat map effect can be added by using software (such as Vicon Workbench).
[0049] In the specific implementation process, the steps for specifically obtaining the joint heatmap tensor can be as follows:
[0050] Obtain the skeletal structure: Use a depth camera to obtain the skeletal structure and depth information of a human or an animal, and generate a 3D model; Software (such as OpenNI) can be used to convert the data into an editable format, and extract the position and pose information of each joint point;
[0051] Calculate the joint acceleration: Calculate the acceleration of each joint point through methods such as the integration method or the Kalman filter;
[0052] Based on the calculated joint acceleration, a heatmap can be generated. The color of the heatmap represents the magnitude of the joint acceleration. The darker the color, the greater the acceleration. The brightness and contrast of the heatmap can be adjusted as needed to better display the stability information of the joint points.
[0053] Furthermore, project the heatmap onto the surface of the skeletal structure, and then apply a color gradient effect to each grid point on the surface of the skeletal structure, so that the darker the color area, the greater the acceleration area, and obtain the joint heatmap tensor of each joint point.
[0054] In the specific implementation process, the joint heatmap tensor is a three-dimensional tensor.
[0055] Step S400: Input all the aggregated tensors into a pre-trained second model, and the second model outputs a second feature vector;
[0056] Step S500: Based on the first feature vector and the second feature vector, use a classification algorithm to fuse the first feature vector and the second feature vector, and calculate the action classification result.
[0057] Adopting the above solution, this solution adopts a human action recognition method based on the joint recognition of a depth camera and a wearable acceleration sensor. After the participant wears the acceleration sensor at the body key points, the background collects the acceleration data of different joint points in real time, and at the same time tracks 25 skeletal joints and their three-dimensional spatial positions through the depth camera. Because there will be some noise in the tracking process, the Gaussian distribution is used to obtain the three-dimensional spatial positions of the skeletal joints. Using a classification algorithm to fuse the first feature vector and the second feature vector can use both skeletal data and acceleration data, so as to perform accurate human action recognition.
[0058] Such as Figure 2 and 3As shown, in some embodiments of the present invention, in the step of calculating the action classification result by fusing the first feature vector and the second feature vector using a classification algorithm, it further includes step S510 of combining and inputting the first feature vector and the second feature vector into a preset third model, outputting the parameters corresponding to each action through the classifier of the third model, and outputting a first classification result based on the parameters corresponding to each action output by the classifier of the third model.
[0059] In the specific implementation process, a 256-dimensional feature vector is extracted from the skeletal data, and a 256-dimensional feature vector is also extracted from the acceleration data. Then these features are concatenated to form a 512-dimensional vector. After that, the concatenated feature vector passes through three structures composed of a fully connected layer, an activation layer, and a dropout layer. The last layer outputs the scores of 15 action types, and the action with the highest score is regarded as the recognized action, that is, the action corresponding to the largest parameter among the parameters corresponding to each action output by the classifier of the third model is used as the first classification result.
[0060] In some embodiments of the present invention, the output action of the first classification result can be used as the final action.
[0061] As Figure 2 and 4 As shown, in some embodiments of the present invention, in the step of calculating the action classification result by fusing the first feature vector and the second feature vector using a classification algorithm, it further includes step S520 of respectively inputting the first feature vector and the second feature vector into a preset first classifier and a second classifier. The first classifier and the second classifier respectively output the parameters corresponding to each action. Based on the parameters corresponding to each action respectively output by the first classifier and the second classifier, a decision fusion algorithm is used to output the fusion parameters of each action, and a second classification result is output based on the fusion parameters of each action.
[0062] In some embodiments of the present invention, the output action of the first classification result can be used as the final action.
[0063] The action corresponding to the largest fusion parameter among the fusion parameters of each action is used as the second classification result.
[0064] In some embodiments of the present invention, in the step of outputting the second classification result using a decision fusion algorithm based on the parameters corresponding to each action respectively output by the first classifier and the second classifier, the decision fusion algorithm of the following formula is used to output the second classification result:
[0065]
[0066] where p I(c) represents the parameter corresponding to action c output by the first classifier, p S (c) represents the parameter corresponding to action c output by the second classifier, and P(c) represents the fusion parameter of action c.
[0067] As Figure 2 and 5 shown, in some embodiments of the present invention, the step of calculating the action classification result by fusing the first feature vector and the second feature vector using a classification algorithm further includes:
[0068] Step S530, determining whether the first classification result and the second classification result are consistent;
[0069] Step S540, if they are consistent, output the actions corresponding to the first classification result and the second classification result as the final action;
[0070] Step S550, if they are inconsistent, obtain the parameter output by the classifier of the third model when outputting the first classification result and the fusion parameter when outputting the second classification result, compare the magnitudes of the two parameters, and use the first classification result or the second classification result corresponding to the larger parameter as the final action.
[0071] In some embodiments of the present invention, both the first model and the second model include a plurality of first convolutional units, and each first convolutional unit includes a convolutional layer, an activation layer, and a pooling layer connected in sequence.
[0072] In some embodiments of the present invention, the third model includes a plurality of second convolutional units and a classifier connected in sequence, and the second convolutional unit includes a fully connected layer, an activation layer, and a dropout layer connected in sequence.
[0073] In some embodiments of the present invention, the acceleration data segment includes the acceleration parameters of each joint point, and the acceleration parameter of each joint point includes the acceleration signals of three axes, the angular velocity signals of three axes, the overall acceleration signal, and the overall angular velocity signal at each time point. In the step of constructing the intercepted acceleration data segment into an acceleration parameter matrix,
[0074] calculate the overall acceleration signal based on the acceleration signals of three axes, and calculate the overall angular velocity signal based on the angular velocity signals of three axes;
[0075] average the acceleration data segment into a first preset number of segments based on the time sequence, and calculate the matrix parameters of each acceleration parameter of each joint point using a data cleaning algorithm;
[0076] construct the matrix parameters into an acceleration parameter matrix, and construct an acceleration parameter matrix with the first preset number of rows and the number of columns of the number of joint points * the number of acceleration parameters.
[0077] In the specific implementation process, the overall acceleration signal (A t ) is calculated from the accelerations of three axes as follows:
[0078]
[0079] The overall angular velocity signal (G t ) is calculated from the angular velocities of three axes as follows:
[0080]
[0081] In some embodiments of the present invention, in the step of calculating the matrix parameters of each acceleration parameter of each joint point by using a data cleaning algorithm, the average value of each acceleration parameter of each joint point in each time period is calculated as the matrix parameter.
[0082] In the specific implementation process, the first preset number can be 100, and the acceleration parameters include the acceleration signals of three axes, the angular velocity signals of three axes, the overall acceleration signal and the overall angular velocity signal, that is, 8 kinds of acceleration parameters. If there are 10 joint points, an acceleration parameter matrix of 100×80 is constructed.
[0083] Adopting the above scheme, since the duration of each action is different, in order to align the input data, we obtain a total of 100 samples for each sequence through downsampling. These signals are sent to a two-dimensional convolutional neural network in the form of superimposed inertial signal images after being normalized. Because general convolution operations and pooling operations can expand the receptive field of feature extraction but may cause information loss in the feature representation. Therefore, we use dilated convolutions to extract features, so that the receptive field can be exponentially expanded without changing the field map or pooling size, and no information loss will occur. In the whole network, the image is output to a fully connected layer after passing through three two-dimensional dilated convolutional layers, and finally the scores of different actions are output through a classification function to determine the action type.
[0084] In some embodiments of the present invention, in the step of processing each joint point heatmap tensor into an aggregated tensor by using a time aggregation strategy, the maximum time aggregation strategy or the sum time aggregation strategy is used to process each joint point heatmap tensor into an aggregated tensor.
[0085] In the specific implementation process, when tracking the data of 25 skeleton joint points collected by a depth camera, due to occlusion or the position of the camera, there will be certain noise in the data. Therefore, it is very difficult to accurately mark if the spatial positions of the joints are directly projected. Therefore, a heatmap centered on the skeleton joints is obtained through a Gaussian distribution, which greatly improves the robustness.
[0086] In the step of obtaining the heatmap centered on the skeletal joints through the Gaussian distribution, for each pixel of the heatmap, its acceleration magnitude and direction are used as the parameters of the Gaussian distribution, and then the probability density function of the Gaussian distribution is applied to this pixel. The closer the probability density function is to the center, the darker the color; the farther the probability density function is from the center, the lighter the color.
[0087] Since any action lasts for a period of time, no single frame can fully reflect the information of the action. The relationship between different frames is an important condition for human action recognition. To utilize the temporal information of the skeletal sequence, we use the maximum temporal aggregation strategy or the sum temporal aggregation strategy to find the temporal information in the skeletal sequence.
[0088] If the maximum temporal aggregation strategy is adopted, it is calculated according to the following formula:
[0089]
[0090] where is the heatmap of joint j at the t-th frame, and W t is the weight coefficient of the t-th frame. This weight coefficient is related to its order in the entire time series. For a skeletal sequence of N frames, the weight coefficient represents the aggregated heatmap of joint j at the t-th frame.
[0091] If the sum temporal aggregation strategy is adopted, it is calculated according to the following formula:
[0092]
[0093] represents the aggregated heatmap of joint j for N frames.
[0094] Adopting the above scheme, as Figure 6 shown, in the maximum temporal aggregation strategy, since the weight of the subsequent frame is higher, the high value of the subsequent frame can replace the high value of the previous frame, which highlights the temporal nature of the skeletal sequence. The sum temporal aggregation strategy takes the temporal information into account through weighting. Therefore, the higher the value at a certain position, the longer the time the joint appears here. Through the temporal aggregation strategy, we stack the temporal aggregation results of each joint to obtain a three-dimensional weighted skeleton motion map. Then, the three-dimensional weighted skeleton motion map is input into a three-dimensional convolutional neural network. The output is passed through a three-dimensional max pooling layer. After passing through four three-dimensional convolutional neural networks, it is output to a fully connected layer, and then through a classification function to output the scores representing 15 possible actions. The action with the highest score is the predicted action type.
[0095] Existing human action recognition mainly falls into two methods. One uses RGB cameras or depth cameras to visually capture human activities. The image data contains sufficient information but also a lot of noise. The application range of the image data is also limited because the shooting angle of the camera is fixed or restricted, which may cause occlusion or background clutter. The other method is to use wearable inertial sensors. Although it can effectively address the above problems, inertial data is sensitive to the position of the sensor on the body, and the wearable inertial sensor may drift during long-term recording. Current human action recognition systems usually use single-modal sensors, but no single sensor modality can handle all situations that may occur under real-world conditions. Combining sensor data from multiple different modalities is a feasible method to improve the performance of human action recognition systems in the absence of redundant modalities. The data of one sensor modality can provide complementary information for other modalities.
[0096] In multi-modal fusion, according to the stage of the fusion modality, the fusion strategies are divided into data-level, feature-level, and decision-level fusion. In data-level fusion, the raw data provided by the sensors is merged before any processing. Since the raw data collected from different and non-uniform sensors may vary greatly, the data-level fusion method has poor effects and is rarely used. Feature-level fusion is the most widely used method for integrating information from different modalities currently. Due to the hierarchical structure of the deep neural network, feature-level fusion can be performed at any layer and can make full use of the information of the modalities. Decision-level fusion makes the final decision by fusing the decisions from each perception modality, which allows for a clear inspection of each modality and reduces the chance of one modality dominating another. In addition, it is relatively easy to add a new modality in decision-level fusion. We combined depth image data and inertial acceleration data through multiple fusion strategies to achieve higher-precision human action recognition.
[0097] Atrous convolution is a convolution idea proposed for the problem of image semantic segmentation that downsampling reduces the image resolution and loses information. Compared with ordinary convolution, atrous convolution has an additional dilation rate parameter in addition to the size of the convolution kernel, which is mainly used to represent the size of the dilation. Its convolution kernel size is the same as that of ordinary convolution, that is, the number of parameters remains unchanged in the neural network, but atrous convolution has a larger receptive field. Therefore, atrous convolution increases the receptive field without losing information and can better obtain the global information of the image.
[0098] In the present invention, a human action recognition method based on the joint recognition of a depth camera and a wearable acceleration sensor is adopted. After the participants wear the acceleration sensors at the body key points, the background collects the acceleration data of different joint points in real time and performs sampling and cleaning. At the same time, 25 skeletal joints and their three-dimensional spatial positions are tracked by the depth camera. Since there will be some noise in the tracking process, the Gaussian distribution is used to obtain the three-dimensional spatial positions of the skeletal joints. For the acceleration data, a two-dimensional dilated convolutional network and a classification function are used to generate action scores for different action categories, and the action with the highest action score is the recognized action. The three-dimensional human joint positions are projected to obtain the projections of each joint in each coordinate system, and the temporal aggregation results of each joint are stacked through a temporal aggregation strategy to obtain a three-dimensional weighted skeleton motion map. Then, the scores of each action are obtained through three-dimensional convolution and classification. Finally, various fusion strategies are used to fuse the obtained acceleration data and skeletal joint data, so as to perform accurate human action recognition.
[0099] Among the two existing human action recognition methods, the image-based method has a low per capita deployment cost and is more common, but it is greatly restricted by the experimental environment and is easily affected by occlusion, light changes, camera angles, etc.; while the method based on wearable acceleration sensors has higher flexibility, but the acceleration data is sensitive to the position of the sensor on the body, and the wearable acceleration sensor may drift during long-term recording. The prior art usually uses single-modal sensors, rarely uses skeletal data and acceleration data at the same time, and it is also difficult to handle various situations that may occur under real conditions. In addition, the data structures of the existing two modalities are different, so how to fuse them is a problem.
[0100] The present invention performs decision-level fusion and feature-level fusion on the skeletal data and acceleration data. One sensor modality data can provide supplementary information for other modalities, thereby effectively improving the accuracy of human action recognition. For example, when occlusion occurs, the acceleration data can provide supplementary information for the skeletal data, and when drift occurs, the skeletal data can provide correct information for the acceleration data. And through a variety of fusion methods, the two types of data are organically fused, improving the accuracy and efficiency of human action recognition. In addition, the present invention projects the skeleton data into image data, and also projects the inertial data into image data, so both can be processed by a mature convolutional neural network, which also makes the fusion method of the two modalities more flexible.
[0101] The present invention combines two common human action recognition methods, fusing the skeletal joint point data collected by a depth camera and the acceleration data collected by an acceleration sensor through a deep learning model, achieving higher-precision human action recognition. In previous technologies, 3D skeletal data was usually regarded as 3D vectors, but in this way, the spatial information of joints cannot be fully utilized. We project the position of each joint in 3D space through a skeletal projection technique to obtain its joint projection in each coordinate system, and then through temporal aggregation, obtain the temporal information between different frames, thereby obtaining a 3D weighted skeletal motion map. Then it is fed into a 3D convolutional neural network for feature learning and action classification to obtain different action scores. For the acceleration data, after data cleaning, we use a 2D dilated convolutional neural network to learn features. Through dilated convolution, we can better obtain the global information of the data, thereby achieving better feature extraction and higher precision. Finally, the scores obtained from these two sensing modes are fused at the decision level and the feature level to obtain the action label with the highest probability, thereby judging the action type and realizing human action recognition.
[0102] The beneficial effects of this solution include:
[0103] 1. The present invention performs decision-level fusion and feature-level fusion using two types of data, depth image data and acceleration data, which can effectively avoid the influence of the environment, enhance the robustness and adaptability of the method, and avoid the influence caused by lighting conditions, lighting changes, or the background. The two types of data complement and confirm each other, effectively improving the performance of the human action recognition system;
[0104] 2. The present invention effectively eliminates background noise using the skeletal projection technique and the heat map, and can accurately extract the spatial information of skeletal joint points. And by using different temporal aggregation strategies to find the temporal information in the entire skeletal sequence, it can better extract actions, and also has good recognition results for actions with long duration, fast speed, and large position changes;
[0105] 3. The present invention performs decision-level fusion and feature-level fusion on depth image data and acceleration sensor data, reducing the influence of external factors on human action recognition and improving the accuracy of human action recognition;
[0106] 4. The present invention uses the skeletal projection technique and the temporal aggregation strategy, which can better obtain the spatial information and temporal information of depth image data, and has better adaptability to actions with faster speed or larger changes.
[0107] An embodiment of the present invention further provides a human action recognition device based on multimodal fusion. The device includes a computer device, and the computer device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps implemented by the method described above.
[0108] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps implemented by the foregoing human action recognition method based on multimodal fusion are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium well known in the technical field.
[0109] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0110] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0111] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0112] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A human action recognition method based on multimodal fusion, characterized in that, The steps of the method include: Based on the acceleration parameters in the acceleration data collected by the acceleration sensor, an acceleration data segment is intercepted from the acceleration data, and a joint point data frame corresponding to the corresponding time period is intercepted from the joint point data frames collected by the depth camera; The intercepted acceleration data segment is constructed into an acceleration parameter matrix, and the acceleration parameter matrix is input into a pre-trained first model, and the first model outputs a first feature vector; A joint point heat map tensor corresponding to each joint point in the joint point data frame is generated, and each joint point heat map tensor is processed into an aggregated tensor by using a time aggregation strategy; All the aggregated tensors are input into a pre-trained second model, and the second model outputs a second feature vector; Based on the first feature vector and the second feature vector, a classification algorithm is used to fuse the first feature vector and the second feature vector, and an action classification result is calculated. The first feature vector and the second feature vector are combined and input into a preset third model, and parameters corresponding to each action are output through the classifier of the third model. Based on the parameters corresponding to each action output by the classifier of the third model, a first classification result is output. The first feature vector and the second feature vector are respectively input into a preset first classifier and a second classifier, and the first classifier and the second classifier respectively output parameters corresponding to each action. Based on the parameters corresponding to each action respectively output by the first classifier and the second classifier, a decision fusion algorithm is used to output fusion parameters for each action, and a second classification result is output based on the fusion parameters for each action. The decision fusion algorithm of the following formula is used to output the second classification result: where p I (c) represents the parameter corresponding to the output action c of the first classifier, p S (c) represents the parameter corresponding to the output action c of the second classifier, and P(c) represents the fusion parameter of action c; Determine whether the first classification result and the second classification result are consistent; if they are consistent, output the actions corresponding to the first classification result and the second classification result as the final actions; if they are inconsistent, obtain the parameters output by the classifier of the third model when the first classification result is output and the fusion parameters when the second classification result is output, compare the magnitudes of the two parameters, and use the first classification result or the second classification result corresponding to the larger parameter as the final action.
2. The human motion recognition method based on multi-modal fusion according to claim 1, wherein Both the first model and the second model include a plurality of first convolutional units, and each first convolutional unit includes a convolutional layer, an activation layer, and a pooling layer connected in sequence; the third model includes a plurality of second convolutional units and a classifier connected in sequence, and the second convolutional unit includes a fully connected layer, an activation layer, and a random inactivation layer connected in sequence.
3. The human action recognition method based on multi-modal fusion according to claim 1, characterized in that The acceleration data segment includes the acceleration parameters of each joint point, and the acceleration parameters of each joint point include the acceleration signals of three axes, the angular velocity signals of three axes, the overall acceleration signal, and the overall angular velocity signal at each time point. In the step of constructing the intercepted acceleration data segment into an acceleration parameter matrix, The overall acceleration signal is calculated based on the acceleration signals of three axes, and the overall angular velocity signal is calculated based on the angular velocity signals of three axes; The acceleration data segment is evenly divided into a first preset number of segments based on the time sequence, and a matrix parameter of each acceleration parameter of each joint point is calculated by using a data cleaning algorithm. Construct the matrix parameters as an acceleration parameter matrix, constructing an acceleration parameter matrix with the first preset number of rows and the number of columns equal to the number of joint points * the number of acceleration parameters.
4. The human action recognition method based on multi-modal fusion according to claim 3, wherein In the step of calculating the matrix parameters of each acceleration parameter of each joint point using a data cleaning algorithm, calculate the average value of each acceleration parameter of each joint point in each time period as the matrix parameter.
5. The human motion recognition method based on multimodal fusion according to claim 1, characterized in that In the step of processing the heat map tensor of each joint point into an aggregated tensor using a time aggregation strategy, use a maximum time aggregation strategy or a sum time aggregation strategy to process the heat map tensor of each joint point into an aggregated tensor.
6. A human action recognition device based on multi-modal fusion, characterized in that, The device includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps implemented by the method according to any one of claims 1-5.