Behavior Recognition Method and Apparatus, Computer-Readable Medium, and Electronic Device

Through multi-device acquisition and multi-modal data fusion, deep learning models are used to identify human behavior, solving the problem of low recognition accuracy of a single IMU device, and achieving high accuracy recognition of complex behaviors and actions.

CN114662606BActive Publication Date: 2025-07-22GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210325383.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-07-22
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

In the prior art, the motion data collected by the inertial measurement unit (IMU) of a single device is relied on the recognition of human behavior, and there is a problem that the accuracy is low and the complex human behavior and movement cannot be recognized.

Method used

User motion data and multi-modal data are collected by multiple devices, and data fusion and recognition are fusion and recognition using the multi-modal transformer fusion model and deep learning hybrid convolutional neural network-long and short-term memory neural network-action classifier model, expanding the scope and accuracy of behavior recognition.

Benefits of technology

It realizes high accuracy recognition of complex human behaviors and actions, expands the scope of recognition of behavior types, reduces the dependence on professional prior knowledge, and improves the compatibility and fineness of equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114662606B_ABST
    Figure CN114662606B_ABST
Patent Text Reader

Abstract

The present disclosure provides a behavior recognition method, a behavior recognition device, a computer-readable medium, and an electronic device, which relate to the technical field of behavior recognition and are applied to a behavior recognition system including a first device and a second device. The method includes: collecting first motion data and / or first multimodal data of a user by the first device, and collecting second motion data and second multimodal data of the user by the second device; performing behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data to obtain the behavior type of the user. The present disclosure collects motion data through multiple devices, and at the same time uses the collected multimodal data as a supplement to the motion data to provide richer human behavior data, and then performs behavior recognition based on the rich human behavior data to expand the range of behavior types that can be recognized by human behavior recognition, and at the same time improve the accuracy of human behavior recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] As an important detection means, human activity recognition (HAR) can make the interaction and monitoring functions of intelligent devices more in line with users' lives. Currently, HAR has been widely applied to current mainstream smartphones and smart wearable devices, such as the lift-to-wake function of mobile phones and the recognition of exercise types of smartwatches. Common HAR technologies mainly use the acceleration, angular velocity, and orientation changes measured by the inertial measurement unit (IMU) carried by a single device, and use corresponding data processing methods and recognition models to obtain different behaviors of users at different times. Summary of the Invention

[0003] The purpose of the present disclosure is to provide a behavior recognition method, a behavior recognition device, a computer-readable medium, and an electronic device, thereby at least to a certain extent improving the recognition range and accuracy of human activity recognition.

[0004] According to a first aspect of the present disclosure, there is provided a behavior recognition method, which is applied to a behavior recognition system including a first device and a second device, and includes: collecting first motion data and / or first multimodal data of a user through the first device, and collecting second motion data and second multimodal data of the user through the second device; wherein, the first multimodal data and the second multimodal data include other modal data of the user in addition to the motion modality; performing behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data to obtain the behavior type of the user.

[0005] According to a second aspect of the present disclosure, there is provided a behavior recognition method, which is applied to the first device, and includes: collecting first motion data and / or first multimodal data of a user, and obtaining second motion data and second multimodal data of the user sent by the second device; wherein, the first multimodal data and the second multimodal data include other modal data of the user in addition to the motion modality; performing behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data to obtain the behavior type of the user.

[0006] According to a third aspect of the present disclosure, there is provided a behavior recognition method, which is applied to the second device, and includes: collecting second motion data and second multimodal data of a user, and obtaining first motion data and / or first multimodal data of the user sent by the first device; wherein, the first multimodal data and the second multimodal data include other modal data of the user in addition to the motion modality; performing behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data to obtain the behavior type of the user.

[0007] According to a fourth aspect of the present disclosure, there is provided a behavior recognition device, which is applied to a behavior recognition system including a first device and a second device, and includes: a first acquisition module, configured to acquire first motion data and / or first multimodal data of a user through the first device, and acquire second motion data and second multimodal data of the user through the second device; wherein the first multimodal data and the second multimodal data include other modal data of the user except for the motion modality; a first recognition module, configured to perform behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data, so as to obtain the behavior type of the user.

[0008] According to a fifth aspect of the present disclosure, there is provided a behavior recognition device, which is applied to the first device and includes: a second acquisition module, configured to acquire first motion data and / or first multimodal data of the user, and obtain second motion data and second multimodal data of the user sent by the second device; wherein the first multimodal data and the second multimodal data include other modal data of the user except for the motion modality; a second recognition module, configured to perform behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data, so as to obtain the behavior type of the user.

[0009] According to a sixth aspect of the present disclosure, there is provided a behavior recognition device, which is applied to the second device and includes: a third acquisition module, configured to acquire second motion data and second multimodal data of the user, and obtain first motion data and / or first multimodal data of the user sent by the first device; wherein the first multimodal data and the second multimodal data include other modal data of the user except for the motion modality; a third recognition module, configured to perform behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data, so as to obtain the behavior type of the user.

[0010] According to a seventh aspect of the present disclosure, there is provided a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.

[0011] According to an eighth aspect of the present disclosure, there is provided an electronic device, which is characterized by including: a processor; and a memory, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.

[0012] A behavior recognition method provided by an embodiment of the present disclosure collects first motion data and / or first multimodal data of a user through a first device, and at the same time collects second motion data and second multimodal data of the user through a second device. Then, behavior recognition is jointly performed based on the motion data and the multimodal data to identify the behavior type of the user. The present disclosure collects motion data through multiple devices, and at the same time uses the collected multimodal data as a supplement to the motion data to provide richer human behavior data. Then, behavior recognition is performed based on the rich human behavior data to expand the range of behavior types that can be recognized by human behavior recognition, and at the same time improve the accuracy of human behavior recognition.

[0013] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0015] Figure 1 A schematic diagram showing an exemplary system architecture to which the embodiments of the present disclosure can be applied is shown;

[0016] Figure 2 A flowchart schematically showing a behavior recognition method in an exemplary embodiment of the present disclosure is shown;

[0017] Figure 3 A flowchart schematically showing another behavior recognition method in an exemplary embodiment of the present disclosure is shown;

[0018] Figure 4 A schematic diagram of the model structure of a multimodal transformer fusion model in an exemplary embodiment of the present disclosure is schematically shown;

[0019] Figure 5 A schematic diagram of the structure of a transformer encoder in an exemplary embodiment of the present disclosure is schematically shown;

[0020] Figure 6 A schematic diagram of the model structure of an action recognition model in an exemplary embodiment of the present disclosure is schematically shown;

[0021] Figure 7 A flowchart schematically showing an action recognition method in an exemplary embodiment of the present disclosure is shown;

[0022] Figure 8Flow chart schematically showing another action recognition method in an exemplary embodiment of the present disclosure;

[0023] Figure 9 Flow chart schematically showing yet another behavior recognition method in an exemplary embodiment of the present disclosure;

[0024] Figure 10 Schematic diagram showing data flow during a behavior recognition process in an exemplary embodiment of the present disclosure;

[0025] Figure 11 Flow chart schematically showing yet another behavior recognition method in an exemplary embodiment of the present disclosure;

[0026] Figure 12 Flow chart schematically showing still another behavior recognition method in an exemplary embodiment of the present disclosure;

[0027] Figure 13 Schematic diagram showing the composition of a behavior recognition device in an exemplary embodiment of the present disclosure;

[0028] Figure 14 Schematic diagram showing an electronic device to which embodiments of the present disclosure can be applied. Detailed implementation manners

[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.

[0030] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0031] Figure 1 Schematic diagram showing the system architecture of an exemplary application environment of a behavior recognition method and device to which embodiments of the present disclosure can be applied.

[0032] As Figure 1As shown, the system architecture 100 may include one or more of the first devices 101, 102, 103, one or more of the second devices 104, 105, 106, a network 107, and a server 108. The network 107 is used to provide a medium for communication links among the first devices 101, 102, 103, the second devices 104, 105, 106, and the server 108. The network 107 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The first devices 101, 102, 103 and the second devices 104, 105, 106 may be movable devices or wearable devices equipped with sensors, including but not limited to mobile phones, tablet computers, portable computers, watches, glasses, earphones, sports shoes, etc. It should be understood that Figure 1 the numbers of the first device, the second device, the network, and the server in

[0033] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. For example, the server 105 may be a server cluster composed of multiple servers, etc.

[0034] Based on the above one or more problems, the present exemplary embodiment provides a behavior recognition method. The behavior recognition method can be applied to a behavior recognition system including a first device and a second device. Referring to Figure 2 as shown, the behavior recognition method may include the following steps S210 and S220:

[0035] In step S210, the first device collects the first motion data and / or the first multimodal data of the user, and the second device collects the second motion data and the second multimodal data of the user.

[0036] Among them, the number of the first devices included in the behavior recognition system may be 1 or more, and the number of the second devices may also be 1 or more. The present disclosure does not make special limitations on this. For example, the behavior recognition system may include 1 first device and multiple second devices; for another example, the behavior recognition system may include multiple first devices and 1 second device; and for yet another example, the behavior recognition system may include multiple first devices and multiple second devices at the same time.

[0037] Among them, the first motion data and the second motion data can respectively include data collected when the user carries or wears the first device or the second device, which are used to characterize the movement currently occurring of the first device or the second device. For example, data such as acceleration and angular velocity. It should be noted that both the first device and the second device are equipped with sensors to facilitate the collection of the motion data corresponding to the first device and the second device when the user carries or wears the first device or the second device.

[0038] Among them, the first multimodal data and the second multimodal data can respectively include data collected when the user carries or wears the first device or the second device, which are used to characterize other modal data collected by the first device or the second device except for the motion modality, such as multimodal data such as sound, video, environmental data, and physiological data.

[0039] In an exemplary embodiment, the first device may include a movable device carried by the user, including but not limited to a smart phone, a tablet computer, a portable computer, etc.; the second device may include a wearable device worn by the user, including but not limited to a smart watch, smart glasses, smart headphones, smart sports shoes, etc.

[0040] In step S220, based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data, behavior recognition is performed to obtain the behavior type of the user.

[0041] In an exemplary embodiment, when performing behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data, recognition can be performed based on a deep learning model. Specifically, the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data can be input into a behavior recognition model for behavior recognition to obtain the behavior type of the user.

[0042] In an exemplary embodiment, referring to Figure 3 As shown, when inputting the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data into the behavior recognition model for behavior recognition, the following steps S310 and S320 may be included:

[0043] In step S310, the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data are fused to obtain fused data.

[0044] In step S320, behavior recognition is performed on the fused data to obtain the behavior type of the user.

[0045] In an exemplary embodiment, after inputting the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data into the behavior recognition model, the behavior recognition model can first perform data fusion on the three or four types of data to obtain fused data, and then perform behavior recognition based on the fused data, thereby identifying the behavior type of the user. For example, data fusion can be performed through a multimodal transformer model, and the attention mechanism can be used to effectively learn the behavior information contained in the data of each modality, thereby obtaining the behavior type of the user.

[0046] In an exemplary embodiment, the behavior recognition model can include a multimodal transformer fusion model. Referring to Figure 4 as shown, the multimodal transformer fusion model is composed of a multimodal data input layer, a linear mapping layer, a position embedding layer, a transformer encoder, and a classifier. Among them, the multimodal data input layer inputs the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data into the model; the linear mapping layer linearly maps the data of each modality (including the motion modality) to an equal-dimensional space; the position embedding layer performs position embedding according to the behavior type; the transformer encoder performs multiple encoding transformations on the embedded data and learns the mapping relationship between multiple types of data and the behavior type through the attention mechanism (the structure of the transformer encoder can be referred to Figure 5 as shown); the classifier outputs the mapping relationship to the behavior label to obtain the behavior type of the user.

[0047] By means of the multimodal transformer fusion model, without the need for steps such as feature extraction and feature screening while ensuring high accuracy of behavior recognition; at the same time, the compatibility with different wearable devices is ensured without relying on a large amount of professional prior knowledge for data feature engineering.

[0048] It should be noted that in the behavior recognition model, classifiers with multiple recognition scopes can be set to respectively recognize user behaviors in different scopes. For example, the types of user behaviors can be classified as: real-time status type, specific body part behavior type, interaction behavior type, and scenario behavior type. Among them, the real-time status type can be used to determine the overall state of the user, such as walking, running, standing still, etc.; the specific body part behavior type is used to determine the state of a specific body part of the user, such as the state of the hand, the state of the head, the state of the mouth (whether speaking), etc.; the interaction behavior type can be used to determine the interaction object of the user's current behavior, such as the user interacting with a pet, other users, the user operating household appliances, musical instruments, etc.; the scenario behavior type is used to determine the scenario where the user's current behavior is located. For example, when the user is running, whether it is indoors or outdoors, or when the user is standing on a means of transportation. Correspondingly, classifiers with corresponding recognition scopes can be set to realize the recognition of user behaviors in different scopes.

[0049] In addition, in an exemplary embodiment, after determining the type of user behavior, complex behavior recognition and user behavior prediction can also be achieved by combining and predicting associations of the recognition results of the above multiple recognition scopes. For example, through a context action joint model, complex behaviors in daily life such as cleaning the room, running in the gym, taking the subway to work, etc. can be recognized, and at the same time, prediction associations of user behaviors such as getting up to washing, getting off work to taking the subway, taking a walk to going home, etc. can be realized.

[0050] In an exemplary embodiment, in order to achieve the recognition from simple motion states to complex human behaviors, when the first device collects the first motion data, or when the first device collects the first motion data and the first multi-modal data, action recognition can also be performed based on the first motion data and the second motion data to obtain the action type of the user.

[0051] In an exemplary embodiment, when performing action recognition based on the first motion data and the second motion data, the first motion data and the second motion data can be input into the first action recognition model for action recognition to obtain the action type of the user.

[0052] In an exemplary embodiment, the first action recognition model may include a first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. The structure of the first action recognition model may include a first convolutional neural network, a first long short-term memory neural network, and a first action classifier connected in series. It should be noted that since the first motion data or the second motion data input into the first action recognition model may include multiple types of sensor data, the input branches of the first action recognition model need to be adjusted according to the number of sensor data. For example, when the input first motion data includes three types of sensor data, namely gyroscope sensor data, acceleration sensor data, and magnetometer sensor data, the input branch structure for the first motion data refers to Figure 6 as shown, and includes three input branches (i.e., three groups of first convolutional neural network-first long short-term memory neural network connected in series), corresponding to the gyroscope sensor data, acceleration sensor data, and magnetometer sensor data respectively. Finally, the first action classifier processes the output results of the three input branches. In addition, the model structure needs to be further adjusted according to the number of sensor data in the second motion data.

[0053] At this time, based on the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model for action recognition, referring to Figure 7 as shown, it may include the following steps S710 to S740:

[0054] In step S710, based on the first convolutional neural network, feature extraction is respectively performed on the first motion data and the second motion data to obtain a first spatial feature and a second spatial feature.

[0055] Specifically, the first motion data and the second motion data are input into the first convolutional neural network, and feature extraction is performed through a series of multi-layer convolutional layers, batch normalization layers, pooling layers, and dropout layers to respectively obtain the first spatial feature corresponding to the first motion data and the second spatial feature corresponding to the second motion data.

[0056] In step S720, based on the first long short-term memory neural network, feature extraction is respectively performed on the first spatial feature and the second spatial feature to obtain a first temporal feature and a second temporal feature.

[0057] Specifically, the first spatial feature and the second spatial feature are input into the first long short-term memory neural network for further feature extraction, and respectively obtain the first temporal feature corresponding to the first spatial feature and the second temporal feature corresponding to the second spatial feature.

[0058] It should be noted that the first action recognition model may include multiple layers of first long short-term memory neural network connected in series, and the specific number of layers can be set differently according to the complexity of the sensor data. For example,Figure 6 The model structure in

[0059] In step S730, the first temporal feature and the second temporal feature are subjected to feature fusion to obtain a fused feature.

[0060] Specifically, the first temporal feature and the second temporal feature can be subjected to feature fusion in a high-dimensional data space through a deep neural network to obtain a fused feature.

[0061] In step S740, the fused data is subjected to action classification based on the first action classifier to obtain the action type of the user.

[0062] Specifically, the fused feature after fusion is passed through the fully connected layer and the softmax layer in the first action classifier, and the category with the maximum probability after probability mapping is output, that is, the user action type.

[0063] By using the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model for action recognition, action recognition can be achieved while ensuring high recognition accuracy, based on a method that does not adopt steps such as feature extraction and feature screening; at the same time, this method can perform data feature engineering without relying on a large amount of professional prior knowledge, and can also ensure compatibility with different first devices and second devices. In addition, when the first device is a terminal device and the second device is a wearable device, the terminal device can be used as the core, and by adding a wearable device, a deep learning method can be used to fuse multi-sensor data, which has good scalability.

[0064] In an exemplary embodiment, in order to achieve the recognition from a simple motion state to complex human behaviors, when the first device only collects the first multi-modal data, action recognition can also be performed based on the second motion data to obtain the action type of the user.

[0065] In an exemplary embodiment, when performing action recognition based on the second motion data, the second motion data can be input into the second action recognition model for action recognition to obtain the action type of the user.

[0066] In an exemplary embodiment, the second action recognition model may include a second deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. The structure of the second action recognition model may include a second convolutional neural network, a second long short-term memory neural network, and a second action classifier connected in series. It should be noted that the specific details of the second deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model are similar to those of the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model, which have been described in detail in the implementation manner of the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. The undisclosed details can be referred to the implementation manner of this part, so they will not be elaborated here.

[0067] At this time, based on the second deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model for action recognition, referring to Figure 8 as shown, the following steps S810 to S830 may be included:

[0068] In step S810, based on the second convolutional neural network, feature extraction is performed on the second motion data to obtain third spatial features.

[0069] In step S820, based on the second long short-term memory neural network, feature extraction is performed on the third spatial features to obtain third temporal features.

[0070] In step S830, based on the second action classifier, action classification is performed on the third temporal features to obtain the action type of the user.

[0071] Specifically, the second motion data is input into the second convolutional neural network, and feature extraction is performed through a series of convolutional layers, batch normalization layers, pooling layers, and dropout layers connected in series to obtain third spatial features; then the third spatial features are input into the second long short-term memory neural network for further feature extraction to obtain third temporal features corresponding to the third spatial features; afterwards, the third temporal features pass through the fully connected layer and the softmax layer in the second action classifier, and the category with the maximum probability after probability mapping is output, that is, the user action type.

[0072] In an exemplary embodiment, the above-mentioned action types may include continuous action types and / or transitional action types. Among them, the action type refers to the type that can characterize a specific action performed based on part or all of the limbs, rather than the motion state (state type) achieved based on multiple actions; the continuous action type can characterize the type of the user's continuous action. For example, when the user is stationary (state type), it is determined whether the user is standing or sitting (continuous action type); when the user is on a vehicle (state type), it is determined whether the user is driving (continuous action type); in addition, common continuous actions may include eating, smoking, typing, sweeping the floor, and specific part actions such as nodding / shaking the head; the transitional action type can identify the transitions between continuous actions, such as standing-sitting, sitting-standing up, sitting-squatting, standing-squatting and other changing actions.

[0073] In an exemplary embodiment, when the action type includes continuous action types and / or transitional action types, when performing action recognition based on the first motion data and the second motion data, continuous action recognition and / or transitional action recognition may be performed based on the first motion data and the second motion data, so as to obtain the continuous action type and / or transitional action type of the user.

[0074] Correspondingly, when performing the recognition of the action type based on the first action recognition model or the second action recognition model, a continuous action classifier may be used as the first action classifier or the second action classifier, or a transitional action classifier may be used as the first action classifier or the second action classifier for recognition, and the continuous action type and the transitional action type are correspondingly obtained.

[0075] It should be noted that since the processing processes of the above-mentioned first action recognition model and the second action recognition model are similar, in some embodiments, the two action recognition models may be combined according to whether the processing processes are the same, and the same processing process may be executed based on the same network structure to compress the size of the action recognition model and reduce the calculation amount of action recognition at the same time.

[0076] In addition, in an exemplary embodiment, the behavior recognition model and the action recognition model may also be used as different branches of the same recognition model, and behavior recognition and action recognition may be performed simultaneously based on this recognition model.

[0077] In an exemplary embodiment, when training the above-mentioned recognition model, behavior recognition model or action recognition model, the Adam optimizer and the cross-entropy loss function may be used, and an evaluation method combining evaluation parameters such as F1 value, accuracy, precision and recall rate may be used to train the model. Among them, the cross-entropy loss function can be calculated by the following formula (1):

[0078]

[0079] Among them, y i represents the value of the true category, p i represents the predicted value output by the model, and N represents the number of possible categories for a single sample.

[0080] The following refers to Figure 9 and Figure 10 , taking a smartphone as the first device, and wearable devices such as smart headphones, smart watches, and smart glasses as the second device. Taking the first device to collect first motion data and first multimodal data at the same time, taking the inertial measurement unit IMU as the sensor for collecting motion data, and taking microphones, camera modules, photoplethysmography sensors, and electromyogram sensors as the sensors for collecting multimodal data as an example, the behavior recognition process is described as follows:

[0081] Step S901, collect first motion data through the IMU carried by the smartphone, and collect first multimodal data through the microphones, camera modules, photoplethysmography sensors, and electromyogram sensors carried by the smartphone;

[0082] Step S903, collect second motion data through the IMU carried by the wearable device, and collect second multimodal data through the microphones, camera modules, photoplethysmography sensors, and electromyogram sensors carried by the wearable device;

[0083] Specifically, the sampling frequency of the IMU sensor is 50HZ. The first motion data and the second motion data can respectively include the x, y, and z axis values of the acceleration sensor, gyroscope sensor, and magnetometer sensor, for a total of 9-dimensional data.

[0084] Specifically, the sound data collected by the microphone can include voice semantic data, sound source recognition data, reflection positioning data, etc.; the visual data collected by the camera module can include human body information, user's field of view, light data, etc.; the data collected by the photoplethysmography sensor can include heart rate data, blood oxygen data, etc.; the data collected by the electromyogram sensor can include electromyogram data, etc.

[0085] It should be noted that due to the different sensors carried in different wearable devices, the modalities of the corresponding second multimodal data collected are also different; for example, if a microphone and a camera module are set in the wearable device, the microphone can be used to collect sound data in the sound modality, and the camera module can be used to collect visual data in the visual modality.

[0086] Step S905, input the first motion data, the first multimodal data, the second motion data, and the second multimodal data into the multimodal transformer fusion model, and identify the above data based on classifiers with different recognition ranges to obtain recognition results with different recognition ranges;

[0087] Among them, the multi-modal Transformer fusion model is used to process multi-IMU data and multi-modal data of smartphones and wearable devices. The multi-modal data can include data of various modalities collected by sensors such as microphones, camera modules, photoplethysmography sensors, and electromyogram sensors.

[0088] Specifically, it is necessary to train the multi-modal Transformer fusion model in advance. During training, it is necessary to manually annotate each modality data under different behavior patterns, and then perform supervised model training.

[0089] In addition, the first motion data and the second motion data can also be input into the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model to obtain the user's action type.

[0090] Among them, the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model is used to process multi-IMU data of smartphones and wearable devices, and identify specific actions of users through IMU data, including continuous actions such as eating, typing, smoking, etc., and transitional actions such as standing-up changes.

[0091] Specifically, it is necessary to train the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model in advance:

[0092] First, when the sensor collects motion data, the starting point and ending point of each sample data can be manually annotated; it should be noted that the beginning and ending parts of each sample data may contain information of other interfering actions, so it is necessary to delete a certain length of data at the beginning and a certain length of data at the end of each group of data to ensure data quality.

[0093] Secondly, in order to extract the key information of the action, the original data can be sampled by using a sliding window. For example, when the sensor sampling frequency is 50HZ, a sliding window of size 128 can be used to collect action data within 2.56s. By setting the sensor sampling frequency and the sliding window, it can be ensured that the time span of the sampled data can cover the execution time of most daily actions. For example, after a section of three-axis sensor data with a length of 256 is processed by the sliding window and the key information is extracted, the dimension of each window data is 128×9, which can be used as sample data to input into the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model for training.

[0094] After that, the sampled sliding window data is input into the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. Each input branch of the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model performs feature extraction and fusion respectively, and conducts action recognition to obtain the action type of the user.

[0095] Meanwhile, the recognition results with multiple different recognition ranges can also be input into the context action joint model to recognize complex behaviors and predict and associate the user's behaviors simultaneously.

[0096] In an exemplary embodiment, when the behavior recognition model and the action recognition model are two branches of an identification model, the data flow process can refer to Figure 10 as shown. Specifically, the first motion data and the first multi-modal data are collected through a mobile phone, and the second motion data and the second multi-modal data are collected through wearable devices such as smart headphones, smart watches, and smart glasses. The above four types of data are input into the multi-modal transformer fusion model, and the corresponding behavior types are respectively output based on classifiers with multiple different recognition ranges. Then, the multiple behavior types are input into the context action joint model to output complex behavior types, realizing complex behavior recognition; meanwhile, the behavior prediction and association results are output. In addition, the first motion data and the second motion data are input into the first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. Through the processing of the first convolutional neural network and the first long short-term memory neural network, the first temporal feature and the second temporal feature are obtained. Then, through feature fusion and action classifiers (continuous action classifier and transition action classifier), the action types of the user (continuous action type and transition action type) are obtained.

[0097] In summary, in this exemplary embodiment, on the one hand, if only IMU data is used for behavior recognition, there will be many limitations, and specific scenarios (such as a cinema), special behaviors (such as talking), and specific objects (such as other users, pets) cannot be recognized. Therefore, through the multi-modal transformer fusion model, based on the attention mechanism, the key data concerned by different actions is learned from multi-modal data, so as to extract the abstract expression of complex behavior patterns from IMU data and multi-modal data, and then expand the range of recognizable behavior types and improve the refinement degree of recognition results. For example, the user behavior type can include performing behavior C with object B in scenario A.

[0098] On the other hand, the behavior recognition can be completed by the data collected through the mobile phone and wearable devices without the need to use additional sensors worn. Complex human behaviors can be recognized in real time in various environments, and there are few limitations, low cost, and high feasibility in real-life applications.

[0099] On the other hand, with the mobile phone as the core of computing and communication and wearable devices as an extension, sensor data fusion is carried out in a deep learning manner, which has good scalability and can realize the recognition from simple motion states to complex human actions. It can also be applied to the recognition and prediction of user behavior context.

[0100] In addition, in addition to using IMU data, multi-modal data is also fused for behavior recognition, which not only improves the types and scope of behavior recognition, but also realizes the recognition of behavior recognition scenarios and objects, effectively expanding the boundaries of behavior recognition.

[0101] Reference Figure 11 As shown, another behavior recognition method is also provided in the exemplary embodiments of the present disclosure, which can be applied to the first device. The behavior recognition method includes the following steps S1110 and S1120:

[0102] Step S1110, collect the user's first motion data and / or first multi-modal data, and obtain the user's second motion data and second multi-modal data sent by the second device.

[0103] Step S1120, perform behavior recognition based on the first motion data and / or first multi-modal data, and the second motion data and second multi-modal data to obtain the user's behavior type.

[0104] Wherein, the first multi-modal data and the second multi-modal data include other modal data of the user except the motion modality.

[0105] It should be noted that in an exemplary embodiment, when there are multiple first devices, any one of the first devices can be used as the execution subject. At this time, the first device as the execution subject not only needs to collect the user's first motion data and / or first multi-modal data, obtain the user's second motion data and second multi-modal data sent by the second device, but also needs to obtain the first motion data and / or first multi-modal data sent by other first devices to ensure the integrity of the data.

[0106] Reference Figure 12 As shown, yet another behavior recognition method is also provided in the exemplary embodiments of the present disclosure, which can be applied to the second device. The behavior recognition method includes the following steps S1210 and S1220:

[0107] Step S1210, collect the user's second motion data and second multi-modal data, and obtain the user's first motion data and / or first multi-modal data sent by the first device;

[0108] Step S1220, perform behavior recognition based on the first motion data and / or first multi-modal data, and the second motion data and second multi-modal data to obtain the user's behavior type.

[0109] Among them, the first multimodal data and the second multimodal data include other modal data of the user except for the motion modality.

[0110] Similarly, in an exemplary embodiment, when there are multiple second devices, any one of the second devices can be used as the execution entity. At this time, the second device serving as the execution entity not only needs to collect the second motion data and the second multimodal data of the user, obtain the first motion data and / or the first multimodal data of the user sent by the first device, but also needs to obtain the second motion data and the second multimodal data sent by other second devices to ensure the integrity of the data.

[0111] The specific details of each step in the above method have been described in detail in the embodiments applied to the behavior recognition system part. The details not disclosed can be referred to the content of this part of the embodiments, so they will not be elaborated here.

[0112] It should be noted that the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed, for example, synchronously or asynchronously in multiple modules.

[0113] Further, referring to Figure 13 As shown, an exemplary embodiment of the present disclosure provides a behavior recognition device 1300, which is applied to a behavior recognition system including a first device and a second device, and includes a first acquisition module 1310 and a first recognition module 1320. Among them:

[0114] The first acquisition module 1310 can be used to collect the first motion data and / or the first multimodal data of the user through the first device, and collect the second motion data and the second multimodal data of the user through the second device; among them, the first multimodal data and the second multimodal data include other modal data of the user except for the motion modality.

[0115] The first recognition module 1320 can be used to perform behavior recognition based on the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data, to obtain the behavior type of the user.

[0116] In an exemplary embodiment, the first recognition module 1320 can be used to input the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data into a behavior recognition model for behavior recognition, to obtain the behavior type of the user.

[0117] In an exemplary embodiment, the first recognition module 1320 may be configured to perform data fusion on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data to obtain fused data; and perform behavior recognition on the fused data to obtain the behavior type of the user.

[0118] In an exemplary embodiment, when the first device collects the first motion data, or when the first device collects the first motion data and the first multimodal data, the first recognition module 1320 may further be configured to perform action recognition based on the first motion data and the second motion data to obtain the action type of the user.

[0119] In an exemplary embodiment, the first recognition module 1320 may be configured to input the first motion data and the second motion data into a first action recognition model for action recognition to obtain the action type of the user.

[0120] In an exemplary embodiment, when the first action recognition model includes a first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model, the first recognition module 1320 may be configured to perform feature extraction on the first motion data and the second motion data respectively based on the first convolutional neural network to obtain a first spatial feature and a second spatial feature; perform feature extraction on the first spatial feature and the second spatial feature respectively based on the first long short-term memory neural network to obtain a first temporal feature and a second temporal feature; perform feature fusion on the first temporal feature and the second temporal feature to obtain a fused feature; and perform action classification on the fused data based on the first action classifier to obtain the action type of the user.

[0121] In an exemplary embodiment, when the first device collects the first multimodal data, the first recognition module 1320 may further be configured to perform action recognition based on the second motion data to obtain the action type of the user.

[0122] In an exemplary embodiment, the first recognition module 1320 may be configured to input the second motion data into a second action recognition model for action recognition to obtain the action type of the user.

[0123] In an exemplary embodiment, when the second action recognition model includes a second deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model, the first recognition module 1320 may be configured to perform feature extraction on the second motion data based on the second convolutional neural network to obtain a third spatial feature; perform feature extraction on the third spatial feature based on the second long short-term memory neural network to obtain a third temporal feature; and perform action classification on the third temporal feature based on the second action classifier to obtain the action type of the user.

[0124] In an exemplary embodiment, the action type includes a continuous action type and / or a transitional action type.

[0125] In an exemplary embodiment of the present disclosure, another behavior recognition device is further provided, which is applied to a first device and includes a second acquisition module and a second recognition module. Among them:

[0126] The second acquisition module can be used to acquire the user's first motion data and / or first multimodal data, and obtain the user's second motion data and second multimodal data sent by the second device; wherein, the first multimodal data and the second multimodal data include other modal data of the user except the motion modality.

[0127] The second recognition module can be used to perform behavior recognition based on the first motion data and / or first multimodal data, and the second motion data and second multimodal data to obtain the user's behavior type.

[0128] In an exemplary embodiment of the present disclosure, yet another behavior recognition device is further provided, which is applied to a second device and includes a third acquisition module and a third recognition module. Among them:

[0129] The third acquisition module can be used to acquire the user's second motion data and second multimodal data, and obtain the user's first motion data and / or first multimodal data sent by the first device; wherein, the first multimodal data and the second multimodal data include other modal data of the user except the motion modality.

[0130] The third recognition module can be used to perform behavior recognition based on the first motion data and / or first multimodal data, and the second motion data and second multimodal data to obtain the user's behavior type.

[0131] The specific details of each module in the above device have been described in detail in the method part of the embodiment. For the details not disclosed, please refer to the content of the embodiment in the method part, so it will not be repeated here.

[0132] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0133] In an exemplary embodiment of the present disclosure, an electronic device for implementing the behavior recognition method is further provided, which can be Figure 1 the terminal devices 101, 102, 103, the wearable devices 104, 105, 106, or the server 108 in. The electronic device includes at least a processor and a memory. The memory is used to store the executable instructions of the processor, and the processor is configured to execute the behavior recognition method by executing the executable instructions.

[0134] Below, taking the mobile terminal 1400 in Figure 14 as an example, the structure of the electronic device in the embodiments of the present disclosure will be described exemplarily. Those skilled in the art should understand that, except for the components specifically for mobile purposes, Figure 14 the structure in Figure 14 can also be applied to fixed-type devices. In some other embodiments, the mobile terminal 1400 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware. The interface connection relationships between the components are only schematically shown and do not constitute a limitation on the structure of the mobile terminal 1400. In some other embodiments, the mobile terminal 1400 may also adopt an interface connection method different from that in

[0135] As Figure 14 shown, the mobile terminal 1400 may specifically include: a processor 1410, an internal memory 1421, an external memory interface 1422, a Universal Serial Bus (USB) interface 1430, a charging management module 1440, a power management module 1441, a battery 1442, an antenna 1, an antenna 2, a mobile communication module 1450, a wireless communication module 1460, an audio module 1470, a speaker 1471, a receiver 1472, a microphone 1473, a headphone interface 1474, a sensor module 1480, a display screen 1490, a camera module 1491, an indicator 1492, a motor 1493, a button 1494, and a subscriber identification module (SIM) card interface 1495, etc. Among them, the sensor module 1480 may include a gyroscope sensor 14801, an acceleration sensor 14802, a magnetometer sensor 14803, etc.

[0136] The processor 1410 may include one or more processing units. For example, the processor 1410 may include an Application Processor (AP), a modem processor, a Graphics Processing Unit (GPU), an Image Signal Processor (ISP), a controller, a video codec, a Digital Signal Processor (DSP), a baseband processor, and / or a Neural-Network Processing Unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0137] The NPU is a Neural-Network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between human brain neurons, it can quickly process input information and can also continuously learn on its own. Through the NPU, applications such as intelligent cognition of the mobile terminal 1400 can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc. In some embodiments, the NPU may be used to input the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data into a behavior recognition model for behavior recognition, and input the first motion data and the second motion data into a first action recognition model, or input the second motion data into a second action recognition model for action recognition.

[0138] A memory is provided in the processor 1410. The memory may store instructions for implementing six modular functions: detection instructions, connection instructions, information management instructions, analysis instructions, data transmission instructions, and notification instructions, and is controlled by the processor 1410 for execution.

[0139] The wireless communication function of the mobile terminal 1400 can be implemented through Antenna 1, Antenna 2, the mobile communication module 1450, the wireless communication module 1460, the modulation and demodulation processor, and the baseband processor. Among them, Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals; the mobile communication module 1450 can provide solutions for wireless communications such as 14G / 3G / 4G / 5G applied to the mobile terminal 1400; the modulation and demodulation processor can include a modulator and a demodulator; the wireless communication module 1460 can provide solutions for wireless communications such as Wireless Local Area Networks (WLAN) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), etc. applied to the mobile terminal 1400. In some embodiments, Antenna 1 of the mobile terminal 1400 is coupled to the mobile communication module 1450, and Antenna 2 is coupled to the wireless communication module 1460, so that the mobile terminal 1400 can communicate with the network and other devices through wireless communication technologies.

[0140] In some embodiments, the first motion data and / or the first multimodal data collected by the first device, and the second motion data and the second multimodal data collected by the second device can be transmitted to the execution entity (behavior recognition system, the first device, the second device, the server, etc.) of the behavior recognition method through the wireless communication function, so that the execution entity can perform behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data.

[0141] The gyroscope sensor 14801 can be used to determine the motion posture of the mobile terminal 1400. In some embodiments, the angular velocity of the mobile terminal 1400 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 14803. The gyroscope sensor 14803 can be used for anti-shake shooting, navigation, somatosensory game scenarios, etc.

[0142] The acceleration sensor 14802 can detect the magnitude of the acceleration of the mobile terminal 1400 in various directions (generally three axes). When the mobile terminal 1400 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0143] The magnetometer sensor 14803 is used to locate the orientation of the device. The included angles between the electronic device and the four directions of east, south, west, and north can be measured.

[0144] In addition, sensors with other functions can be set in the sensor module 1480 according to actual needs, such as depth sensors, air pressure sensors, magnetic sensors, acceleration sensors, distance sensors, proximity light sensors, fingerprint sensors, temperature sensors, touch sensors, ambient light sensors, bone conduction sensors, etc.

[0145] The mobile terminal 1400 may also include other devices that provide auxiliary functions. For example, the key 1494 includes a power key, a volume key, etc., and the user can input through the key to generate key signal input related to the user settings and function control of the mobile terminal 1400. Another example is the indicator 1492, the motor 1493, the SIM card interface 1495, etc.

[0146] In addition, the exemplary embodiments of the present disclosure further provide a computer-readable storage medium on which a program product capable of implementing the above-mentioned method of the present specification is stored. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes a program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of the present specification, for example, the program product may be executed. Figure 2 , Figure 3 , Figure 7 , Figure 8 , Figure 9 , Figure 11 as well as Figure 12 Any one or more steps in .

[0147] It should be noted that the computer-readable medium shown in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0148] In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0149] In addition, program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages - such as Java, C++, etc., and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on a user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or, it can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0150] Those skilled in the art will readily think of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0151] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A behavior recognition method, characterized in that, Applied to a behavior recognition system including a first device and a second device, comprising: Collecting the user's first motion data and / or first multimodal data through the first device, and collecting the user's second motion data and second multimodal data through the second device; Wherein, the first multimodal data and the second multimodal data include the user's other modal data except the motion modality; Performing behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data to obtain the user's behavior type, including: inputting the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data into a behavior recognition model for behavior recognition to obtain the user's behavior type; wherein, performing data fusion on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data to obtain fused data; performing behavior recognition on the fused data to obtain the user's behavior type; Wherein, when the first device collects the first multimodal data, the method further includes: performing action recognition based on the second motion data to obtain the user's action type; Wherein, the behavior recognition model includes a first action recognition model, and the first action recognition model includes a first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. Obtaining the user's action type includes: extracting features from the first motion data and the second motion data respectively based on a first convolutional neural network to obtain a first spatial feature and a second spatial feature; extracting features from the first spatial feature and the second spatial feature respectively based on a first long short-term memory neural network to obtain a first temporal feature and a second temporal feature; performing feature fusion on the first temporal feature and the second temporal feature to obtain a fused feature; performing action classification on the fused data based on a first action classifier to obtain the user's action type.

2. The method according to claim 1, wherein When the first device collects the first motion data, or when the first device collects the first motion data and the first multimodal data, the method further includes: Performing action recognition based on the first motion data and the second motion data to obtain the user's action type.

3. The method according to claim 2, wherein The performing action recognition based on the first motion data and the second motion data to obtain the user's action type includes: Inputting the first motion data and the second motion data into a first action recognition model for action recognition to obtain the user's action type.

4. The method according to claim 1, wherein The performing action recognition based on the second motion data to obtain the user's action type includes: Inputting the second motion data into a second action recognition model for action recognition to obtain the user's action type.

5. The method according to claim 4, wherein The second action recognition model includes a second deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model; Inputting the second motion data into an action recognition model to perform action recognition to obtain the action type of the user includes: Performing feature extraction on the second motion data based on a second convolutional neural network to obtain third spatial features; Performing feature extraction on the third spatial features based on a second long short-term memory neural network to obtain third temporal features; Performing action classification on the third temporal features based on a second action classifier to obtain the action type of the user.

6. The method according to any one of claims 2 to 5, characterized in that, The action type includes a continuous action type and / or a transitional action type.

7. A behavior recognition method, characterized in that, When applied to a first device, the method includes: Collecting the first motion data and / or the first multi-modal data of the user, and obtaining the second motion data and the second multi-modal data of the user sent by a second device; Wherein, the first multi-modal data and the second multi-modal data include other modal data of the user except for the motion modality; Performing behavior recognition based on the first motion data and / or the first multi-modal data, and the second motion data and the second multi-modal data to obtain the behavior type of the user, including: inputting the first motion data and / or the first multi-modal data, and the second motion data and the second multi-modal data into a behavior recognition model to perform behavior recognition to obtain the behavior type of the user; wherein, performing data fusion on the first motion data and / or the first multi-modal data, and the second motion data and the second multi-modal data to obtain fusion data; performing behavior recognition on the fusion data to obtain the behavior type of the user; Wherein, when the first device collects the first multi-modal data, the method further includes: performing action recognition based on the second motion data to obtain the action type of the user; Wherein, the behavior recognition model includes a first action recognition model, and the first action recognition model includes a first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. Obtaining the action type of the user includes: performing feature extraction on the first motion data and the second motion data respectively based on a first convolutional neural network to obtain a first spatial feature and a second spatial feature; performing feature extraction on the first spatial feature and the second spatial feature respectively based on a first long short-term memory neural network to obtain a first temporal feature and a second temporal feature; performing feature fusion on the first temporal feature and the second temporal feature to obtain a fusion feature; performing action classification on the fusion data based on a first action classifier to obtain the action type of the user.

8. A behavior recognition method, characterized in that When applied to a second device, the method includes: Collecting the second motion data and the second multi-modal data of the user, and obtaining the first motion data and / or the first multi-modal data of the user sent by a first device; Wherein, the first multi-modal data and the second multi-modal data include other modal data of the user except for the motion modality; Performing behavior recognition based on the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data to obtain the behavior type of the user, including: inputting the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data into a behavior recognition model for behavior recognition to obtain the behavior type of the user; wherein, performing data fusion on the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data to obtain fused data; performing behavior recognition on the fused data to obtain the behavior type of the user; Wherein, when the first device collects the first multimodal data, the method further includes: performing action recognition based on the second motion data to obtain the action type of the user; Wherein, the behavior recognition model includes a first action recognition model, and the first action recognition model includes a first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. Obtaining the action type of the user includes: respectively extracting features from the first motion data and the second motion data based on a first convolutional neural network to obtain a first spatial feature and a second spatial feature; respectively extracting features from the first spatial feature and the second spatial feature based on a first long short-term memory neural network to obtain a first temporal feature and a second temporal feature; performing feature fusion on the first temporal feature and the second temporal feature to obtain a fused feature; performing action classification on the fused data based on a first action classifier to obtain the action type of the user.

9. A behavior recognition device, characterized in that, Applied to a behavior recognition system including a first device and a second device, including: A first acquisition module, configured to collect the first motion data and / or the first multimodal data of the user through the first device, and collect the second motion data and the second multimodal data of the user through the second device; wherein, the first multimodal data and the second multimodal data include other modal data of the user except the motion modality; A first recognition module, configured to perform behavior recognition based on the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data to obtain the behavior type of the user, including: inputting the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data into a behavior recognition model for behavior recognition to obtain the behavior type of the user; wherein, performing data fusion on the first motion data and / or the first multimodal data, as well as the second motion data and the second multimodal data to obtain fused data; performing behavior recognition on the fused data to obtain the behavior type of the user; Wherein, when the first device collects the first multimodal data, the apparatus is further configured to: perform action recognition based on the second motion data to obtain the action type of the user; Among them, the behavior recognition model includes a first action recognition model, and the first action recognition model includes a first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. Obtaining the action type of the user includes: extracting features from the first motion data and the second motion data respectively based on a first convolutional neural network to obtain a first spatial feature and a second spatial feature; extracting features from the first spatial feature and the second spatial feature respectively based on a first long short-term memory neural network to obtain a first temporal feature and a second temporal feature; performing feature fusion on the first temporal feature and the second temporal feature to obtain a fused feature; performing action classification on the fused data based on a first action classifier to obtain the action type of the user.

10. An action recognition device, characterized in that, Applied to a first device, the apparatus includes: A second acquisition module, configured to acquire the first motion data and / or the first multimodal data of the user, and obtain the second motion data and the second multimodal data of the user sent by a second device; wherein, the first multimodal data and the second multimodal data include other modal data of the user except the motion modality; A second recognition module, configured to perform behavior recognition based on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data to obtain the behavior type of the user, including: inputting the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data into a behavior recognition model for behavior recognition to obtain the behavior type of the user; wherein, performing data fusion on the first motion data and / or the first multimodal data, and the second motion data and the second multimodal data to obtain fused data; performing behavior recognition on the fused data to obtain the behavior type of the user; Among them, when the first device acquires the first multimodal data, the apparatus is further configured to: perform action recognition based on the second motion data to obtain the action type of the user; Among them, the behavior recognition model includes a first action recognition model, and the first action recognition model includes a first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. Obtaining the action type of the user includes: extracting features from the first motion data and the second motion data respectively based on a first convolutional neural network to obtain a first spatial feature and a second spatial feature; extracting features from the first spatial feature and the second spatial feature respectively based on a first long short-term memory neural network to obtain a first temporal feature and a second temporal feature; performing feature fusion on the first temporal feature and the second temporal feature to obtain a fused feature; performing action classification on the fused data based on a first action classifier to obtain the action type of the user.

11. A behavior recognition device, characterized in that, Applied to a second device, the apparatus includes: A third acquisition module, configured to acquire the user's second motion data and second multi-modal data, and obtain the user's first motion data and / or first multi-modal data sent by a first device; wherein, the first multi-modal data and the second multi-modal data include the user's other modal data except the motion modality; A third recognition module, configured to perform behavior recognition based on the first motion data and / or the first multi-modal data, and the second motion data and the second multi-modal data to obtain the user's behavior type, including: inputting the first motion data and / or the first multi-modal data, and the second motion data and the second multi-modal data into a behavior recognition model for behavior recognition to obtain the user's behavior type; wherein, data fusion is performed on the first motion data and / or the first multi-modal data, and the second motion data and the second multi-modal data to obtain fusion data; and behavior recognition is performed on the fusion data to obtain the user's behavior type; Wherein, when the first device acquires the first multi-modal data, the apparatus is further configured to: perform action recognition based on the second motion data to obtain the user's action type; Wherein, the behavior recognition model includes a first action recognition model, and the first action recognition model includes a first deep learning hybrid convolutional neural network-long short-term memory neural network-action classifier model. Obtaining the user's action type includes: respectively performing feature extraction on the first motion data and the second motion data based on a first convolutional neural network to obtain a first spatial feature and a second spatial feature; respectively performing feature extraction on the first spatial feature and the second spatial feature based on a first long short-term memory neural network to obtain a first temporal feature and a second temporal feature; performing feature fusion on the first temporal feature and the second temporal feature to obtain a fusion feature; and performing action classification on the fusion data based on a first action classifier to obtain the user's action type.

12. A computer-readable medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 8.

13. An electronic device, characterized in that, Including: A processor; And A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the method according to any one of claims 1 to 8 by executing the executable instructions.

Citation Information

Patent Citations

  • Behavior identification system and identification method of multi-modal sensor

    CN110807471A