Action recognition method, apparatus, and electronic device

By using knowledge distillation and classification neural network models in motion-sensing interactive devices to separate the motion attribute features and sample attribute features of motion signals, the problem of low recognition accuracy caused by differences in motion signals between different users is solved, achieving higher cross-user motion recognition accuracy and user experience.

CN115862139BActive Publication Date: 2025-12-19NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211610630.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-14
Publication Date
2025-12-19
Estimated Expiration
2042-12-14

AI Technical Summary

Technical Problem

The motion signals of different users in motion-sensing interactive devices vary greatly, resulting in low accuracy of motion recognition across users.

Method used

By using knowledge distillation, the action attribute features and sample attribute features in the action signal are separated. A classification neural network model is then used for action recognition, ignoring the user's own attributes and relying solely on action attributes for classification.

Benefits of technology

It improves the accuracy of action recognition between users, alleviates the problem of low accuracy of user action recognition during motion-sensing interaction, and provides a better user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862139B_ABST
    Figure CN115862139B_ABST
Patent Text Reader

Abstract

The present disclosure provides a motion recognition method and device and electronic equipment, and relates to the technical field of interaction. The technical problem of low accuracy of user motion recognition in somatosensory interaction is solved. The method comprises: in response to the operation of the somatosensory interaction device, acquiring the motion signal collected by the somatosensory interaction device; separating the motion attribute features in the motion signal and sample attribute features by means of knowledge distillation to obtain the motion attribute features corresponding to the motion signal; wherein the sample attribute features are used to represent the self attribute features of the user; and performing motion recognition by a classification neural network model according to the motion attribute features to obtain the motion recognition result of the user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of somatosensory interaction, and in particular to a motion recognition method and device and electronic equipment. BACKGROUND

[0002] Somatosensory interaction is to control human-computer interaction by somatosensory, and users can control the system through their own limbs to realize the interaction process. Somatosensory interaction can be applied to various scenes such as consumption, entertainment, and education and training. For example, when somatosensory interaction is applied to somatosensory games, it simulates a three-dimensional scene through a simulator, and users can hold a special game controller or stand in front of a large screen to control the actions of characters in the game through their own body movements, such as lifting, shaking, and turning hands to control virtual characters in the game, so that players are fully immersed in the game and provide somatosensory interactive game experience.

[0003] Currently, for somatosensory interaction devices, multiple users can use them simultaneously or at different times, and somatosensory interaction devices can collect motion signals of different users. However, the motion signals of different users differ greatly, which affects the motion recognition accuracy across users and makes the accuracy of user motion recognition in the somatosensory interaction process low. SUMMARY

[0004] The purpose of the present disclosure is to provide a motion recognition method, device, and electronic equipment to alleviate the technical problem of low accuracy of user motion recognition in the somatosensory interaction process.

[0005] In a first aspect, the embodiments of the present disclosure provide a motion recognition method, which collects motion signals of a user through a somatosensory interaction device; the method comprises:

[0006] In response to the operation of the somatosensory interaction device, the motion signals collected by the somatosensory interaction device are acquired;

[0007] The motion attribute features in the motion signals and sample attribute features are separated by knowledge distillation to obtain the motion attribute features corresponding to the motion signals; wherein the sample attribute features are used to represent the self attribute features of the user;

[0008] According to the motion attribute features, motion recognition is performed through a classification neural network model to obtain the motion recognition result of the user.

[0009] In a second aspect, a motion recognition device is provided, which collects motion signals of a user through a somatosensory interaction device; the device comprises:

[0010] An acquisition module is configured to acquire the motion signals collected by the somatosensory interaction device in response to the operation of the somatosensory interaction device.

[0011] a separation module configured to separate action attribute features in the action signals and sample attribute features in a manner of knowledge distillation, to obtain the action attribute features corresponding to the action signals, wherein the sample attribute features are used to represent self attribute features of the user;

[0012] a recognition module configured to perform action recognition according to the action attribute features through a classification neural network model, to obtain an action recognition result of the user.

[0013] In a third aspect, the embodiments of the present disclosure further provide an electronic device, including a memory and a processor, the memory stores a computer program which can be run on the processor, and the processor implements the method of the first aspect when executing the computer program.

[0014] In a fourth aspect, the embodiments of the present disclosure further provide a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions, when invoked and run by a processor, cause the processor to run the method of the first aspect.

[0015] The embodiments of the present disclosure bring the following beneficial effects:

[0016] The action recognition method, device and electronic device provided by the embodiments of the present disclosure can separate action attribute features in action signals collected by a somatosensory interaction device and sample attribute features used to represent self attribute features of a user in a manner of knowledge distillation, to obtain action attribute features corresponding to the action signals, and then perform action recognition according to the action attribute features through a classification neural network model, to obtain an action recognition result of the user. In the present solution, the action attribute features in the action signals and the self attribute features of the user are separated in a manner of knowledge distillation, so that the subsequent classification neural network model only relies on the action attribute for classification and recognition, the self attribute of the user is ignored, the difference between users no longer affects the recognition of actions in somatosensory interaction, the accuracy of cross-user action recognition is improved, and the technical problem of low accuracy of user action recognition in the somatosensory interaction process is alleviated.

[0017] In order to make the above objectives, characteristics and advantages of the present disclosure more apparent and easy to understand, the following preferred embodiments are specifically described below, and the accompanying drawings are referred to for detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the specific embodiments or the prior art of the present disclosure, the drawings needed to be used in the specific embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present disclosure, and other drawings can be obtained by a person of ordinary skill in the art without creative effort based on these drawings.

[0019] Figure 1 An application scenario diagram provided by an embodiment of the present disclosure is shown.

[0020] Figure 2 A flow diagram of a motion recognition method provided by an embodiment of the present disclosure is shown.

[0021] Figure 3 A diagram of a motion signal processing process in the motion recognition method provided by an embodiment of the present disclosure is shown.

[0022] Figure 4 A diagram of a new class data expansion process in the motion recognition method provided by an embodiment of the present disclosure is shown.

[0023] Figure 5 A structural diagram of a motion recognition device provided by an embodiment of the present disclosure is shown.

[0024] Figure 6 A structural diagram of an electronic device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0025] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the present disclosure will be described clearly and completely below with reference to the drawings, obviously, the described embodiments are some of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present disclosure.

[0026] The terms "comprise" and "have" and any variations thereof mentioned in the embodiments of the present disclosure are intended to cover the inclusions without exclusivity. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but optionally further includes other steps or units not listed, or optionally further includes other steps or units inherent to the process, method, product or device.

[0027] Artificial intelligence has been widely applied in image processing, natural language processing and other fields. However, the current recognition scheme is mostly based on the closed set assumption, that is, the training set and the test set use the same categories. This leads to the fact that the finally trained model can only recognize the fixed categories preset in advance. If the user makes a motion gesture that does not appear in the training process during the interaction process using the body sensing device, an incorrect response will inevitably be obtained, which significantly reduces the user experience. Secondly, for these unknown gestures, if the model needs to be expanded to recognize this category, historical data is usually needed, which has a large storage cost.

[0028] The current technical solution divides the open world recognition into two aspects, namely open set recognition and incremental learning. Open set recognition can be simply divided into two categories of traditional machine learning and deep learning. Most of the traditional machine learning is based on support vector machine (SVM), K nearest neighborhood (KNN) and extreme value machine (EVM) to model the distribution of data. Deep learning can be divided into generative and discriminative. The generative mostly uses variational auto-encoder (VAE) or generate adversarial network (GAN) to generate unknown samples, and trains a K+1 classification model, where K is the training category. The discriminative directly models the unknown boundary, for example, reciprocal point learning (RPL) directly models the center of unknown categories, and all known categories are far enough from it. In order to classify, the known categories need to be clustered to limit the range, and the distance from the corresponding point is used as the classification standard.

[0029] As can be seen, the existing scheme divides the open world recognition into two sub-problems, and lacks a unified solution. At the same time, for the body sensing interaction scene, the time sequence of the signal cannot be utilized in the above scheme. Secondly, the signal difference between different users is very large, and the existing scheme lacks a general model to achieve high accuracy recognition between users. Therefore, the accuracy of user motion recognition in the current body sensing interaction process is low.

[0030] Based on this, the embodiments of the present disclosure provide a motion recognition method, device and electronic equipment, which can alleviate the technical problem of low accuracy of user motion recognition in the body sensing interaction process.

[0031] The action recognition method in the embodiments of the present disclosure can run on a local terminal device or a server. When the action recognition method runs on the server, the method can be implemented and executed based on a cloud interaction system, where the cloud interaction system includes a server and a client device.

[0032] In an optional embodiment, various cloud applications can run on the cloud interaction system, such as cloud games. Taking cloud games as an example, cloud games refer to a game mode based on cloud computing. In the running mode of cloud games, the running subject of the game program and the presentation subject of the game picture are separated, and the storage and running of the action recognition method are completed on the cloud game server. The client device is used for receiving and sending data and presenting the game picture. For example, the client device can be a display device close to the user side with data transmission function, such as a mobile terminal, a television, a computer, a palm computer, etc. However, the cloud game server in the cloud end performs information processing. When playing the game, the player operates the client device to send operation instructions to the cloud game server, the cloud game server runs the game according to the operation instructions, encodes and compresses the game picture and other data, returns the data to the client device through the network, and finally decodes and outputs the game picture through the client device.

[0033] In an optional embodiment, taking games as an example, the local terminal device stores a game program and is used for presenting a game picture and performing somatosensory interaction. The local terminal device is used for interacting with the player through a graphical user interface and a somatosensory interaction device, that is, downloading and installing the game program through an electronic device with somatosensory interaction function and running the game program. The local terminal device can provide the graphical user interface to the player in various ways, for example, the graphical user interface can be rendered and displayed on the display screen of the terminal, or the graphical user interface can be provided to the player through holographic projection. For example, the local terminal device can include a display screen and a processor, the display screen is used for presenting a graphical user interface including a game picture, and the processor is used for running the game, generating the graphical user interface and controlling the display of the graphical user interface on the display screen.

[0034] In a possible embodiment, the embodiments of the present disclosure provide an action recognition method, which provides a graphical user interface and a somatosensory interaction function through a terminal device. The terminal device can be the aforementioned local terminal device or the aforementioned client device in the cloud interaction system.

[0035] For example, as shown in FIG. 1, Figure 1 Figure 1 ​An application scenario schematic diagram is provided for the embodiments of the present disclosure. The application scenario can include a terminal device 102 and a server 101. The terminal device can communicate with the server 101 through a wired network or a wireless network. The terminal device is configured to run a virtual desktop and a somatosensory interaction device. Through the virtual desktop and the somatosensory interaction device, the terminal device can interact with the server 101 to control content in the server 101.

[0036] The embodiments of the present disclosure are further described below with reference to the accompanying drawings.

[0037] Figure 2 A flowchart of a motion recognition method provided by an embodiment of the present disclosure is shown. The method can be applied to an electronic device that can provide a somatosensory interaction function, such as a somatosensory interaction device. The somatosensory interaction device collects motion signals of a user. Figure 2 As shown, the method includes the following steps.

[0038] In step S210, in response to the running of the somatosensory interaction device, motion signals collected by the somatosensory interaction device are obtained.

[0039] In actual application, when a user wears a collection device to make a motion, the somatosensory interaction device can collect the motion signals. The size of the motion signals can be T x N, where T is the length of the signal, and N is the number of channels.

[0040] In an optional embodiment, the signal form of the collected motion signals can be various. For example, the signal form of the motion signals includes any one or more of the following: electromyography signals, gyroscope signals, and acceleration signals. Through various signals such as electromyography signals, gyroscope data, and acceleration data, the signal collection and processing are more flexible.

[0041] In step S220, the motion attribute features in the motion signals and the sample attribute features are separated by means of knowledge distillation to obtain motion attribute features corresponding to the motion signals.

[0042] The sample attribute features are used to represent the self attribute features of the user.

[0043] In actual application, a signal can contain two parts of content, i.e., the attribute of a user (sample attribute features) and the category attribute (motion attribute features). For this type of interaction device, the signal difference between different users is very large. Generally, this type of model requires a user to collect his own signal to fine-tune the initial model. The method provided by the embodiments of the present disclosure separates the two features by means of knowledge decoupling, so that the subsequent classification network only relies on the motion attribute for classification and ignores the user information.

[0044] It should be noted that knowledge distillation, i.e., knowledge decoupling, is to make the final action classification only rely on the action attribute, and is irrelevant to the specific user. For example, as shown in FIG. 3, the signal X is divided into the category Y and the irrelevant attribute Z, Y can be classified and Y is subject to the distribution N(u, s). By using the variational autoencoder and the decoupled representation of knowledge, the input signal is encoded into sample features and category features, i.e., the intermediate layer feature is represented as sample features and category features, which can be implemented by using beta-VAE and GAN. Figure 3

[0045] In step S230, the action recognition result of the user is obtained by performing action recognition on the action attribute features through the classification neural network model.

[0046] Optionally, the classification neural network model can only rely on the action attribute features for classification, and the overall architecture can adopt c-VAE to map each action attribute feature to the prior distribution in the space for classification with respect to the distance of each prior, and the VAE itself is a generative model, and the existing prior distribution can generate the corresponding input. As for the prior mean, i.e., the classification center, the distance between each batch of samples and it can be minimized.

[0047] In the embodiments of the present disclosure, the action attribute features and the user's own attribute features in the action signal are separated by the knowledge distillation, so that the subsequent classification neural network model only relies on the action attribute for classification and recognition, the user's own attribute is ignored, the difference between users no longer affects the recognition of the action in the somatosensory interaction, the accuracy of the action recognition across users is improved, and the technical problem of low accuracy of user action recognition in the somatosensory interaction process is alleviated.

[0048] By open set recognition to prevent false response, a user-friendly interactive system can be finally obtained, the user can make irrelevant actions at will without worrying about misoperation, the difference between users can be solved, and the user can avoid consuming too much familiarization time. The knowledge decoupling is used to separate the sample features and the irrelevant user features, and the generalization recognition accuracy is improved.

[0049] The above steps will be described in detail below.

[0050] In some embodiments, the feature separation process can use a variational autoencoder to accurately determine the action attribute features corresponding to the action signal. As an example, the above step S220 can include the following steps:

[0051] In step a), the variational autoencoder is used to encode the action signal into action attribute features and sample attribute features.

[0052] ​Step b), determining the action attribute feature corresponding to the action signal by representing the intermediate layer feature of the action signal as the action attribute feature and the sample attribute feature through the representation of knowledge distillation.

[0053] Optionally, the detailed process of processing the collected action signal and action recognition is as shown in Figure 3 As shown, the input electromyographic signal is encoded as sample features (sample attribute features) and category features (action attribute features) through decoupling representation of the variational autoencoder and knowledge. Taking handwritten digits as an example, the so-called category features refer to the key information required for digit classification, i.e., the shape of the digit. The sample features refer to noise irrelevant to classification, such as the thickness and angle of the digit.

[0054] For knowledge decoupling (knowledge distillation), it needs to be explained that the intermediate layer feature is represented as sample features and category features (action features). As shown in Figure 3 As shown, the sample attribute z is first decoupled from the label y by using the beta-VAE, and a coefficient greater than 1 is added to the KL divergence, as shown in the following formula:

[0055]

[0056] Then, the label is decoupled from the sample attribute, because the generation cannot be guaranteed to be based on the label through the above method, and the input can be completely reconstructed by relying on the sample attribute regardless of the label. This is because if the model decoupling ability is insufficient, the label attribute can be completely contained in the sample attribute. For the same sample attribute, it can be understood that the reconstruction result corresponding to the category label will be better than the results of other categories:

[0057]

[0058] As can be seen, the same label attribute under different sample attributes should belong to the category, and the objective function of the VAE only models the prior evidence lower bound (ELBO), which will destroy this property, so the Wasserstein GAN is used to solve the problem. The specific method is to randomly sample sample attributes with normal distribution, and reconstruct the input jointly with the category attribute. The model needs to distinguish that the reconstructed input is false and the original input is true, as shown in the following formula:

[0059]

[0060] By using the variational autoencoder to encode the action signal into the action attribute feature and the sample attribute feature, and then representing the intermediate layer feature of the action signal as the action attribute feature and the sample attribute feature through the representation of knowledge distillation, the action attribute feature corresponding to the action signal can be more accurately determined.

[0061] In some embodiments, to facilitate the user to expand the action, incremental learning can be used for fast class expansion. As an example, the method can further include the following steps:

[0062] Step c), obtaining pseudo-known data based on the center features of the historical action classes in the historical training samples of the classification neural network model.

[0063] Step d), regenerating new samples based on the labeled unknown data and the pseudo-known data, and training the classification neural network model by means of incremental learning using the new samples to obtain a new classification neural network model after incremental learning.

[0064] Wherein, the label is the label of the expanded new action class.

[0065] For example, there are originally five actions, and if three new actions are to be added, the prior art usually requires the data of the original five actions (saved or newly collected) and the three new actions to retrain the model. As shown in Figure 4 In the embodiment of the present disclosure, the method of incremental learning is used to model the classification neural network model of the signal using c-VAE, so that each action conforms to the prior distribution. After the training of the classification neural network model is completed, the historical five data can be generated by sampling, and no additional storage of historical action data is required. By sampling N and random Z, X can be generated, and the model can generate new classes without storing historical data.

[0066] In actual application, incremental learning can be used to expand the model without using historical data, meeting the needs of users to quickly add custom gestures. As shown in Figure 4 In an open scenario, a model not only needs to recognize unknown classes that do not appear during training, but also needs to perform incremental learning on the model only with the historical model and the label of the new class if there are enough labels of unknown classes.

[0067] It should be noted that there are many types of incremental learning, such as knowledge distillation, bias correction, and elastic weight, etc. Knowledge distillation requires a small amount of historical template data, and the new model needs to keep the activation unchanged on the template data. Elastic weight does not require historical data at all, and it uses the second derivative of the gradient as the importance of the parameter, and adds a regularization term to make the important parameters of the historical task change as small as possible to avoid forgetting.

[0068] In the embodiments of the present disclosure, a unified open world recognition model is provided, two problems are integrated together, open set recognition and incremental learning tasks can be realized at the same time, through the body gesture interaction scheme in the open scene, open set recognition and incremental learning can be carried out at the same time, open set recognition prevents false response, incremental learning carries out rapid category expansion, and the user can customize actions, and the model can quickly learn and expand.

[0069] Based on the above steps c) and d), the center features of the action categories can be stored by the memory model, and pseudo known data can be obtained by adding random noise to the center features and the decoding process, and then the data of the known action categories is no longer needed. As an example, the above step c) can include the following steps:

[0070] Step e) stores the center features of the historical action categories in the historical training samples of the classification neural network model by the memory model.

[0071] Step f) adds random sample attribute feature noise to the center features to obtain target features.

[0072] Step g) decodes the target features by the decoder, and the decoded data is used as pseudo known data.

[0073] In the embodiments of the present disclosure, the "center" features of the categories are stored by the memory model. As shown in Figure 3 After the memory storage is obtained, pseudo known data can be obtained by adding random sample attribute noise and then through the decoder. When the unknown data is provided with labels, the model can be quickly fine-tuned by the sample regeneration method, and the data of the known categories is not needed.

[0074] For incremental learning, it should be noted that, as shown in Figure 4 The model of the known historical task has decoupled the knowledge representation of the intermediate layer, and the sample attribute obeys the normal distribution, so the hidden representation of the intermediate layer can be obtained by sampling the stored centroid, and the data of the historical known categories can be reconstructed:

[0075]

[0076] Finally, the model is incrementally learned by the regeneration method. It should be noted that the centroid of the historical categories does not need to be updated during learning. As shown in Figure 3 The loss function used is the same as above.

[0077] For the classifier, it should be noted that the centroid with discriminative ability is stored, and M represents the memory of all training data. This memory storage is directly learned together with the hidden layer features, and it simultaneously considers the compactness within the class and the discriminability between the classes.

[0078] In practical applications, the centroid can be calculated in two steps. The first step is neighborhood sampling: in the training process, intra-class and inter-class examples are sampled simultaneously to form a small batch. These samples are grouped according to their class labels, and the centroid ci of each group is updated by the features of the batch data. The second step is propagation: alternately update the hidden layer class features and the centroids to minimize the distance between each feature and its class centroid and maximize the distance from other centroids. The distance-based classification loss is defined as:

[0079]

[0080]

[0081] In the embodiments of the present disclosure, a memory model is proposed, as shown in the formula (1), the "center" features of the classes are stored, and then the memory storage is obtained. By adding random sample attribute noise, pseudo-known data can be obtained through the decoder. When the label of unknown data is provided, the method of sample regeneration can be used for rapid fine-tuning, and the data of known classes are not required. Figure 3

[0082] In some embodiments, the attention mechanism is used to fuse the intermediate layer features according to the time sequence characteristics of the signal, so as to improve the data accuracy. As an example, before step S220, the method can further include the following steps:

[0083] Step h), performing time sequence processing on the sequence of action signals to obtain time sequence features of the action signals.

[0084] Step i), searching for key frames of the action signals based on the time sequence features.

[0085] Step j), fusing the intermediate layer features of the action signals based on the key frames through the attention mechanism to obtain fused action signals, so that the classification neural network model focuses on the features at the key frames in the fused action signals.

[0086] Optionally, after the action signal is obtained, the Encoder can be used to encode to obtain a feature vector corresponding to a high-level semantic. The self attention structure is used in the Encoder, which can be adaptive filtering to make the model pay more attention to the core signal.

[0087] In practical applications, the attention method can refer to the standard self Attention. For time sequence features X, respectively project to Q, K, and V three spaces, QK dot product to get the similarity of V to accumulate to get the feature vector X', and at the same time, linearly activate X to get the importance I of each dimension, I dot product X' output the final vector.

[0088] ​In the embodiments of the present disclosure, the attention mechanism is used for fusing the intermediate layer features according to the timing characteristics of the signal, so as to improve the accuracy of the data.

[0089] Based on the above steps h), i) and j), the intermediate layer features of the action signal can also be weighted by the adaptive attention mechanism to enhance the discrimination ability of the features at the key moment, and solve the problem of too large signal difference between users. As an example, the above step j) can include the following steps:

[0090] In step h), the intermediate layer features of the action signal are weighted by the adaptive attention mechanism based on the key frame, to obtain the weighted action signal.

[0091] For the codec, it should be noted that due to the timing characteristics of the signal, the features are weighted by the adaptive attention mechanism to enhance the discrimination ability of the features at the key moment:

[0092]

[0093] By processing the entire signal sequence, the key frame of the signal is found, which is normalized by softmax, and the sum of the unknown weights of each sequence is 1. Using the attention mechanism can make different gesture categories pay attention to different positions in the sequence.

[0094] In the embodiments of the present disclosure, the timing characteristics of the signal are utilized, and the features are weighted by the adaptive attention mechanism to enhance the discrimination ability of the features at the key moment, so as to solve the problem of too large signal difference between users.

[0095] Figure 5 A structural schematic diagram of an action recognition device is provided. The action signal of a user is collected by a somatosensory interaction device, such as Figure 5 As shown in the figure, the action recognition device 500 includes:

[0096] The acquisition module 501 is configured to acquire the action signal collected by the somatosensory interaction device in response to the running of the somatosensory interaction device.

[0097] The separation module 502 is configured to separate the action attribute features and sample attribute features in the action signal by knowledge distillation, to obtain the action attribute features corresponding to the action signal; wherein the sample attribute features are used to represent the self attribute features of the user.

[0098] The recognition module 503 is configured to perform action recognition according to the action attribute features by a classification neural network model, to obtain the action recognition result of the user.

[0099] By the above manner, the action attribute features and the user's own attribute features in the action signals are separated by the knowledge distillation manner, so that the subsequent classification neural network model only relies on the action attribute for classification and recognition, the user's own attribute is ignored, the difference between users no longer affects the recognition of the action in the somatosensory interaction, the cross-user action recognition accuracy is improved, and the technical problem of low accuracy of user action recognition in the somatosensory interaction process is alleviated.

[0100] In an implementable embodiment, the separation module 502 is specifically configured to:

[0101] encoding the action signals into action attribute features and sample attribute features by using a variational autoencoder;

[0102] representing the intermediate layer features of the action signals as the action attribute features and the sample attribute features by using a knowledge distillation manner, and determining the action attribute features corresponding to the action signals.

[0103] In an implementable embodiment, the apparatus further comprises:

[0104] a training module configured to: obtain pseudo-known data based on the center features of historical action categories in historical training samples of the classification neural network model; regenerate new samples based on the labeled unknown data and the pseudo-known data, and train the classification neural network model by using the new samples in an incremental learning manner to obtain a new classification neural network model after incremental learning; wherein the label is a label of an expanded new action category.

[0105] In an implementable embodiment, the training module is specifically configured to:

[0106] store the center features of historical action categories in historical training samples of the classification neural network model by using a memory model;

[0107] add noise of random sample attribute features to the center features to obtain target features;

[0108] decode the target features by using a decoder, and use the decoded data as pseudo-known data.

[0109] In an implementable embodiment, the apparatus further comprises:

[0110] a fusion module configured to: perform timing processing on the sequence of action signals to obtain timing features of the action signals; search for a key frame of the action signals based on the timing features; and perform fusion on intermediate layer features of the action signals based on the key frame through an attention mechanism to obtain fused action signals, so that the classification neural network model focuses on features at the key frame in the fused action signals.

[0111] In an implementation, the fusion module is specifically configured to:

[0112] perform weighting processing on the intermediate layer features of the action signals based on the key frame through an adaptive attention mechanism to obtain weighted action signals.

[0113] In an implementation, the signal form of the action signals includes any one or more of the following:

[0114] electromyographic signals, gyroscope signals, and acceleration signals.

[0115] The action recognition apparatus provided by the embodiments of the present disclosure has the same technical features as the action recognition method provided by the above embodiments, and can therefore solve the same technical problems and achieve the same technical effects.

[0116] Figure 6 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown, which includes a processor 601, a storage medium 602, and a bus 603. The storage medium 602 stores machine-readable instructions executable by the processor 601. When the electronic device runs an action recognition method as in an embodiment, the processor 601 and the storage medium 602 communicate through the bus 603. The processor 601 executes the machine-readable instructions, and the processor 601 performs the pre-sequential part of the method to perform the following steps:

[0117] In response to running of the somatosensory interaction device, the action signals collected by the somatosensory interaction device are obtained.

[0118] The action attribute features and sample attribute features in the action signals are separated through knowledge distillation to obtain the action attribute features corresponding to the action signals. The sample attribute features are used to represent the self attribute features of the user.

[0119] The action recognition result of the user is obtained through action recognition by a classification neural network model based on the action attribute features.

[0120] In an implementable embodiment, when the processor 601 performs feature separation of the action attribute feature and the sample attribute feature in the action signal by means of knowledge distillation to obtain the action attribute feature corresponding to the action signal, the processor 601 is specifically configured to:

[0121] encoding the action signal into the action attribute feature and the sample attribute feature by means of a variational autoencoder;

[0122] representing the intermediate layer feature of the action signal as the action attribute feature and the sample attribute feature by means of a knowledge distillation representation to determine the action attribute feature corresponding to the action signal.

[0123] In an implementable embodiment, the processor is further configured to:

[0124] obtaining pseudo-known data based on the center feature of the historical action category in the historical training sample of the classification neural network model;

[0125] regenerating new samples based on the labeled unknown data and the pseudo-known data, and training the classification neural network model by means of incremental learning using the new samples to obtain a new classification neural network model after incremental learning; wherein the label is a label of an expanded new action category.

[0126] In an implementable embodiment, when the pseudo-known data is obtained based on the center feature of the historical action category in the historical training sample of the classification neural network model, the processor is specifically configured to:

[0127] storing the center feature of the historical action category in the historical training sample of the classification neural network model by means of a memory model;

[0128] adding noise of a random sample attribute feature to the center feature to obtain a target feature;

[0129] decoding the target feature by means of a decoder, and taking the decoded data as pseudo-known data.

[0130] In an implementable embodiment, before the processor 601 performs feature separation of the action attribute feature and the sample attribute feature in the action signal by means of knowledge distillation to obtain the action attribute feature corresponding to the action signal, the processor 601 is further configured to:

[0131] performing time sequence processing on the sequence of the action signal to obtain time sequence features of the action signal;

[0132] finding key frames of the action signal based on the time sequence features;

[0133] The intermediate layer features of the action signal are fused based on the key frame through an attention mechanism to obtain a fused action signal, so that the classification neural network model focuses on the features at the key frame in the fused action signal.

[0134] In a feasible implementation, when the processor 601 performs the fusion of the intermediate layer features of the action signal based on the key frame through the attention mechanism to obtain the fused action signal, the processor 601 is specifically configured to:

[0135] The intermediate layer features of the action signal are weighted based on the key frame through an adaptive attention mechanism to obtain a weighted action signal.

[0136] In a feasible implementation, the signal form of the action signal includes any one or more of the following:

[0137] An electromyographic signal, a gyroscope signal, and an acceleration signal.

[0138] In the foregoing manner, the action attribute features and the user's own attribute features in the action signal are separated through knowledge distillation, so that the subsequent classification neural network model only relies on the action attribute for classification and recognition, the user's own attribute is ignored, the difference between users no longer affects the recognition of actions in somatosensory interaction, the accuracy of cross-user action recognition is improved, and the technical problem of low accuracy of user action recognition in somatosensory interaction is alleviated.

[0139] In actual application, the memory 601 can include a high-speed random access memory (RAM) and can also include a non-volatile memory such as at least one disk memory. The communication connection between the system network element and at least one other network element is implemented through at least one communication interface 604 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.

[0140] The bus 603 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one bidirectional arrow is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0141] The memory 601 is configured to store a program, and the processor 602 executes the program after receiving an execution instruction. The method performed by the device defined by the process of any embodiment of the present disclosure can be applied to the processor 602 or implemented by the processor 602.

[0142] The processor 602 can be an integrated circuit chip with a processing capability for signals. In implementation, each step of the above method can be completed by integrated logic circuits of hardware in the processor 602 or instructions in the form of software. The processor 602 described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block disclosed in the embodiments of the present disclosure can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The storage medium is located in the storage 601, and the processor 602 reads the information in the storage 601 and combines the hardware to complete the steps of the above method.

[0143] The embodiments of the present disclosure further provide a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is run by a processor, the processor executes the following steps:

[0144] In response to running of the somatosensory interaction device, the motion signal collected by the somatosensory interaction device is acquired;

[0145] The motion attribute feature in the motion signal and a sample attribute feature are separated by knowledge distillation, to obtain the motion attribute feature corresponding to the motion signal; wherein the sample attribute feature is used to represent the self attribute feature of the user;

[0146] According to the motion attribute feature, motion recognition is performed through a classification neural network model, to obtain a motion recognition result of the user.

[0147] In an implementable embodiment, when the processor performs feature separation of the action attribute feature and the sample attribute feature in the action signal by means of knowledge distillation to obtain the action attribute feature corresponding to the action signal, the processor is specifically configured to:

[0148] encoding the action signal into the action attribute feature and the sample attribute feature by means of a variational autoencoder;

[0149] representing the intermediate layer feature of the action signal as the action attribute feature and the sample attribute feature by means of a knowledge distillation representation to determine the action attribute feature corresponding to the action signal.

[0150] In an implementable embodiment, the processor is further configured to:

[0151] obtaining pseudo-known data based on the center feature of the historical action category in the historical training sample of the classification neural network model;

[0152] regenerating new samples based on the labeled unknown data and the pseudo-known data, and training the classification neural network model by means of incremental learning using the new samples to obtain a new classification neural network model after incremental learning; wherein the label is a label of an expanded new action category.

[0153] In an implementable embodiment, when the pseudo-known data is obtained based on the center feature of the historical action category in the historical training sample of the classification neural network model, the processor is specifically configured to:

[0154] storing the center feature of the historical action category in the historical training sample of the classification neural network model by means of a memory model;

[0155] adding noise of a random sample attribute feature to the center feature to obtain a target feature;

[0156] decoding the target feature by means of a decoder, and taking the decoded data as pseudo-known data.

[0157] In an implementable embodiment, before the processor performs feature separation of the action attribute feature and the sample attribute feature in the action signal by means of knowledge distillation to obtain the action attribute feature corresponding to the action signal, the processor is further configured to:

[0158] performing time sequence processing on the sequence of the action signal to obtain time sequence features of the action signal;

[0159] finding key frames of the action signal based on the time sequence features;

[0160] The intermediate layer features of the action signal are fused based on the key frame through an attention mechanism to obtain a fused action signal, so that the classification neural network model focuses on the features at the key frame in the fused action signal.

[0161] In a feasible implementation, when the processor performs the fusion of the intermediate layer features of the action signal based on the key frame through the attention mechanism to obtain the fused action signal, the processor is specifically configured to:

[0162] The intermediate layer features of the action signal are weighted based on the key frame through an adaptive attention mechanism to obtain a weighted action signal.

[0163] In a feasible implementation, the signal form of the action signal includes any one or more of the following:

[0164] An electromyographic signal, a gyroscope signal, and an acceleration signal.

[0165] In the foregoing manner, the action attribute features and the user's own attribute features in the action signal are separated through knowledge distillation, so that the subsequent classification neural network model only relies on the action attribute for classification and recognition, the user's own attribute is ignored, the difference between users no longer affects the recognition of actions in somatosensory interaction, the cross-user action recognition accuracy is improved, and the technical problem of low accuracy of user action recognition in somatosensory interaction is alleviated.

[0166] In the embodiments of the present disclosure, the computer program, when executed by the processor, can also execute other machine-readable instructions to perform the methods described in other embodiments. For specific method steps and principles, refer to the description of the embodiments, which will not be described in detail here.

[0167] In the embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0168] For example, the flow diagrams and the block diagrams in the drawings are presented to provide illustrations of the architectures, functions, and operations of possible implementations of apparatuses, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or combinations of special purpose hardware and computer instructions.

[0169] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one place, or they may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present disclosure.

[0170] In addition, the functional units in the embodiments provided by the present disclosure can be integrated into one processing unit, or each unit can exist alone physically, or two or more units can be integrated into one unit.

[0171] The functions, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present disclosure, essentially or the part that contributes to the prior art, or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the action recognition method described in the various embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0172] It should be noted that similar reference numerals and letters refer to like items in the following drawings and that once an item is defined in one drawing, it should not be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", and the like are used only to distinguish descriptions and should not be understood as indicating or implying relative importance.

[0173] Finally, it should be noted that the above-described embodiments are merely specific embodiments of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than limit the same. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, it should be understood by those skilled in the art that any modification or easy-to-think change or equivalent replacement of some technical features of the technical solutions recorded in the foregoing embodiments can still be made within the technical range disclosed by the present disclosure by any person skilled in the art. The essence of the corresponding technical solution does not deviate from the scope of the technical solutions of the embodiments of the present disclosure. All should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method of action recognition, characterized by, The method comprises: In response to the operation of the somatosensory interaction device, the motion signal collected by the somatosensory interaction device is obtained; The motion attribute features in the motion signal and the sample attribute features are separated by knowledge distillation, and the motion attribute features corresponding to the motion signal are obtained; wherein the sample attribute features are used to represent the self attribute features of the user; According to the motion attribute features, the motion recognition is performed through a classification neural network model to obtain the motion recognition result of the user; Based on the center features of the historical motion categories in the historical training samples of the classification neural network model, pseudo-known data is obtained; based on the labeled unknown data and the pseudo-known data, new samples are regenerated, and the classification neural network model is trained by incremental learning to obtain a new classification neural network model after incremental learning; wherein the label is the label of the new motion category. The motion attribute features in the motion signal and the sample attribute features are separated by knowledge distillation, and the motion attribute features corresponding to the motion signal are obtained; wherein the sample attribute features are used to represent the self attribute features of the user; 2. The method of claim 1, wherein, The center features of the historical motion categories in the historical training samples of the classification neural network model are stored by a memory model; Random sample attribute feature noise is added to the center features to obtain target features; The target features are decoded by a decoder, and the decoded data is used as pseudo-known data. Before the motion attribute features in the motion signal and the sample attribute features are separated by knowledge distillation to obtain the motion attribute features corresponding to the motion signal, the method further comprises:

3. The method of claim 1, wherein, The sequence of the motion signal is processed in time sequence to obtain the time sequence features of the motion signal; Based on the time sequence features, the key frames of the motion signal are found; Based on the key frames, the intermediate layer features of the motion signal are fused by an attention mechanism to obtain fused motion signal, so that the classification neural network model focuses on the features at the key frames in the fused motion signal. Based on the key frames, the intermediate layer features of the motion signal are fused by an attention mechanism to obtain fused motion signal, so that the classification neural network model focuses on the features at the key frames in the fused motion signal.

4. The method of claim 3, wherein, The signal form of the motion signal includes any one or more of the following: Electromyographic signal, gyroscope signal, acceleration signal.

5. The method of claim 1, wherein, The somatosensory interaction device is used to collect motion signals for a user; the device comprises: ​ 6. An action recognition apparatus characterized by comprising: ​ An acquisition module is configured to acquire the motion signal collected by the somatosensory interaction device in response to running of the somatosensory interaction device; A separation module is configured to separate the motion attribute feature and the sample attribute feature in the motion signal by means of knowledge distillation to obtain the motion attribute feature corresponding to the motion signal; the sample attribute feature is used to represent the self attribute feature of the user; An identification module is configured to perform motion identification according to the motion attribute feature by means of a classification neural network model to obtain a motion identification result of the user. The device further includes a training module configured to: obtain pseudo-known data based on a center feature of a historical motion category in a historical training sample of the classification neural network model; regenerate new samples based on unknown data with labels and the pseudo-known data, and train the classification neural network model by means of incremental learning using the new samples to obtain a new classification neural network model after incremental learning; the labels are labels of new motion categories that are expanded; The separation module is specifically configured to: encode the motion signal into the motion attribute feature and the sample attribute feature by means of a variational autoencoder; and represent an intermediate layer feature of the motion signal as the motion attribute feature and the sample attribute feature by means of a knowledge distillation representation to determine the motion attribute feature corresponding to the motion signal.

7. An electronic device comprising a memory, a processor, the memory having stored therein a computer program executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions, and the computer executable instructions, when called and executed by the processor, cause the processor to execute the method in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Human body action recognition system based on attention mechanism feature fusion and irrelevant to position

    CN114639169A