Method, device and system for action recognition, and computer readable medium
By adopting active infrared sensor arrays and lightweight neural network frameworks in action recognition technology, combining convolution modules, GRU modules and attention mechanism modules, the problems of sensitivity to environmental lighting conditions and high computational volume in the existing technology are solved, and high sensitivity and precision action recognition is achieved, and suitable for embedded environments with limited resources.
Patent Information
- Application Number
- PCT/CN2024/110010
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-08-06
- Publication Date
- 2025-05-22
AI Technical Summary
The existing action recognition technology has problems such as being sensitive to ambient lighting conditions, high computational volume, needing to wear additional equipment, low recognition accuracy, and many parameters of deep learning algorithms, large computational volume, and not suitable for resource-constrained environments.
Active infrared sensor array is used to combine with lightweight neural network framework, including convolution module, GRU module and attention mechanism module, for action recognition. The framework performs feature extraction through multi-scale convolution kernel and MB3 module, the GRU module captures time series features, and the attention mechanism module enhances context information.
It realizes high sensitivity and precision action recognition under different lighting conditions, with the advantages of high-speed response, low power consumption and low cost, and is suitable for embedded environments with resource limitations.
Smart Images

Figure CN2024110010_22052025_PF_FP_ABST
Abstract
Description
Method, device, system and computer-readable medium for action recognition Technical Field
[0001] The present invention relates to action recognition using an active infrared sensor array, and in particular to a method, device, system and medium for action recognition. Background Art
[0002] Action recognition, including gesture recognition, head movement recognition, and body posture recognition, is a key research area in human-computer interaction. Currently, several approaches are available: Vision-based methods use cameras to capture images or videos of actions, then identify them through image processing and machine learning algorithms. However, these methods are sensitive to ambient lighting conditions and require a high computational load. Sensor-based methods use sensors such as accelerometers and gyroscopes on devices like gloves and head-mounted devices to capture motion data, then analyze and identify actions. However, they require additional equipment. Radar-based methods use millimeter-wave radar to capture point cloud information of body parts in the air and identify actions by analyzing motion characteristics. However, these methods suffer from low recognition accuracy. Furthermore, existing deep learning-based algorithms suffer from numerous parameters and high computational load, making them unsuitable for deployment on resource-constrained endpoint devices, such as microcontrollers (MCUs). This, in turn, limits the practical application of action recognition solutions.
[0003] Summary of the Invention
[0004] It is to be understood that both the foregoing general description and the following detailed description of the present invention are exemplary and explanatory and are intended to provide further explanation of the invention as claimed.
[0005] According to one aspect of the present invention, a method for action recognition is provided, comprising: receiving training data, the training data comprising first infrared signals corresponding to a plurality of infrared receivers and action recognition results corresponding to the first infrared signals; constructing an action recognition model, the action recognition model comprising a data preprocessing module, a convolution module, a gated recurrent unit (GRU) module, an attention mechanism module and a classification module, wherein the data preprocessing module is used to divide the input data of the action recognition model into a plurality of data blocks, each of the plurality of data blocks comprising infrared signals collected k times consecutively by a plurality of infrared receivers; wherein, for each of the plurality of data blocks: the convolution module is used to perform spatial feature extraction on a current data block to obtain a current convolution module output; the attention mechanism module is used to perform context information extraction on the current convolution module output and the previous GRU module output to obtain a current spatial attention feature; the GRU module is used to perform feature scanning on the current spatial attention feature and the previous GRU module output to obtain a current GRU module output; and wherein the classification module is used to perform action classification on the current GRU module output for the last data block in the plurality of data blocks to obtain an action recognition result; and the action recognition model is trained using the training data to obtain a trained action recognition model.
[0006] In the above method, the convolution module performs spatial feature extraction on the current data block, including performing the following operations on the current data block: using the first convolution kernel to perform single-channel feature extraction on the current data block to obtain a first feature; using the second convolution kernel to perform inter-channel feature extraction on the current data block to obtain a second feature; superimposing the first feature and the second feature and applying a first nonlinear activation function to obtain a third feature; applying a point convolution layer, a second nonlinear activation function and a depth convolution layer to the third feature to obtain a fourth feature; applying a pooling layer to the current data block to obtain a fifth feature; superimposing the fourth feature and the fifth feature to obtain a sixth feature; and applying the MB3 module to the sixth feature and superimposing it with the fifth feature, applying a batch normalization layer and a third nonlinear activation function to obtain the current convolution module output.
[0007] In the above method, the first convolution kernel includes one or more of 5×1, 7×1, 9×1, 11×1, and 13×1 convolution kernels, and the second convolution kernel is a 3×3 convolution kernel.
[0008] In the above method, the first nonlinear activation function, the second nonlinear activation function, and the third nonlinear activation function are any one of RELU, RELU6, and Sigmoid functions.
[0009] In the above method, the number of MB3 modules is 3.
[0010] In the above method, the classification module performs action classification by applying a global average pooling layer or a fully connected layer.
[0011] In the above method, it further includes: receiving a second infrared signal collected by a plurality of infrared receivers; and inputting the second infrared signal into a trained motion recognition model as input data to obtain a motion recognition result.
[0012] According to another aspect of the present invention, there is provided an apparatus for motion recognition, comprising: a storage device for storing infrared signals corresponding to a plurality of infrared receivers; and a computing device for executing any one of the above methods.
[0013] According to another aspect of the present invention, a system for action recognition is provided, comprising: an infrared transmitter for transmitting infrared signals; a plurality of infrared receivers for collecting infrared signals; a storage device for storing infrared signals collected by the plurality of infrared receivers; and a computing device for executing a method such as any one of the above methods.
[0014] According to yet another aspect of the present invention, a computer-readable medium is provided, on which computer program instructions are stored. When the computer program instructions are executed, the computer is caused to perform any one of the above methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings are included to provide a further understanding of the present invention, are incorporated into and constitute a part of this application, illustrate embodiments of the present invention, and together with the specification serve to explain the principles of the present invention. In the drawings:
[0016] FIG1 is a schematic diagram of an active infrared sensor array according to an embodiment of the present invention;
[0017] FIG2 is a schematic diagram of an action recognition model according to an embodiment of the present invention;
[0018] FIG3 is a schematic diagram of a convolution module according to an embodiment of the present invention;
[0019] FIG4 is a flow chart of a method for action recognition according to an embodiment of the present invention;
[0020] FIG5 is a block diagram of an apparatus for action recognition according to an embodiment of the present invention; and
[0021] FIG6 is a block diagram of a system for action recognition according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] Embodiments of the present invention will now be described in detail with reference to the accompanying drawings, but the invention is not limited thereto but only by the claims. In the drawings, for illustrative purposes, the dimensions of some of the elements may be exaggerated and not drawn to scale. Wherever possible, the same reference numerals will be used throughout the drawings to refer to the same or similar parts.
[0023] Although the terms used in the present invention are selected from commonly known and commonly used terms, some of the terms mentioned in the present invention specification may be selected by the applicant at his or her discretion, and their detailed meanings are explained in the relevant parts of the description herein. In addition, it is required to understand the present invention not only by the actual terms used, but also by the meaning implied by each term.
[0024] In the description provided herein, numerous specific details are set forth. However, it should be understood that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques are not shown in detail to avoid obscuring an understanding of the present invention.
[0025] In view of the above problems, this application proposes a solution of using an active infrared sensor array to perform action recognition.
[0026] Figure 1 is a schematic diagram of an active infrared sensor array 100 according to an embodiment of the present invention. The active infrared sensor array 100 is composed of an infrared transmitter 102 and infrared receivers 104A-104D. The infrared transmitter 102 can transmit infrared signals, and the infrared receivers 104A-104D can receive infrared signals reflected from a human body or an object. When the head, hand or body of a human body moves within the range of the infrared sensor array, for example, when the hand in the figure moves from right to left, the human body will reflect the infrared light emitted from the infrared transmitter 102, causing the infrared receivers 104A-104D to receive signals of different light intensities. The infrared receivers 104A-104D can collect infrared signals at a frequency of THz, which can be in the range of 10-1000Hz. The collected infrared signals are stored as x1, x2, x3, and x4 respectively. The infrared signal collected within 1 second is X = [x1, x2, x3, x4], X∈R 1×T×4 , where T represents the length of the collected signal, and 4 represents the four data channels x1, x2, x3, and x4, corresponding to the four infrared receivers 104A-104D.
[0027] By analyzing the intensity changes of the infrared signal, different actions can be identified. Although Figure 1 shows one infrared transmitter and four infrared receivers, the present application is not limited to this. The active infrared sensor array 100 may include at least one infrared transmitter and at least three infrared receivers. The infrared receivers are arranged to receive infrared light reflected by a human body or object and cannot be arranged in the same straight line. When the number of infrared receivers changes, the number of data channels of the collected infrared signal also changes accordingly.
[0028] Compared to passive infrared sensor arrays, active infrared sensor array motion recognition solutions have many advantages. Passive infrared sensor arrays sense infrared light emitted by the object itself, which has a relatively low intensity. Active infrared sensor arrays, on the other hand, sense infrared light emitted by an infrared emitter and reflected by the human body or objects. Because active infrared sensor arrays do not rely solely on infrared radiation in the environment, they can adapt to different lighting conditions, providing higher sensitivity, accuracy, and a longer recognition distance. Compared to existing vision-based, radar-based, and wearable-based solutions, this application has the advantages of high-speed response, low power consumption, and low cost.
[0029] Although there are solutions for motion recognition using infrared sensor arrays, the typical process requires calculating the spatial distance of the target part based on the reflected infrared light, extracting motion-related features, and inputting the extracted feature vectors into a pre-trained motion classification model (such as a deep neural network) to identify specific actions. Therefore, existing deep learning algorithms have the disadvantages of having many parameters, high computational complexity, difficulty in algorithm deployment, and unsuitability for embedded environments.
[0030] In order to solve the above problems existing in the application of existing deep learning algorithms to action recognition, this application proposes an action recognition model including a lightweight neural network framework.
[0031] 2 is a schematic diagram of an action recognition model 200 according to an embodiment of the present invention. The action recognition model 200 includes a data preprocessing module 202, a convolution module 204, a gated recurrent unit (GRU) module 206, an attention mechanism module 208, and a classification module 210.
[0032] Existing neural network algorithms process all the data collected by the infrared receiver as a whole to achieve action recognition. For example, the neural network frameworks such as Mobilenetv2, LSTM, and GRU are good at processing the data of the form X∈R 1×T×4 When classifying data, a tensor of size 1×T×4 will be used as the calculation object. On the one hand, there are disadvantages such as many model parameters, large amount of calculation, and large storage requirements. On the other hand, it does not fully consider the characteristic that infrared signals are generated in a time series manner.
[0033] To solve this problem, according to an embodiment of the present application, the data pre-processing module 202 can process the infrared signal X∈R collected by multiple infrared receivers. 1×T×4 The data is used as input data of the action recognition model 200, and the input data is divided into multiple data blocks {X1, X2, X3, ...}, X i ∈R 1×k×4 , i = 1, 2, ..., that is, each data block includes infrared signals collected by multiple infrared receivers for k consecutive times, and each data block X i ∈R 1×k×4 It may be provided to the convolution module 204 as an input to the convolution module 204. k may be 48, 64, 128, 256, a multiple of 8, a power of 2, and so on.
[0034] Next, for each data block in the multiple data blocks from the data pre-processing module 202 , the convolution module 204 may perform spatial feature extraction on the current data block to obtain a current convolution module output.
[0035] Figure 3 is a schematic diagram of a convolution module 300 according to an embodiment of the present invention. For each data block in a plurality of data blocks, the convolution module 300 may perform the following operations on the current data block: First, a first convolution kernel 302 is used to perform single-channel feature extraction on the data block to obtain a first feature. In one embodiment, the first convolution kernel 302 may be a single convolution kernel or multiple convolution kernels, including one or more of a 5×1, 7×1, 9×1, 11×1, or 13×1 convolution kernel. For example, the first convolution kernel 302 may be a 5×1 or a 7×1 convolution kernel. By using a large convolution kernel to extract features from the data block, the large receptive field and rich spatial information of the large convolution kernel can be fully utilized, thereby maximizing the extraction of data features from each infrared receptor. Then, a second convolution kernel 304 is used to perform inter-channel feature extraction on the data block to obtain a second feature. In one embodiment, the second convolution kernel 304 may be a 3×3 convolution kernel. By using a cross-channel convolution kernel, the overfitting tendency of large convolution kernels can be mitigated while also extracting inter-layer information between infrared receptors. Next, the first feature and the second feature are superimposed and the first nonlinear activation function 306 is applied to obtain the third feature. In one embodiment, the first nonlinear activation function 306 can be a nonlinear activation function such as RELU, RELU6, or Sigmoid function. Subsequently, a point convolution layer 308, a second nonlinear activation function 310, and a depthwise convolution layer 312 are applied to the third feature to obtain the fourth feature. In one embodiment, the second nonlinear activation function 310 can be a nonlinear activation function such as RELU, RELU6, or Sigmoid function. The positions of the point convolution layer 308 and the depthwise convolution layer 312 are interchangeable. Then, a pooling layer 314 is applied to the data block to obtain the fifth feature. The pooling operation corresponding to the pooling layer 314 can preserve the original numerical features of the data block. Next, the fourth feature and the fifth feature are superimposed to obtain the sixth feature. The superposition of the fourth and fifth features can improve the performance, robustness, and generalization ability of the neural network, making it suitable for different data forms. Next, the MB3 module 316 is applied to the sixth feature and superimposed with the fifth feature, and a batch normalization (BN) layer 320 and a third nonlinear activation function 322 are applied to obtain the convolution module output F t ∈R c×m×n , where t is the number of the data block, c is the number of convolution channels, and m and n are feature dimensions. In one embodiment, the third nonlinear activation function 322 can be a nonlinear activation function such as RELU, RELU6, Sigmoid function, etc.
[0036] MB3 module 316 can be obtained by modifying the expansion factor of MB6 module in Mobilenetv2 neural network from 6 to 3. The specific operation of MB3 module is as follows: 1×1 convolution kernel, batch normalization layer and ReLU function are applied to the input to increase the dimension according to the expansion factor 3, then 3×3 convolution kernel, batch normalization layer and ReLU function are applied, and finally 1×1 convolution kernel and batch normalization layer are applied to reduce the dimension according to the expansion factor 3. The number of MB3 modules 316 can be optional. In one embodiment, the number of MB3 modules is 3. MB3 modules can increase the nonlinear ability and expressive power of neural network models while achieving higher accuracy and lower computational complexity.
[0037] Because each infrared receiver is affected by factors such as technology and cost, the accuracy of the infrared signal fluctuates significantly, meaning the signal is noisy. However, from a longer time scale, the infrared signal received by each infrared receiver can generally reflect the motion trajectory of the action. To address this issue, this application uses multiple convolution kernels in the convolution module to perform multi-scale feature extraction on the modular input, and also introduces the MB3 module to ultimately achieve feature extraction.
[0038] Compared with the traditional single-core convolutional neural network, the multi-scale convolution module proposed in this application is conducive to learning infrared signal information and improving the model representation ability. In addition, since the constructed convolution module has the advantage of being lightweight, the combination of the multi-scale convolution kernel and the MB3 module in this application has advantages for real-time applications with limited resources. The convolution module architecture of this application also includes multi-scale convolution kernels and depth-separable (depth convolution, point convolution) convolution, which can better balance the relationship between accuracy and computational cost. It should be understood that the above is only an example of a convolution module, and those skilled in the art can use various convolution modules to perform convolution operations on data blocks in the action recognition model of this application.
[0039] Returning to Figure 2, although convolution module 204 can achieve action recognition, it is primarily used for feature extraction and only extracts infrared signal features within a small time period, failing to capture signal changes within an entire period (e.g., 1 second). Therefore, this application introduces a lightweight variant of the recurrent neural network model, GRU module 206, to capture inter-segment information and use the infrared signal's temporal scanning features to complete action classification.
[0040] The GRU module 206 can perform feature scanning on the current convolution module output from the convolution module 204 and the previous GRU module output from the GRU module 206 to obtain the current GRU module output. This application uses a variant of the GRU for sequence modeling, which uses a one-dimensional convolution instead of a fully connected layer. The specific settings of the GRU module 206 are as follows: t =σ(convr ([F t :H t-1 ])) U t =σ(conv u ([F t :H t-1 ])) H t ∈R c×m×n
[0041] Where: R t represents the reset gate, σ(·) represents the sigmoid activation function, F t represents the convolution module output from the convolution module, H t-1 Represents the previous hidden state (i.e., the output of the previous GRU module), conv r (·),conv u (·),conv h (·) represents the convolution operation, : is the tensor connector, U t represents the update gate, represents the candidate hidden state, H t represents the current hidden state (i.e., the current GRU module output), tanh(·) is the activation function, and ⊙ represents element-by-element multiplication.
[0042] Specifically, for each of the multiple data blocks, the GRU module receives the output F of the current convolution module t , and the hidden state H representing the output of the previous GRU module t-1 Through the convolution operation, F t and H t-1 Mapping is performed to obtain the reset gate R t and update gate U t . t and H after being filtered by the reset gate t-1 Perform convolution mapping to obtain candidate hidden states Will and the previous hidden state H t-1 Multiply element-wise to get the current hidden state H representing the output of the current GRU module t Among them, the update gate U t Control the degree of new information entering and reset gate R t Controls the degree of forgetting of past information. Finally, when scanning the current row, the current hidden state H tThis contains sequential information from the top of the image to the current row, which is used to predict the classification. In the neural network architecture, the weights of the update gate and reset gate are shared across pixels to facilitate learning of common rules. The GRU module 206 extracts contextual features of the entire data by cyclically processing each row of data. Using the GRU module 206 enables the model of the present application to maintain signal feature invariance.
[0043] In order to more accurately capture the key information in the time series and provide more context information, this application introduces the attention mechanism module 208 to improve the classification performance. The attention mechanism module 208 can extract context information from the current convolution module output from the convolution module 204 and the previous GRU module output from the GRU module 206 to obtain the current spatial attention feature, and provide the current spatial attention feature to the GRU module 206. The specific settings of the attention mechanism module 208 are as follows: K t =conv(F t ;Θ k ) V t =conv(F t ;Θ v )
[0044] Among them: K t represents the spatial attention weight, F t is the output of the convolution module, Θ k 、Θ v is the convolution kernel, conv(·) represents the convolution operation, S t is the correlation matrix, H t-1 is the previous hidden state from the GRU module 206 (i.e., the previous GRU module output), is the attention map, V t represents the information context of the input sequence at time point t, Represents spatial attention features.
[0045] Specifically, the attention mechanism module 208 is concerned with the current convolution module output F from the convolution module 204. t Perform 1D convolution to get K t , calculate K t and the previous hidden state H from GRU module 206 t-1 The correlation matrix S is obtained by t , for S t Apply softmax to get the attention map F t Perform 1D convolution to get V t , and finally V tand attention map Multiply and add F t Add together to get the spatial attention feature The spatial attention feature is weighted by calculating the correlation between each spatial position in the current feature map and the previous hidden state of the GRU module, thereby enhancing the feature representation related to the target.
[0046] Then, the attention mechanism module 208 provides the spatial attention feature to the GRU module 206 instead of the convolution module output. Thus, the GRU module 206 can perform feature scanning on the current spatial attention feature from the attention mechanism module 208 and the previous GRU module output from the GRU module 206 to obtain the current GRU module output.
[0047] When the last data block among the plurality of data blocks is processed, the GRU module 206 may provide the current GRU module output for the last data block as the last GRU module output to the classification module 210 .
[0048] Finally, the classification module 210 performs action classification on the current GRU module output for the last data block in the plurality of data blocks from the GRU module 206 to obtain an action recognition result. In one embodiment, the action classification performed by the classification module 210 may include applying a global average pooling (GAP) layer or a fully connected layer.
[0049] This application realizes the lightweighting of the model by designing a lightweight convolution module based on multi-scale convolution kernel and MB3 module, introducing GRU module to realize the classification of time series, and introducing spatial attention mechanism to realize context connection.
[0050] After constructing the motion recognition model 200, the motion recognition model 200 can be trained using training data to obtain a trained motion recognition model. The training data may include infrared signals corresponding to the plurality of infrared receivers and motion recognition results corresponding to the infrared signals. The infrared signals collected by the plurality of infrared receivers can then be input into the trained motion recognition model as input data to obtain motion recognition results.
[0051] Table 1 below shows a comparison of the recognition performance of the proposed action recognition model with the lightweight neural network algorithm Mobilenetv2_0.1 and the time series processing method GRU. Taking hand gestures as an example, there are 13 gestures (up and down, down and up, left and right, right and left, left and right, right and left, right and left, right and left, right and right, forward rotation, reverse rotation, far and near waving, left and right waving, and up and down waving). Each action is collected about 20 times, and data enhancement technology is used to enhance each action to 100 times.
[0052] Table 1 Gesture recognition performance comparison
[0053] As can be seen from Table 1, the action recognition model proposed in this application has a high recognition rate and can make full use of the gesture signals collected by the infrared sensor array. On the one hand, the Mobilenetv2_0.1 network is a neural network framework designed for images. When the time series is converted into two-dimensional information without time dimension, it will cause interference. For example, the two actions of up and down, and down and up have high similarity in the two-dimensional spatial dimension, but have large differences in time sequence. On the other hand, the recognition accuracy of the GRU neural network is lower than that of the action recognition model of this application. This is mainly because the convolution module added in this application improves the local feature extraction ability and generalization ability. In addition, the attention mechanism module introduced in this application further improves the generalization ability and recognition accuracy of the model. At the same time, the model parameters of this application are slightly more than those of the GRU model and far less than those of the Mobilenetv2_0.1 model. The main reason is that although the convolution module introduced in this application increases the parameters of the model, the convolution module parameters are relatively small, so the parameter amount is not significantly increased. In terms of MAC (multiplication and accumulation operation), the computational complexity of the action recognition model of this application is slightly higher than that of Mobilenetv2_0.1 and much lower than that of GRU. The main reason is that after introducing the convolution module, this application reduces the number of parameters input to the GRU module, thereby reducing the computational complexity of the GRU module.
[0054] FIG4 is a flowchart of a method 400 for action recognition according to an embodiment of the present invention.
[0055] At step 402, training data is received. The training data may include first infrared signals corresponding to a plurality of infrared receivers and action recognition results corresponding to the first infrared signals.
[0056] At step 404, an action recognition model is constructed. The action recognition model may include a data preprocessing module, a convolution module, a gated recurrent unit (GRU) module, an attention mechanism module, and a classification module, wherein the data preprocessing module is used to divide the input data of the action recognition model into multiple data blocks, each of the multiple data blocks includes infrared signals collected k times continuously by multiple infrared receivers; wherein, for each data block in the multiple data blocks: the convolution module is used to extract spatial features of the current data block to obtain the current convolution module output; the attention mechanism module is used to extract context information from the current convolution module output and the previous GRU module output to obtain the current spatial attention feature; the GRU module is used to perform feature scanning on the current spatial attention feature and the previous GRU module output to obtain the current GRU module output; and wherein the classification module is used to perform action classification on the current GRU module output for the last data block in the multiple data blocks to obtain an action recognition result.
[0057] At step 406 , the action recognition model is trained using the training data to obtain a trained action recognition model.
[0058] In one embodiment, the convolution module performing spatial feature extraction on the current data block may include performing the following operations on the current data block: using a first convolution kernel to perform single-channel feature extraction on the current data block to obtain a first feature; using a second convolution kernel to perform inter-channel feature extraction on the current data block to obtain a second feature; superimposing the first feature and the second feature and applying a first nonlinear activation function to obtain a third feature; applying a point convolution layer, a second nonlinear activation function and a depth convolution layer to the third feature to obtain a fourth feature; applying a pooling layer to the current data block to obtain a fifth feature; superimposing the fourth feature and the fifth feature to obtain a sixth feature; and applying an MB3 module to the sixth feature and superimposing it with the fifth feature, applying a batch normalization layer and a third nonlinear activation function to obtain the current convolution module output.
[0059] In one embodiment, the first convolution kernel may include one or more of 5×1, 7×1, 9×1, 11×1, and 13×1 convolution kernels, and the second convolution kernel may be a 3×3 convolution kernel.
[0060] In one embodiment, the first nonlinear activation function, the second nonlinear activation function, and the third nonlinear activation function may be any one of RELU, RELU6, and Sigmoid functions.
[0061] In one embodiment, the number of MB3 modules may be 3.
[0062] In one embodiment, the classification module performing action classification may include applying a global average pooling layer or a fully connected layer.
[0063] In one embodiment, the method 400 may further include: receiving a second infrared signal collected by a plurality of infrared receivers; and inputting the second infrared signal into a trained motion recognition model as input data to obtain a motion recognition result.
[0064] 5 is a block diagram of an apparatus 500 for action recognition according to an embodiment of the present invention. Apparatus 500 may include a storage device 502 and a computing device 504. Storage device 502 may store infrared signals corresponding to a plurality of infrared receivers and provide the infrared signals to computing device 504. Computing device 504 may execute method 400 for action recognition to perform action recognition. In one embodiment, computing device 504 may be a general-purpose computing device, such as a graphics processing unit (GPU), a central processing unit (CPU), a tensor processing unit (TPU), etc. In one embodiment, storage device 502 may be a hard disk, a flash memory, a random access memory (RAM), etc., or a combination thereof. In one embodiment, apparatus 500 may be implemented in an embedded environment, such as a mobile phone, a smart watch, a wearable device, a switch panel, a single-chip microcomputer, a microcontroller unit (MCU), etc.
[0065] FIG6 is a block diagram of a system 600 for action recognition according to an embodiment of the present invention. System 600 may include an infrared transmitter 602, an infrared receiver 604, a storage device 606, and a computing device 608. The infrared signal emitted by infrared transmitter 602 is reflected by the human body to infrared receiver 604, which collects the infrared signal and provides the collected infrared signal to storage device 606. Storage device 606 can store infrared signals collected by multiple infrared receivers and provide the infrared signals to computing device 608. Computing device 608 can execute method 400 for action recognition to perform action recognition. In one embodiment, computing device 608 can be a general-purpose computing device, such as a graphics processing unit (GPU), a central processing unit (CPU), a tensor processing unit (TPU), etc. In one embodiment, storage device 606 can be a hard disk, flash memory, RAM, etc., or a combination thereof. In one embodiment, storage device 606 and computing device 608 can be implemented in an embedded environment, such as a mobile phone, a smart watch, a wearable device, a switch panel, a single-chip microcomputer, a microcontroller unit (MCU), etc.
[0066] This application proposes a lightweight neural network action recognition framework based on an active infrared sensor array. Compared with the existing deep neural network framework, this application introduces a convolution module, a GRU module, and an attention mechanism module, which fully considers the temporal and spatial features and better models the time series data. This application improves the expressiveness of the model by combining the convolution module and the GRU module, while considering both local and global features; by introducing the attention mechanism module, the model is more focused on key temporal features and spatial positions, thereby improving the accuracy of the model; the combination of the three can adapt to different types of time series data, different numbers of infrared receivers, different positions of infrared transmitters, etc., and flexibly adjust the complexity and performance of the model. For example, by adjusting the parameters, action recognition with different sampling frequencies and different motion speeds can be achieved.
[0067] References throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all references to the same embodiment, but may refer to the same embodiment. Furthermore, in one or more embodiments, as will be apparent to one of ordinary skill in the art from this disclosure, the particular features, structures, or characteristics may be combined in any suitable manner.
[0068] Similarly, it should be appreciated that in the description of exemplary embodiments of the present invention, various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. On the contrary, as reflected in the appended claims, inventive aspects lie in fewer features than all of the features of a single preceding disclosed embodiment. Accordingly, the claims appended hereto are hereby expressly incorporated into this detailed description, with each claim itself representing a separate embodiment of the present invention.
[0069] Furthermore, although some embodiments described herein include some features included in other embodiments but not other features included in other embodiments, combinations of features from different embodiments are intended to fall within the scope of the present invention and form different embodiments as will be understood by those skilled in the art. For example, in the appended claims, any of the claimed embodiments may be used in any combination.
[0070] As used herein, module refers to any combination of hardware, software, and / or firmware. As an example, a module includes hardware such as a microcontroller associated with a non-transient medium, and the non-transient medium is used to store code suitable for being executed by the microcontroller. Therefore, in one implementation, reference to a module refers to hardware that is specifically configured to identify and / or execute code to be stored on a non-transient medium. In addition, in another implementation, the use of a module refers to a non-transient medium comprising code that is specifically adapted to be executed by a microcontroller to perform a predetermined operation. And as can be inferred, in yet another implementation, the term module can refer to a combination of a microcontroller and a non-transient medium. Typically, the boundaries of modules illustrated as separate can vary and potentially overlap. For example, a first module and a second module can share hardware, software, firmware, or a combination thereof while potentially retaining some independent hardware, software, or firmware.
[0071] Certain portions of the embodiments may be provided as a computer program product, which may include a computer-readable medium having computer program instructions stored thereon, which may be used to program a computer (or other electronic device) to be executed by one or more processors to perform processes according to certain embodiments. The computer-readable medium may include, but is not limited to, a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic or optical card, a flash memory, or other types of computer-readable media suitable for storing electronic instructions. In addition, the embodiments may also be downloaded as a computer program product, wherein the program may be transferred from a remote computer to a requesting computer. In some embodiments, a non-transitory computer-readable storage medium has data stored thereon representing a sequence of instructions that, when executed by a processor, causes the processor to perform certain operations.
[0072] Embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementations. Embodiments of the present invention may be implemented as a computer program or program code executed on a programmable system comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0073] It will be apparent to those skilled in the art that various modifications and variations may be made to the above exemplary embodiments of the present invention without departing from the spirit and scope of the present invention. Therefore, it is intended that the present invention cover modifications and variations of the present invention that fall within the scope of the appended claims and their equivalent technical solutions.
Claims
1. A method for action recognition, characterized in that: include: receiving training data, the training data comprising first infrared signals corresponding to a plurality of infrared receivers and action recognition results corresponding to the first infrared signals; Constructing an action recognition model, the action recognition model includes a data preprocessing module, a convolution module, a gated recurrent unit (GRU) module, an attention mechanism module and a classification module, The data preprocessing module is used to divide the input data of the action recognition model into multiple data blocks, each of the multiple data blocks includes infrared signals collected by multiple infrared receivers for k consecutive times; Wherein, for each data block in the multiple data blocks: The convolution module is used to extract spatial features of the current data block to obtain the current convolution module output; The attention mechanism module is used to extract context information from the current convolution module output and the previous GRU module output to obtain the current spatial attention feature; The GRU module is used to perform feature scanning on the current spatial attention feature and the previous GRU module output to obtain the current GRU module output; and The classification module is used to perform action classification on the current GRU module output for the last data block among the multiple data blocks to obtain an action recognition result; and The action recognition model is trained using the training data to obtain a trained action recognition model.
2. The method according to claim 1, characterized in that The convolution module extracts spatial features from the current data block by performing the following operations on the current data block: Using a first convolution kernel to perform single-channel feature extraction on the current data block to obtain a first feature; Using a second convolution kernel to perform inter-channel feature extraction on the current data block to obtain a second feature; Superimposing the first feature and the second feature and applying a first nonlinear activation function to obtain a third feature; Apply a point convolution layer, a second nonlinear activation function and a depth convolution layer to the third feature to obtain The fourth characteristic; Applying a pooling layer to the current data block to obtain a fifth feature; superimposing the fourth feature and the fifth feature to obtain a sixth feature; as well as The MB3 module is applied to the sixth feature and superimposed with the fifth feature, and a batch normalization layer and a third non-linear activation function are applied to obtain the current convolution module output.
3. The method according to claim 2, characterized in that The first convolution kernel includes one or more of 5×1, 7×1, 9×1, 11×1, and 13×1 convolution kernels, and the second convolution kernel is a 3×3 convolution kernel.
4. The method according to claim 2, characterized in that The first non-linear activation function, the second non-linear activation function, and the third non-linear activation function are any one of RELU, RELU6, and Sigmoid functions.
5. The method according to claim 2, characterized in that The number of the MB3 modules is 3.
6. The method according to any one of claims 1, characterized in that The classification module performs action classification including applying a global average pooling layer or a fully connected layer.
7. The method according to any one of claims 1 to 6, characterized in that Further including: receiving a second infrared signal collected by a plurality of infrared receivers; as well as The second infrared signal is input into the trained motion recognition model as the input data to obtain a motion recognition result.
8. A device for action recognition, comprising: A storage device for storing infrared signals corresponding to a plurality of infrared receivers; as well as A computing device, configured to execute the method according to any one of claims 1 to 7.
9. A system for action recognition, comprising: Infrared transmitter, used for transmitting infrared signals; Multiple infrared receivers for collecting infrared signals; A storage device, used for storing infrared signals collected by the plurality of infrared receivers; as well as A computing device, configured to execute the method according to any one of claims 1 to 7.
10. A computer readable medium having stored thereon computer program instructions which, when executed, cause a computer to perform the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Human body posture recognition method and system based on convolution and gated recurrent neural network
CN110610158A
Attention mechanism convolutional neural network-based infrared target classification method
CN111401473A
Infrared target identification method based on multi-channel convolutional neural network
CN115422968A
Method, device and system for action recognition and computer readable medium
CN117612250A
Systems and methods for video paragraph captioning using hierarchical recurrent neural networks
US20170127016A1
Cited By
Human body activity identification method and system based on lightweight Transform encoder
CN122332732A