Gesture recognition method and device, intelligent wearable equipment and storage medium

Through the gesture recognition method of event stream data conversion and event sparseness information processing, the problem of low accuracy and real-time in high-speed dynamic gesture recognition is solved, and efficient gesture recognition is achieved in complex environments.

CN120340112APending Publication Date: 2025-07-18ZHUHAI MOJIE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510217181.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has low accuracy and real-time performance in gesture recognition under high-speed dynamic gestures, complex backgrounds and low light.

Method used

The event stream data is used to convert the event image to be identified, and the accuracy and real-timeness of gesture recognition are improved through feature extraction and feature encoding processing based on event sparseness information.

Benefits of technology

Reduce the loss of gesture action information in high-speed dynamic scenarios, improve gesture recognition accuracy, and reduce invalid calculations by focusing on significantly changing areas, improving real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340112A_ABST
    Figure CN120340112A_ABST
Patent Text Reader

Abstract

The invention provides a gesture recognition method and device, intelligent wearable equipment and a computer readable storage medium. The gesture recognition method comprises the steps that a to-be-recognized event image of a target scene is acquired, the to-be-recognized event image is obtained by converting event flow data of the target scene, and the target scene comprises a gesture; performing feature extraction processing on the to-be-identified event image to obtain a preliminary feature map of the to-be-identified event image; on the basis of event sparseness information of the preliminary feature map, performing feature coding processing on the preliminary feature map to obtain image features of the event image to be identified; and performing gesture classification based on the image features to obtain a gesture category of the event image to be recognized. According to the invention, the gesture recognition accuracy and real-time performance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and particularly relates to a gesture recognition method, device, smart wearable device, and computer-readable storage medium. Background Art

[0002] With the development of artificial intelligence and computer vision technologies, gesture recognition technology has been widely applied in multiple fields such as augmented reality, virtual reality, intelligent interaction, and robot control. In related technologies, RGB images or depth images are collected and gesture recognition is performed using image classification models such as neural networks. However, the inventors of the present application have found in the actual research and development process that: this method of gesture recognition using image classification models has low accuracy and real-time performance in situations such as high-speed dynamic gestures, complex backgrounds, fast movements, and low light. Summary of the Invention

[0003] The present application provides a gesture recognition method, device, smart wearable device, and computer-readable storage medium, which can improve the accuracy and real-time performance of gesture recognition.

[0004] In a first aspect, the present application provides a gesture recognition method, the method including:

[0005] Obtaining a to-be-recognized event image of a target scene, where the to-be-recognized event image is obtained by converting event stream data of the target scene, and the target scene includes gestures;

[0006] Performing feature extraction processing on the to-be-recognized event image to obtain a preliminary feature map of the to-be-recognized event image;

[0007] Based on the event sparsity information of the preliminary feature map, performing feature encoding processing on the preliminary feature map to obtain an image feature of the to-be-recognized event image;

[0008] Performing gesture classification based on the image feature to obtain a gesture category of the to-be-recognized event image.

[0009] In a second aspect, the present application provides a gesture recognition device, the gesture recognition device including:

[0010] An obtaining unit, configured to obtain a to-be-recognized event image of a target scene, where the to-be-recognized event image is obtained by converting event stream data of the target scene, and the target scene includes gestures;

[0011] An extraction unit, configured to perform feature extraction processing on the to-be-recognized event image to obtain a preliminary feature map of the to-be-recognized event image;

[0012] An encoding unit, configured to perform feature encoding processing on the preliminary feature map based on the event sparsity information of the preliminary feature map, so as to obtain the image features of the image of the event to be recognized;

[0013] A classification unit, configured to perform gesture classification based on the image features to obtain the gesture category of the image of the event to be recognized.

[0014] In a third aspect, the present application further provides an intelligent wearable device, which includes a processor and a memory. A computer program is stored in the memory. When the processor calls the computer program in the memory, it executes any gesture recognition method provided by the present application.

[0015] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. The computer program is loaded by a processor to execute the gesture recognition method.

[0016] In the present application, in a first aspect, by extracting image features from the image of the event to be recognized for gesture recognition, since the image of the event to be recognized is obtained by converting the event stream data of the target scene, and the event stream data can capture the pixel changes of high-speed moving objects in a high-speed dynamic scene, the complete dynamic information of the target scene in the high-speed dynamic scene can be captured through the image of the event to be recognized, thereby reducing the problem of inaccurate gesture recognition caused by the loss of gesture action information that is likely to occur in a high-dynamic scene and improving the accuracy of gesture recognition. In a second aspect, by performing feature encoding processing based on the event sparsity information, since the event sparsity information can reflect the temporal and spatial distribution characteristics of pixel point changes, the regions with significant changes can be focused on based on the event sparsity information, the sparse features in the image of the event to be recognized can be adaptively captured, important spatio-temporal regions can be concerned, and the processing in the spatio-temporal regions with fewer pixel point changes can be reduced, thereby avoiding a large number of invalid calculation problems caused by the sparse characteristics of the event image and improving the real-time performance of gesture recognition. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 It is a structural schematic diagram of an intelligent wearable device provided by an embodiment of the present application;

[0019] Figure 2 It is a schematic flowchart of a gesture recognition method provided by an embodiment of the present application;

[0020] Figure 3 It is a schematic diagram for explaining the process of event image conversion provided in an embodiment of the present application;

[0021] Figure 4 It is a schematic diagram of the principle structure of a trained gesture classification model provided in an embodiment of the present application;

[0022] Figure 5 It is a schematic diagram of the principle framework of an attention encoding structure provided in an embodiment of the present application;

[0023] Figure 6 It is a schematic diagram of the principle framework of a scoring sub-module and a selection sub-module provided in an embodiment of the present application;

[0024] Figure 7 It is a schematic diagram of the principle framework of an attention sub-module provided in an embodiment of the present application;

[0025] Figure 8 It is a schematic diagram for explaining the window division of a preliminary feature map provided in an embodiment of the present application;

[0026] Figure 9 It is a schematic diagram of the structural embodiment of a gesture recognition device provided in an embodiment of the present application. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0028] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may change according to the actual situation.

[0029] In the description of the embodiments of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the embodiments of the present application, "a plurality" means two or more, unless otherwise specifically defined.

[0030] The following description is provided to enable any person skilled in the art to implement and use this application. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that this application can be implemented without using these specific details. In other instances, well-known processes are not elaborated in detail to avoid obscuring the description of the embodiments of this application with unnecessary details. Therefore, this application is not intended to be limited to the embodiments shown, but rather to be in line with the broadest scope consistent with the principles and features disclosed in the embodiments of this application.

[0031] Embodiments of this application provide a gesture recognition method, apparatus, intelligent wearable device, and computer-readable storage medium. Among them, the gesture recognition apparatus can be integrated in the intelligent wearable device. Among them, the intelligent wearable device can be an intelligent glasses, an intelligent helmet, etc. The intelligent glasses can be AR (augmented reality) glasses, VR (Virtual Reality) glasses, MR (Mixed Reality) glasses, XR (eXtended Reality) glasses, etc. The intelligent helmet can be an AR helmet, etc.

[0032] The execution subject of the gesture recognition method in the embodiments of this application can be the gesture recognition apparatus provided in the embodiments of this application, or an intelligent wearable device integrated with the gesture recognition apparatus. Among them, the gesture recognition apparatus can be implemented in a hardware or software manner.

[0033] The following will describe in detail some embodiments of this application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0034] Figure 1 It is a structural schematic block diagram of an intelligent wearable device provided by an embodiment of this application.

[0035] As Figure 1 shown, the intelligent wearable device 100 includes a processor 101 and a memory 102. The processor 101 and the memory 102 are connected through a bus 103, and this bus is, for example, an I2C (Inter-integrated Circuit) bus.

[0036] Specifically, the processor 101 is used to provide computing and control capabilities to support the operation of the entire intelligent wearable device 100. The processor 101 can be a Central Processing Unit (CPU), and this processor 101 can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or this processor can also be any conventional processor, etc.

[0037] Specifically, the memory 102 can be a Flash chip, Read-Only Memory (ROM), magnetic disk, optical disc, USB flash drive, or mobile hard disk, etc.

[0038] Those skilled in the art can understand that Figure 1 the structure shown in

[0039] is only a block diagram of some structures related to the solution of the embodiment of the present application, and does not constitute a limitation on the intelligent wearable device to which the solution of the embodiment of the present application is applied. The specific intelligent wearable device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0040] Obtain the image of the event to be recognized in the target scene, where the image of the event to be recognized is obtained by converting the event stream data of the target scene, and the target scene contains gestures; perform feature extraction processing on the image of the event to be recognized to obtain a preliminary feature map of the image of the event to be recognized; based on the event sparsity information of the preliminary feature map, perform feature encoding processing on the preliminary feature map to obtain the image features of the image of the event to be recognized; perform gesture classification based on the image features to obtain the gesture category of the image of the event to be recognized.

[0041] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described intelligent wearable device can refer to the corresponding process in the following embodiments of the gesture recognition method, and will not be elaborated herein.

[0042] Hereinafter, taking Figure 1 the intelligent wearable device shown in

[0043] as the execution subject of this gesture recognition method as an example, the gesture recognition method provided by the embodiments of the present application will be introduced in detail. For the sake of simplicity and convenience of description, the execution subject will be omitted in the subsequent method embodiments. Figure 2 , Figure 2 Please refer to

[0044] which is a schematic flowchart of a gesture recognition method provided by the embodiments of the present application. The gesture recognition method includes steps 201 to 203, where:

[0045] 201. Obtain the image of the event to be recognized in the target scene.

[0046] The image of the event to be recognized is obtained by converting the event stream data of the target scene.

[0047] The target scene is a scene where gesture recognition is to be performed. For example, it can be a virtual reality scene for interactive control based on gestures. The target scene contains gestures.

[0048] (1) In some embodiments, the image of the event to be recognized is obtained by converting the event stream data collected by the event camera in real time. At this time, step 201 may specifically include the following steps 2011A to 2013A:

[0049] 2011A. Obtain the latest event of the event stream data of the target scene.

[0050] The polarity (polarity, p) is used to represent the type of brightness change of the pixel point (such as the direction of brightness change). Exemplarily, it can be divided into positive polarity and negative polarity. For example, when the polarity of the latest event is positive polarity, it means that the brightness change of the event pixel point is an increase (i.e., the brightness becomes brighter); another example is that when the polarity of the previous event is negative polarity, it means that the brightness change of the pixel point corresponding to the previous event is a decrease (i.e., the brightness becomes darker).

[0051] Among them, event stream data is used to record events of brightness changes of each pixel point in the target scene over time. When the brightness of a pixel point in the target scene changes, a data point (i.e., an event) will be generated and output. Each event (x, y, p, t) includes the pixel point position (x, y), polarity p, and timestamp t, so as to record the detailed information of the dynamic changes of the scene. Event stream data can be collected by an event camera. An event camera is an event-driven vision sensor. Different from traditional frame cameras, an event camera does not continuously capture image frames at fixed time intervals, but dynamically responds based on the brightness change of each pixel j. Whenever the brightness change of a certain pixel j exceeds a set threshold, it will generate a data point (i.e., an event). Each generated event includes the pixel point position (x, y), polarity p, and timestamp t, so that each event carries accurate spatio-temporal information about the brightness change of that pixel j. Exemplarily, every time an event camera generates a data point (i.e., an event), it is added to a data queue or buffer for storage. All the data points recorded in the data queue or buffer area are event stream data. The data queue or buffer area can store data points (i.e., events) in a time-increasing data format, so that the event stream data cached in the data queue or buffer area is an ordered set of data points (i.e., events). In step 2011A, the data point (i.e., event) with the largest corresponding timestamp can be read from the event stream data recorded in the data queue or buffer area as the latest event.

[0052] 2012A. Update the pixel values of the maintenance image of the target scene according to the pixel point position and polarity of the latest event to obtain an updated maintenance image.

[0053] Among them, the maintenance image refers to an image used to record event stream data.

[0054] Exemplarily, the conversion process of the event image is as follows:

[0055] 1. Initialize the image: Initialize a blank RGB image (for example, the R-channel pixel value, G-channel pixel value, and B-channel pixel value of all pixel point positions in the RGB image are initialized to 0). The resolution of this RGB image is the same as that of the event camera (for example, the resolution is 1280*720).

[0056] 2> Pixel value generation: Whenever an event (x, y, p, t) is generated, the pixel position (x, y) of the event is correspondingly converted to the corresponding pixel position in the RGB image, and the pixel value at the corresponding pixel position in the RGB image is modified according to the polarity p of the event. The pixel value at each pixel position in the RGB image can be represented in the format (r, g, b), where r represents the pixel value of the R channel, g represents the pixel value of the G channel, and b represents the pixel value of the B channel.

[0057] In some embodiments, the modification rule of the pixel value can be as follows:

[0058] When the polarity p of the event (x, y, p, t) is positive (e.g., p = 1), the pixel value of the R channel at the pixel position (x, y) corresponding to the event in the RGB image is modified to a first preset value (e.g., 255);

[0059] When the polarity p of the event (x, y, p, t) is negative (e.g., p = -1), the pixel value of the G channel at the pixel position (x, y) corresponding to the event in the RGB image is modified to a second preset value (e.g., 255).

[0060] In some embodiments, the modification rule of the pixel value can be as follows:

[0061] When the polarity p of the event (x, y, p, t) is positive (e.g., p = 1), the pixel value of the G channel at the pixel position (x, y) corresponding to the event in the RGB image is modified to a first preset value (e.g., 255);

[0062] When the polarity p of the event (x, y, p, t) is negative (e.g., p = -1), the pixel value of the R channel at the pixel position (x, y) corresponding to the event in the RGB image is modified to a second preset value (e.g., 255).

[0063] In some embodiments, the modification rule of the pixel value can be as follows:

[0064] When the polarity p of the event (x, y, p, t) is positive (e.g., p = 1), the pixel value of the G channel at the pixel position (x, y) corresponding to the event in the RGB image is modified to a first preset value (e.g., 255);

[0065] When the polarity p of the event (x, y, p, t) is negative (e.g., p = -1), the pixel value of the B channel at the pixel position (x, y) corresponding to the event in the RGB image is modified to a second preset value (e.g., 255).

[0066] And so on. It can be understood that the modification rule of the pixel value can be set according to the actual business scenario requirements and is not limited hereby.

[0067] Among them, the specific values of the first preset value and the second preset value can be set according to the requirements of the actual business scenario. The specific values of the first preset value and the second preset value can be the same (for example, both are 255) or different (for example, the first preset value is 255 and the second preset value is 250), and it is not limited here.

[0068] 3> Event image generation: Whenever an event is used to update the pixel value in the RGB image, the cumulative number of update events of the RGB image is detected. When the cumulative number of update events of the RGB image is greater than the preset number threshold (such as greater than 10,000 events), the current RGB image is output as the event image to be recognized, and the next round of processing is continued with reference to the foregoing method. Among them, the cumulative number of update events refers to the cumulative number of events (x, y, p, t) that have been used to update the RGB image. Updating the RGB image means converting according to the pixel point position (x, y) of the event to the corresponding pixel point position in the RGB image.

[0069] The following introduces some implementation methods of step 2012A:

[0070] ① In some embodiments, the modification rule of the pixel value is: if the latest event is positive polarity, the pixel value of the R channel is modified to the first preset value (such as 255); if the latest event is negative polarity, the pixel value of the G channel is modified to the second preset value (such as 255). Please refer to Figure 3 , Figure 3 is an explanatory schematic diagram of the event image conversion process provided in the embodiments of the present application. The initialized RGB image can be used as the maintenance image of the target scene. Refer to Figure 3 as shown in (a) in; each time the latest event of the event stream data of the target scene is obtained, step 2012A is entered. If the polarity p of the latest event is positive polarity and the pixel point position is (x2, y3), then in step 2012A, the pixel value of the R channel at the pixel point position (x2, y3) in the maintenance image is modified to the first preset value (such as 255). Refer to Figure 3 as shown in (b) in, the pixel value at the pixel point position (x2, y3) can be modified from (0, 0, 0) to (255, 0, 0), so as to realize the modification of the pixel value of the R channel to the first preset value. If the polarity p of the latest event is negative polarity and the pixel point position is (x4, y4), then in step 2012A, the pixel value of the G channel at the pixel point position (x4, y4) in the maintenance image is modified to the second preset value (such as 255). Refer to Figure 3 as shown in (c) in, the pixel value at the pixel point position (x4, y4) can be modified from (0, 0, 0) to (0, 255, 0), so as to realize the modification of the pixel value of the G channel to the second preset value. At this time, the modified maintenance image can be used as the updated maintenance image.

[0071] ② In some embodiments, the modification rule of pixel values is as follows: if the latest event is positive polarity, the pixel value of the G channel is modified to a first preset value (such as 255); if the latest event is negative polarity, the pixel value of the B channel is modified to a second preset value (such as 255). The initialized RGB image can be used as the maintenance image of the target scene. Each time the latest event of the event stream data of the target scene is obtained, step 2012A is entered. If the polarity p of the latest event is positive polarity and the pixel point position is (x2, y3), then in step 2012A, the pixel value of the G channel at the pixel point position (x2, y3) in the maintenance image is modified to the first preset value (such as 255). If the polarity p of the latest event is negative polarity and the pixel point position is (x4, y4), then in step 2012A, the pixel value of the B channel at the pixel point position (x4, y4) in the maintenance image is modified to the second preset value (such as 255).

[0072] 2013A. Until the cumulative number of updated events in the updated maintenance image is greater than a preset number threshold, the updated maintenance image is used as the event image to be recognized of the target scene.

[0073] Wherein, the cumulative number of updated events refers to the number of events used for cumulative update of the pixel values of the maintenance image from the initial maintenance image (such as the initialized RGB image) to the updated maintenance image.

[0074] For example, if the latest event in step 2012A is the 10000th event of the event stream data, then the cumulative number of updated events in the updated maintenance image is 10000. If the cumulative number of updated events is greater than a preset number threshold (such as 10000), then the updated maintenance image is used as the event image to be recognized of the target scene, and the gesture recognition process of steps 202 - 204 is entered.

[0075] 202. Perform feature extraction processing on the event image to be recognized to obtain a preliminary feature map of the event image to be recognized.

[0076] Exemplarily, the feature extraction module in the trained gesture classification model in the embodiments of the present application can be used to perform feature extraction processing on the event image to be recognized to obtain a preliminary feature map of the event image to be recognized.

[0077] To better understand this embodiment, the trained gesture classification model in this embodiment is introduced below. As Figure 4 shown,[[]] Figure 4 is a schematic diagram of the principle structure of the trained gesture classification model provided in the embodiments of the present application. The trained gesture classification model can include a feature extraction module, a feature encoding module, and a classification module. The trained gesture classification model can be obtained by training based on a preset gesture classification model. The working principles of each module are as follows:

[0078] 1. Feature extraction module, which is used to perform preliminary feature extraction on the image of the event to be recognized to obtain a preliminary feature map.

[0079] In some embodiments, the feature extraction module includes a plurality of sequentially connected convolutional layers (for example, including 4 convolutional layers, as Figure 4 shown), and uses the plurality of sequentially connected convolutional layers to perform feature extraction to obtain a preliminary feature map, so that low-level spatial features can be extracted from the image of the event to be recognized, providing basic feature information for the subsequent feature encoding module. In this way, on the one hand, through the stacked layers of each convolutional layer of the feature extraction module, the model can perform feature extraction and combination at a higher level, gradually learning more complex image patterns such as gesture contours in the image, etc.; on the other hand, a pooling layer can be added after each convolutional layer of the feature extraction module. By combining the pooling layer for pooling operations, the size of the preliminary feature map can be reduced, the feature calculation amount can be reduced, and the calculation efficiency can be improved.

[0080] 2. Feature encoding module, which is used to perform further feature encoding processing according to the preliminary feature map and the image of the event to be recognized to obtain the image features of the image of the event to be recognized.

[0081] In some embodiments, the feature encoding module may include one or more attention encoding structures. Exemplarily, a sparse Transformer structure can be used as the attention encoding structure.

[0082] In some embodiments, the feature encoding module may include one or more attention encoding structures (as Figure 4 shown, for example, the feature encoding module may include 2 attention encoding structures), and at least one convolutional layer (as Figure 4 shown, for example, the feature encoding module may include 2 convolutional layers). At this time, in the feature encoding module, after the feature encoding processing is performed according to the preliminary feature map and the image of the event to be recognized through the attention encoding structure, the feature map obtained by further processing through the convolutional layer of the feature encoding module is used as the output of the feature encoding module, so that the output of the feature encoding module can be used as the image features of the image of the event to be recognized.

[0083] Among them, there are various implementation methods for the attention encoding structure. Exemplarily, they include Method 1 and Method 2:

[0084] Method 1: In some embodiments, each attention encoding structure may include a scoring sub-module, a selection sub-module, and an attention sub-module. At this time, the attention feature map T a output by the attention sub-module can be directly used as the output of the attention encoding structure (that is, the attention encoding result).

[0085] Mode 2: In some embodiments, such as Figure 5 shown, Figure 5 is a schematic diagram of a principle framework of the attention encoding structure provided in the embodiments of the present application. Each attention encoding structure may include a scoring sub-module, a selection sub-module, an attention sub-module, and an optimization sub-module. At this time, the result processed by the optimization sub-module can be used as the output of the attention encoding structure (i.e., the attention encoding result).

[0086] The scoring sub-module, the selection sub-module, the attention sub-module, and the optimization sub-module will be introduced separately below. Please refer to Figure 6 and Figure 7 , Figure 6 is a schematic diagram of a principle framework of the scoring sub-module and the selection sub-module provided in the embodiments of the present application, Figure 7 is a schematic diagram of a principle framework of the attention sub-module provided in the embodiments of the present application.

[0087] 2.1. Scoring sub-module, which is used to score each pixel coding block in the preliminary feature map according to the preliminary feature map and the image of the event to be recognized, and obtain the preliminary score of each pixel coding block. In some embodiments, the scoring sub-module includes a control unit, a response unit, and a weighting unit.

[0088] Among them, the pixel coding block i represents the i-th pixel coding block among N pixel coding blocks, where i is a positive integer greater than 0 and less than or equal to N, and N represents the total number of multiple pixel coding blocks included in the preliminary feature map.

[0089] Among them, the control unit is used to generate a control factor for each pixel coding block i in the preliminary feature map according to the image of the event to be recognized. The control unit specifically includes a pooling layer and an exponential linear layer. The image of the event to be recognized first undergoes a pooling operation through the pooling layer to obtain a pooled feature map, so that the receptive field of the event data (i.e., the size of the event data area) in the pooled feature map matches the receptive field of each pixel coding block i; then, according to the position of each pixel coding block i, the non-zero pixel ratio at the corresponding position in the pooled feature map is obtained (for example, assuming that the preliminary feature map includes 8*10 pixel coding blocks, and the pixel coding block i is at the m-th row and n-th column position in the preliminary feature map, and the total number of pixels at the position of the pixel coding block corresponding to the m-th row and n-th column in the pooled feature map is 20, among which 5 pixels have a pixel value of 0 and 15 pixels have a non-zero pixel value, then at the position of the pixel coding block corresponding to the m-th row and n-th column in the pooled feature map, the non-zero pixel ratio is counted as 15 / 20, as the event sparsity r of the pixel coding block i). In this way, through the event sparsity of each pixel coding block i, the sparsity of the event subset within a specific time range and polarity can be reflected.

[0090] Among them, the response unit is used to calculate the response value of each pixel coding block i in the preliminary feature map. The response unit includes a linear layer and an activation layer (the activation layer can use the ReLU activation function). The preliminary feature map first passes through the linear layer and then through the activation layer to calculate the response value of each pixel coding block i in the preliminary feature map.

[0091] Among them, the weighting unit is used to calculate the weighted score of each feature window according to the pixel coding blocks included in each feature window.

[0092] Exemplarily, as Figure 6 shown, in the scoring sub-module, on the one hand, through the control unit, a control factor for each pixel coding block i in the preliminary feature map is generated according to the image of the event to be recognized; on the other hand, the response value of each pixel coding block i in the preliminary feature map is calculated by the response unit; then, according to the control factor F of each pixel coding block i and the response value R of each pixel coding block i, the preliminary score of each pixel coding block i in the preliminary feature map is calculated with reference to the following formula 1.

[0093]

[0094] In formula 1, S i represents the preliminary score of pixel coding block i (used to indicate the importance of pixel coding block i), R represents the response value of pixel coding block i (R = Relu(W R ·T + b R ))), F represents the control factor of pixel coding block i (F = exp(W F )·r), W R , b R respectively represent the weight and bias of the linear layer, T represents the coding information of pixel coding block i, a represents the hyperparameter of the control coefficient level, W F represents the weight of the control factor F of pixel coding block i, and r represents the event sparsity of pixel coding block i.

[0095] 2.2. Selection sub-module, which is used to select windows and pixel coding blocks according to the preliminary scores of each pixel coding block, and obtain the selected feature windows and the selected pixel coding blocks of the selected feature windows. In some embodiments, the selection sub-module includes a competition unit composed of p-norm normalization and Softmax activation function, a window selection unit and a pixel block selection unit.

[0096] Among them, the competition unit is used to calculate the normalized score St of each pixel coding block (refer to Figure 6 , abbreviated as coding block score) and the normalized score Sw of each feature window (refer to Figure 6, abbreviated as window scoring). Exemplarily, the competition unit can calculate the normalized score St of each pixel encoding block using p-norm normalization (such as L2 normalization). The competition unit maps the preliminary score of each pixel encoding block i to the norm space through p-norm normalization (for example, when p = 2, that is, L2 norm normalization). The preliminary scores of the 1st, 2nd, …, Nth pixel encoding blocks can be represented as S1, S2, …, Sn respectively. Taking "the p-norm normalization method is L2 normalization" as an example, the normalized score St of each pixel encoding block i can be calculated as follows: where represents the normalized score of the i-th pixel encoding block). And so on. When the normalized scores St of each pixel encoding block are obtained, the normalized scores St of all pixel encoding blocks within each feature window are summed (such as direct summation or weighted summation), and the result is used as the normalized score Sw of each feature window. Further, the competition unit can also set a Softmax activation function to convert the normalized scores into a probability distribution, such that each normalized score is between 0 and 1, the sum of the normalized scores of all pixel encoding blocks is 1, and the sum of the normalized scores of all feature windows is 1. This helps to amplify the differences in the normalized scores St between different pixel encoding blocks and the differences in the normalized scores Sw between different feature windows.

[0097] Among them, the window selection unit is used to select feature windows according to the normalized scores Sw of each feature window output by the competition unit and the weighted scores of each feature window output by the weighting unit to obtain the selected feature windows.

[0098] Among them, the pixel block selection unit is used to select pixel blocks according to the normalized scores St of each pixel encoding block in each selected feature window to obtain the selected pixel encoding blocks of each selected feature window.

[0099] 2.3. The attention sub-module is used to perform attention encoding on the selected pixel encoding blocks of the selected feature window to obtain an attention feature map. In some instances, the attention sub-module can adopt the MS-WSA (Multi-Scale Window-based Self-Attention) attention mechanism structure. In this way, when facing window features T of inconsistent sizes s , it can perform a padding operation on the window features T s to unify window features T of different sizes s to the same scale, so as to effectively perform self-attention operations on window features T of different sizes s , improve the feature extraction ability of the model, and further improve the accuracy of gesture classification using image features. As Figure 7 shown, Figure 7It is a schematic diagram of the working principle of the attention sub-module provided in the embodiments of the present application. Exemplarily, the overall working process of the attention sub-module is as follows: First, the selected pixel coding blocks in each selected feature window are spliced to obtain the window features of each selected feature window; then, for the window feature T of each selected feature window s , as shown in Formula 2 below, the attention sub-module performs a padding operation on the window feature T of the selected feature window s to obtain the padded window feature T of the selected feature window p , thereby unifying the window features T of different sizes s to the same scale, so as to unify the lengths of the selected pixel coding blocks of all selected feature windows to the same scale, and can effectively perform self-attention operations on the window features T of different sizes s , improve the feature extraction ability of the model, and further improve the accuracy of gesture classification using image features. Then, the padded window features T of each selected feature window p are parallelly subjected to multi-head attention operations to obtain the attention feature map T a . Exemplarily, the process of performing multi-head attention operations can refer to Formulas 3 and 4 below, including first mapping the padded window feature T p to three vectors: query vector Q, key vector K, and value vector V, and then passing through a masking operation and a Softmax function; finally, an un-padding operation is performed to obtain the attention feature map T a . In some embodiments, the attention feature map T output by the attention sub-module can be directly a used as the output of the attention encoding structure (i.e., the attention encoding result).

[0100] T p = Pad(T s ) Formula 2

[0101]

[0102] T a = UnPad(Softmax(Mask + QK T )V) Formula 4

[0103] In Formulas 2 to 4, Pad represents the padding operation, T s represents the window feature within the selected feature window, T p represents the padded window feature of the selected feature window, and Q, K, and V are the query vector, key vector, and value vector respectively mapped based on the padded window feature T p , and W Q , W K , W VWeight matrices for Q, K, and V respectively, and UnPad represents the unpadding operation.

[0104] 2.4. Optimization sub-module, which is used to optimize the attention feature map to obtain the image features. Exemplarily, the optimization sub-module includes an MLP (Multilayer Perceptron) layer and a CB (Contextual Broadcasting) layer. The attention feature map is further processed by the MLP layer of the optimization sub-module, and the contextual broadcasting operation is performed by the CB layer. Finally, the result processed by the optimization sub-module is used as the output of the attention encoding structure (i.e., the attention encoding result) to further optimize the expression and transmission of features. Among them, the contextual broadcasting operation through the CB layer can effectively capture the long-range dependence relationship between pixel encoding blocks; the further sparsification operation of the attention feature map through the MLP layer can reduce the computational complexity, so that the extracted features can achieve a balance between efficiency and accuracy. In some embodiments, the output of the attention encoding structure can be used as the output of the feature encoding structure, so as to obtain the image features of the event image to be recognized. In some embodiments, the feature encoding module may include an attention encoding structure and a convolutional layer. At this time, the output of the attention encoding structure can be further processed through the convolutional layer of the feature encoding structure, and the result processed by the convolutional layer of the feature encoding structure is used as the output of the feature encoding structure, so as to obtain the image features of the event image to be recognized.

[0105] 3. Classification module, which is used to perform gesture classification according to the image features of the event image to be recognized, and obtain the gesture category of the target scene. In some embodiments, the classification module may include a fully connected layer and a Softmax activation function. After the image features of the event image to be recognized pass through the fully connected layer and the Softmax activation function of the classification module, the gesture category of the event image to be recognized is output.

[0106] In some embodiments, in step 202, the feature extraction module in the trained gesture classification model can be used to perform feature extraction processing on the event image to be recognized, and obtain the preliminary feature map of the event image to be recognized.

[0107] 203. Based on the event sparsity information of the preliminary feature map, perform feature encoding processing on the preliminary feature map to obtain the image features of the event image to be recognized.

[0108] In some embodiments, step 203 may specifically include the following steps 2031 to 2035:

[0109] 2031. Perform window partitioning on multiple pixel encoding blocks of the preliminary feature map to obtain multiple feature windows.

[0110] For example, as Figure 8 shown, assume that the preliminary feature map includes 8 * 10 pixel encoding blocks. According to the preset window size (or the preset number of windows), such as the preset window size being 4 * 5 pixel encoding blocks, the preliminary feature map can be window-divided to obtain 4 feature windows each containing 4 * 5 pixel encoding blocks, as Figure 8 shown by the feature windows 1, 2, 3, and 4 in

[0111] 2032. Obtain the event sparsity of each pixel encoding block.

[0112] There are multiple implementation manners for step 2032. Exemplarily, they include the following manners ① and ②:

[0113] ① In some embodiments, the non-zero pixel ratio is statistically calculated as the event sparsity. The event sparsity r of each pixel encoding block i can be obtained through the feature encoding module of the trained gesture classification model. At this time, step 2032 can specifically include: inputting the event image to be recognized into the control unit of the scoring sub-module of the feature encoding module. In the control unit, the event image to be recognized first undergoes a pooling operation through the pooling layer to obtain a pooled feature map, so that the event data receptive field of the pooled feature map matches the receptive field of each pixel encoding block i; then, according to the position of each pixel encoding block i, the non-zero pixel ratio at the corresponding position of the pooled feature map is statistically calculated as the event sparsity r of each pixel encoding block i. In this way, through the event sparsity of each pixel encoding block i, the sparsity of the event subset within a specific time range and polarity can be reflected. For example, assume that the preliminary feature map includes 8 * 10 pixel encoding blocks, and the pixel encoding block i is at the m-th row and n-th column position in the preliminary feature map. The total number of pixels at the position of the m-th row and n-th column pixel encoding block in the pooled feature map is 20, among which 5 pixels have a pixel value of 0 and 15 pixels have a pixel value not equal to 0. Then, at the position of the m-th row and n-th column pixel encoding block in the pooled feature map, the non-zero pixel ratio (such as 15 / 20) is statistically calculated as the event sparsity r of the pixel encoding block i.

[0114] Among them, the non-zero pixel ratio refers to the pixel ratio at the position corresponding to the pixel encoding block i (such as the m-th row and n-th column pixel encoding block) in the pooled feature map where the pixel value is not equal to zero.

[0115] ②In some embodiments, the number of non-zero pixels is counted as the event sparsity. The event sparsity r of each pixel encoding block i can be obtained through the feature encoding module of the trained gesture classification model. At this time, step 2032 may specifically include: inputting the event image to be recognized into the control unit of the scoring sub-module of the feature encoding module. In the control unit, the event image to be recognized first undergoes a pooling operation through a pooling layer to obtain a pooled feature map, so that the receptive field of the event data in the pooled feature map matches the receptive field of each pixel encoding block i; then, according to the position of each pixel encoding block i, the number of non-zero pixels at the corresponding position in the pooled feature map is counted as the event sparsity r of each pixel encoding block i.

[0116] Among them, the number of non-zero pixels refers to the number of pixels with non-zero pixel values at the position corresponding to the pixel encoding block i (such as the pixel encoding block in the m-th row and n-th column) in the pooled feature map.

[0117] 2033. Score each pixel encoding block based on the event sparsity of each pixel encoding block to obtain the preliminary score of each pixel encoding block.

[0118] Exemplarily, the preliminary score of each pixel encoding block can be obtained through the feature encoding module of the trained gesture classification model. At this time, step 2033 may specifically include: inputting the divided preliminary feature map into the response unit of the scoring sub-module of the feature encoding module, and calculating the response value of each pixel encoding block i in the preliminary feature map through the response unit; then, according to the control factor F of each pixel encoding block i and the response value R of each pixel encoding block i, the preliminary score of each pixel encoding block i in the preliminary feature map is calculated with reference to formula 1.

[0119] 2034. Determine the selected pixel encoding blocks of the selected feature windows according to the preliminary scores of each pixel encoding block and the multiple feature windows.

[0120] There are various implementation manners for step 2034. Exemplarily, it includes:

[0121] <1>In some embodiments, the selected feature window is determined according to the normalized score of the feature window and the weighted score of the feature window. At this time, step 2034 may specifically include: First, input the divided preliminary feature map into the weighted unit of the selection sub-module of the feature encoding module, and calculate the weighted scores of each feature window in the preliminary feature map through the weighted unit; calculate the normalized score of each feature window and the normalized score of each pixel encoding block according to the preliminary scores of each pixel encoding block through the competition unit (for the detailed method of determining the normalized score, reference can be made to the description of the relevant part of the competition unit above, which will not be elaborated here); then, determine the selected feature window from multiple feature windows of the preliminary feature map through the window selection unit according to the normalized score of each feature window and the weighted score of each feature window. Next, according to the normalized score of each pixel encoding block, obtain the pixel encoding blocks whose normalized scores are greater than the second preset score threshold from the pixel encoding blocks falling into the selected feature window as the selected pixel encoding blocks of the selected feature window. In this way, on the one hand, since the preliminary score of each pixel encoding block is related to the event sparsity of each pixel encoding block, the normalized scores of the feature window and the pixel encoding block can be calculated through the p-norm normalization of the competition unit, and the distribution of the scores can be dynamically adjusted according to the event sparsity, so that the normalized scores match the spatial distribution and sparse characteristics of the events; on the other hand, the difference between the normalized scores can be further amplified through the Softmax activation function of the competition unit, which can improve the discriminability of the normalized scores; on the third hand, through the design of the competition unit, under the action of the normalized scores of the competition unit, in the event information sparse scenario, a small number of important features can be more effectively screened out; while in the event information dense scenario, since the feature windows and pixel encoding blocks with high normalized scores will be selected, more feature windows and pixel encoding blocks can be retained.

[0122] Among them, the higher the normalized score of the feature window, the greater the probability that the feature window is selected; conversely, the lower the normalized score of the feature window, the smaller the probability that the feature window is selected.

[0123] Among them, the higher the weighted score of the feature window, the greater the probability that the feature window is selected; conversely, the lower the weighted score of the feature window, the smaller the probability that the feature window is selected.

[0124] Among them, the higher the preliminary normalized score of the pixel encoding block, the greater the probability that the pixel encoding block is selected; conversely, the lower the normalized score of the pixel encoding block, the smaller the probability that the pixel encoding block is selected.

[0125] <2>In some embodiments, the selected feature window is determined according to the normalized score of the feature window. At this time, step 2034 may specifically include: First, the competition unit determines the normalized score of each pixel coding block according to the preliminary score of each pixel coding block; according to the preliminary score of each pixel coding block, determine the normalized score of each feature window among the multiple feature windows; then, the window selection unit obtains, according to the normalized score of each feature window, the feature windows with a normalized score greater than the first preset score threshold from the multiple feature windows as the selected feature windows. Then, according to the normalized score of each pixel coding block, obtain the pixel coding blocks with a normalized score greater than the second preset score threshold from the pixel coding blocks falling into the selected feature window as the selected pixel coding blocks of the selected feature window.

[0126] 2035. Perform encoding processing on the selected pixel coding blocks of the selected feature window to obtain the image features.

[0127] There are various implementation manners for step 2035. Exemplarily, it includes:

[0128] 1> In some embodiments, the attention encoding feature map is obtained through encoding processing by the attention sub-module as the image feature of the image to be recognized. At this time, step 2035 may specifically include: Through the attention sub-module, perform attention encoding on the selected pixel coding blocks of the selected feature window to obtain an attention feature map as the image feature of the image to be recognized.

[0129] 2> In some embodiments, further encoding processing is respectively performed by the attention sub-module and the optimization sub-module to obtain the image feature of the image to be recognized. At this time, step 2035 may specifically include: Through the attention sub-module, perform attention encoding on the selected pixel coding blocks of the selected feature window to obtain an attention feature map; perform optimization processing on the attention feature map through the optimization sub-module to obtain the image feature of the image to be recognized.

[0130] In some embodiments, "through the attention sub-module, perform attention encoding according to the selected feature window and the selected pixel coding blocks of the selected feature window to obtain an attention feature map" may be as follows: Concatenate the selected pixel coding blocks of the selected feature window to obtain the window feature of the selected feature window; perform a padding operation on the window feature of the selected feature window to obtain the padded window feature of the selected feature window; perform a multi-head attention operation on the padded window feature of the selected feature window to obtain an attention feature map. For specific implementation details, reference may be made to the relevant description of the attention sub-module above, and details will not be elaborated here.

[0131] In some embodiments, the detailed implementation details of "optimizing the attention feature map through the optimization sub-module to obtain the image features of the event image to be recognized" can refer to the relevant description of the optimization sub-module above, and will not be elaborated here.

[0132] 204. Perform gesture classification based on the image features to obtain the gesture category of the event image to be recognized.

[0133] Exemplarily, gesture classification can be performed through a trained gesture classification model. At this time, the image features can be input into the classification module of the trained gesture classification model. After passing through the fully connected layer and the Softmax activation function of the classification module, the gesture category of the event image to be recognized is output.

[0134] Furthermore, after determining the gesture category of the event image to be recognized, interactive control of the target scene can be performed based on the gesture category of the event image to be recognized. For example, interactive control of AR glasses can be performed.

[0135] From the above content, it can be seen that on the one hand, by extracting image features from the event image to be recognized for gesture recognition, since the event image to be recognized is transformed from the event stream data of the target scene, and the event stream data can capture the pixel changes of high-speed moving objects in a high-speed dynamic scene, the complete dynamic information of the target scene in the high-speed dynamic scene can be captured through the event image to be recognized, thereby reducing the problem of inaccurate gesture recognition caused by the loss of gesture action information easily occurring in high-dynamic scenes and improving the accuracy of gesture recognition. On the other hand, by performing feature encoding processing based on the event sparsity information, since the event sparsity information can reflect the temporal and spatial distribution characteristics of pixel point changes, the regions with significant changes can be focused on based on the event sparsity information, the sparse features in the event image to be recognized can be adaptively captured, important spatio-temporal regions can be concerned, and processing in spatio-temporal regions with fewer pixel point changes can be reduced, thereby avoiding a large number of invalid calculation problems caused by the sparse characteristics of the event image and improving the real-time performance of gesture recognition.

[0136] In addition, in order to better implement the gesture recognition method in the embodiments of the present application, based on the gesture recognition method, an embodiment of a gesture recognition device is further provided in the embodiments of the present application, as Figure 9 shown, which is a schematic structural diagram of an embodiment of the gesture recognition device provided by the embodiments of the present application. The gesture recognition device 900 includes:

[0137] An acquisition unit 901, configured to acquire an event image to be recognized of a target scene, where the event image to be recognized is transformed from the event stream data of the target scene, and the target scene includes gestures;

[0138] An extraction unit 902 is configured to perform feature extraction processing on the image of the event to be recognized to obtain a preliminary feature map of the image of the event to be recognized;

[0139] An encoding unit 903 is configured to perform feature encoding processing on the preliminary feature map based on the event sparsity information of the preliminary feature map to obtain the image features of the image of the event to be recognized;

[0140] A classification unit 904 is configured to perform gesture classification based on the image features to obtain the gesture category of the image of the event to be recognized.

[0141] In some embodiments, the obtaining unit 901 is configured to:

[0142] Obtain the latest event of the event stream data of the target scene;

[0143] Update the pixel values of the maintenance image of the target scene according to the pixel point position of the latest event and the polarity of the latest event to obtain an updated maintenance image;

[0144] Until the cumulative number of updated events of the updated maintenance image is greater than a preset number threshold, use the updated maintenance image as the image of the event to be recognized of the target scene.

[0145] In some embodiments, the preliminary feature map includes a plurality of pixel encoding blocks, the event sparsity information includes the event sparsity of each pixel encoding block in the preliminary feature map, and the encoding unit 903 is configured to:

[0146] Perform window division on the multiple pixel encoding blocks of the preliminary feature map to obtain a plurality of feature windows;

[0147] Obtain the event sparsity of each pixel encoding block;

[0148] Score each pixel encoding block based on the event sparsity of each pixel encoding block to obtain a preliminary score of each pixel encoding block;

[0149] Determine the selected pixel encoding blocks of the selected feature windows according to the preliminary scores of each pixel encoding block and the plurality of feature windows;

[0150] Perform encoding processing on the selected pixel encoding blocks of the selected feature windows to obtain the image features.

[0151] In some embodiments, the encoding unit 903 is configured to:

[0152] Perform a pooling operation on the image of the event to be recognized to obtain a pooled feature map;

[0153] According to the position of each pixel encoding block i, obtain the non-zero pixel ratio at the corresponding position of the pooling feature map as the event sparsity of each pixel encoding block i, where the pixel encoding block i represents the i-th pixel encoding block among N pixel encoding blocks, i is a positive integer greater than 0 and less than or equal to N, and N represents the total number of multiple pixel encoding blocks included in the preliminary feature map.

[0154] In some embodiments, the encoding unit 903 is configured to:

[0155] Obtain the response value of each pixel encoding block;

[0156] Based on the event sparsity of each pixel encoding block, obtain the control factor of each pixel encoding block;

[0157] Based on the response value of each pixel encoding block and the control factor of each pixel encoding block, determine the preliminary score of each pixel encoding block.

[0158] In some embodiments, the encoding unit 903 is configured to:

[0159] Determine the normalized score of each pixel encoding block according to the preliminary score of each pixel encoding block;

[0160] Determine the normalized score of each feature window among the multiple feature windows according to the preliminary score of each pixel encoding block;

[0161] According to the normalized score of each feature window, obtain, from the multiple feature windows, the feature window whose normalized score is greater than the first preset score threshold as the selected feature window;

[0162] According to the normalized score of each pixel encoding block, obtain, from the pixel encoding blocks falling into the selected feature window, the pixel encoding block whose normalized score is greater than the second preset score threshold as the selected pixel encoding block of the selected feature window.

[0163] In some embodiments, the encoding unit 903 is configured to:

[0164] Perform attention encoding on the selected pixel encoding blocks of the selected feature window to obtain an attention feature map;

[0165] Perform optimization processing based on the attention feature map to obtain the image feature.

[0166] In some embodiments, the encoding unit 903 is configured to:

[0167] Stitch the encoded pixel blocks of the selected pixels in the selected feature window to obtain the window feature of the selected feature window;

[0168] Perform a filling operation on the window feature of the selected feature window to obtain the filled window feature of the selected feature window;

[0169] Perform a multi-head attention operation on the filled window feature of the selected feature window to obtain an attention feature map.

[0170] In specific implementation, each of the above units can be implemented as an independent entity, or can be combined arbitrarily and implemented as the same or several entities. For the specific implementation of each of the above units, reference can be made to the embodiments of the gesture recognition method described above, which will not be elaborated here.

[0171] Those of ordinary skill in the art can understand that all or part of the steps in the above gesture recognition method can be completed by instructions, or by controlling relevant hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0172] Therefore, an embodiment of the present application provides a computer-readable storage medium, in which multiple computer programs are stored. These computer programs can be loaded by a processor to execute any one of the gesture recognition methods provided by the embodiments of the present application.

[0173] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0174] In the above embodiments of the gesture recognition device, computer-readable storage medium, and smart wearable device, the descriptions of each embodiment have their own emphases. For the parts not elaborated in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and the beneficial effects that can be brought by the above-described gesture recognition device, computer-readable storage medium, smart wearable device and their corresponding units can refer to the description of the gesture recognition method in the above embodiments, which will not be elaborated here specifically.

[0175] The above has introduced in detail a gesture recognition method, device, smart wearable device, and computer-readable storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present application.

Claims

1. A gesture recognition method, characterized in that, The method includes: Obtaining an event image to be recognized in a target scenario, where the event image to be recognized is obtained by converting event stream data of the target scenario, and gestures are included in the target scenario; Performing feature extraction processing on the event image to be recognized to obtain a preliminary feature map of the event image to be recognized; Based on the event sparsity information of the preliminary feature map, performing feature encoding processing on the preliminary feature map to obtain image features of the event image to be recognized; Performing gesture classification based on the image features to obtain a gesture category of the event image to be recognized.

2. The gesture recognition method according to claim 1, wherein The obtaining of the event image to be recognized in the target scenario includes: Obtaining the latest event of the event stream data of the target scenario; Updating the pixel values of the maintained image of the target scenario according to the pixel point position of the latest event and the polarity of the latest event to obtain an updated maintained image; Until the cumulative number of updated events in the updated maintained image is greater than a preset number threshold, taking the updated maintained image as the event image to be recognized in the target scenario.

3. The gesture recognition method according to claim 1, wherein The preliminary feature map includes a plurality of pixel encoding blocks, and the event sparsity information includes the event sparsity of each pixel encoding block in the preliminary feature map; The performing of feature encoding processing on the preliminary feature map based on the event sparsity information of the preliminary feature map to obtain image features of the event image to be recognized includes: Performing window partitioning on the plurality of pixel encoding blocks of the preliminary feature map to obtain a plurality of feature windows; Obtaining the event sparsity of each pixel encoding block; Based on the event sparsity of each pixel encoding block, scoring each pixel encoding block to obtain a preliminary score of each pixel encoding block; According to the preliminary score of each pixel encoding block and the plurality of feature windows, determining the selected pixel encoding blocks of the selected feature windows; Performing encoding processing on the selected pixel encoding blocks of the selected feature windows to obtain the image features.

4. The gesture recognition method according to claim 3, wherein The obtaining of the event sparsity of each pixel encoding block includes: Performing a pooling operation on the event image to be recognized to obtain a pooled feature map; According to the position of each pixel encoding block i, obtaining the non-zero pixel ratio at the corresponding position of the pooled feature map as the event sparsity of each pixel encoding block i, where pixel encoding block i represents the i-th pixel encoding block among N pixel encoding blocks, i is a positive integer greater than 0 and less than or equal to N, and N represents the total number of the plurality of pixel encoding blocks included in the preliminary feature map.

5. The gesture recognition method according to claim 3, wherein The scoring of each pixel encoding block based on the event sparsity of each pixel encoding block to obtain a preliminary score of each pixel encoding block includes: Obtaining the response value of each pixel encoding block; Based on the event sparsity of each pixel encoding block, obtaining the control factor of each pixel encoding block; Based on the response value of each pixel encoding block and the control factor of each pixel encoding block, determining the preliminary score of each pixel encoding block.

6. The gesture recognition method according to claim 3, wherein The determining of the selected pixel encoding blocks of the selected feature windows according to the preliminary score of each pixel encoding block and the plurality of feature windows includes: Determine the normalized score of each pixel coding block according to the preliminary score of each pixel coding block; Determine the normalized score of each feature window in the multiple feature windows according to the preliminary score of each pixel coding block; According to the normalized score of each feature window, obtain the feature windows with normalized scores greater than the first preset score threshold from the multiple feature windows as the selected feature windows; According to the normalized score of each pixel coding block, obtain the pixel coding blocks with normalized scores greater than the second preset score threshold from the pixel coding blocks falling into the selected feature windows as the selected pixel coding blocks of the selected feature windows.

7. The gesture recognition method according to claim 3, wherein The encoding process for the selected pixel coding blocks of the selected feature windows to obtain the image features includes: Perform attention encoding on the selected pixel coding blocks of the selected feature windows to obtain an attention feature map; Perform optimization processing based on the attention feature map to obtain the image features.

8. The gesture recognition method according to claim 7, wherein The performing attention encoding on the selected pixel coding blocks of the selected feature windows to obtain an attention feature map includes: Perform splicing on the selected pixel coding blocks of the selected feature windows to obtain the window features of the selected feature windows; Perform a padding operation on the window features of the selected feature windows to obtain the padded window features of the selected feature windows; Perform a multi-head attention operation on the padded window features of the selected feature windows to obtain an attention feature map.

9. An intelligent wearable device, characterized in that, It includes a processor and a memory, and a computer program is stored in the memory. When the processor calls the computer program in the memory, it executes the gesture recognition method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program is loaded by a processor to execute the gesture recognition method according to any one of claims 1 to 8.