A gesture recognition method, device and storage medium

CN115410274BActive Publication Date: 2026-10-09HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211049900.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2026-10-09
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

通过这种方式进行手势识别时,由于采用双分支网络搭建网络模型的架构,因此在进行手势识别时需要大量的计算量

Benefits of technology

[0024] It should be noted that the aforementioned computer instructions may be stored, in whole or in part, on a computer-readable storage medium. This computer-readable storage medium may be packaged together with the controller's processor, or it may be packaged separately from the controller's processor; this application does not impose any limitations on this.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115410274B_ABST
    Figure CN115410274B_ABST
Patent Text Reader

Abstract

The application discloses a gesture recognition method and device and a storage medium, relates to the technical field of image processing, and can be used for improving the accuracy of gesture recognition and reducing the calculation amount of gesture recognition. The method comprises the following steps: acquiring hand key point feature information of each frame of image in multiple frames of images; the multiple frames of images are obtained by shooting a moving hand; inputting the hand key point feature information of each frame of image in the multiple frames of images, preset spatial sequence serial number information and preset time sequence serial number information into a gesture recognition model based on a graph convolutional neural network to obtain a gesture recognition result; wherein the spatial sequence serial number information is used for marking the serial number of the hand key point of each frame of image in the multiple frames of images, and the time sequence serial number information is used for marking the frame number corresponding to the spatial information of each frame of image in the multiple frames of images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a gesture recognition method, apparatus and storage medium. Background Technology

[0002] With the rapid development of science and technology, gesture recognition is becoming increasingly widespread and has become a research hotspot in the field of human-computer interaction. For example, in the home appliance industry, gestures can be used to control appliances and execute commands corresponding to user gestures; similarly, in the transportation sector, the actions of traffic police can be recognized to obtain corresponding traffic instructions. Typically, gesture recognition in images is achieved through image processing and recognition.

[0003] In related technologies, one approach to gesture recognition is to construct key hand features and directly input these features into a gesture classification model to obtain the gesture recognition result. However, when using this method for gesture recognition, the model input only includes key hand features. If the changes in key hand features are small across multiple consecutive frames, it becomes difficult to recognize dynamic gestures corresponding to multiple frames, resulting in low accuracy in gesture recognition.

[0004] Another approach to gesture recognition is to obtain the hand joint sequence from an image, and then use a two-branch network to obtain the correlation between the hand joints, thereby performing gesture recognition. However, this method requires a significant amount of computation due to the two-branch network architecture. Summary of the Invention

[0005] This application provides a gesture recognition method, device, and storage medium that can achieve high accuracy gesture recognition with low computational load.

[0006] In a first aspect, a gesture recognition method is provided, comprising: acquiring hand key point feature information of each frame of a multi-frame image; the multi-frame images are obtained by capturing images of a moving hand; the hand key point feature information of a frame image includes the coordinates of the hand key points in that frame image and the motion information of the hand key points, wherein the motion information of the hand key points in that frame image is used to characterize the change of the coordinates of the hand key points in that frame image relative to the coordinates of the corresponding hand key points in the previous frame image; inputting the hand key point feature information of each frame of the multi-frame image, a preset spatial sequence number, and a preset temporal sequence number into a gesture recognition model based on a graph convolutional neural network to obtain a gesture recognition result; wherein the spatial sequence number is used to mark the sequence number of the hand key points in each frame of the multi-frame image, and the temporal sequence number is used to mark the frame number corresponding to the spatial information of each frame of the multi-frame image.

[0007] The technical solution provided in this application offers at least the following advantages: by using hand keypoint feature information, spatial sequence number information, and temporal sequence number information from multiple frames of images as input to the gesture recognition model, this multi-information input method incorporates more source data information from multiple frames of images compared to inputting only single hand keypoint features, thus improving the accuracy of gesture recognition. Furthermore, this application does not require the construction of a complex dual-branch network when building the gesture recognition model, thereby reducing the computational load during the gesture recognition process.

[0008] In some embodiments, the gesture recognition model based on graph convolutional neural networks includes a spatial information extraction sub-model, a temporal information extraction sub-model, and a classification sub-model based on graph convolutional neural networks; the above-mentioned inputting the hand key point feature information, preset spatial sequence number information, and preset temporal sequence number information of each frame of multiple images into the gesture recognition model based on graph convolutional neural networks to obtain the gesture recognition result includes: inputting the hand key point feature information and preset spatial sequence number information of each frame of multiple images into the spatial information extraction sub-model to obtain the hand key point adjacency matrix of each frame of the image; each frame of the image The adjacency matrix of hand keypoints in an image is used to represent the connection relationship between any two hand keypoints in each frame. The adjacency matrix of hand keypoints in each frame is processed by a graph convolutional neural network of the spatial information extraction sub-model to obtain the spatial information of each frame in the multi-frame image. The spatial information and time series sequence information of each frame in the multi-frame image are input into the temporal information extraction sub-model to obtain the temporal information of the multi-frame image. The temporal information of the multi-frame image is used to represent the spatial information change of the multi-frame image from the first frame to the last frame. The temporal information of the multi-frame image is input into the classification sub-model to obtain the gesture recognition result.

[0009] It should be understood that the aforementioned adjacency matrix of hand key points is generated based on a graph convolutional neural network. Therefore, compared to manually designed adjacency matrices, this graph convolutional neural network-generated adjacency matrix of hand key points is more adaptable to gesture recognition in different scenarios and has strong generalization capabilities. Furthermore, the spatial information of each frame in the aforementioned multi-frame images contains relevant information about the gesture to be recognized. However, since gestures are a dynamic and continuous process, it is necessary to construct the dependencies between multiple frames through a temporal information extraction sub-model, thereby obtaining the temporal information of the multi-frame images. Further, gesture recognition is performed based on the temporal information of the multi-frame images after establishing the dependencies, in order to improve the accuracy of gesture recognition.

[0010] In some embodiments, the above-mentioned inputting the hand keypoint feature information and preset spatial sequence number information of each frame of the multi-frame images into the spatial information extraction sub-model to obtain the hand keypoint adjacency matrix of each frame image includes: for each frame of the multi-frame images, the hand keypoint feature information and spatial sequence number information of the frame image are fused through the spatial information extraction sub-model to obtain the fused information of the frame image; the fused information of the frame image is input into the first convolutional layer and the second convolutional layer of the spatial information extraction sub-model respectively to obtain the matrix of the fused information of the frame image output by the first convolutional layer and the matrix of the fused information of the frame image output by the second convolutional layer; the matrix of the fused information of the frame image output by the first convolutional layer is transposed and multiplied with the matrix of the fused information of the frame image output by the second convolutional layer to obtain the hand keypoint adjacency matrix of the frame image.

[0011] In some embodiments, the above-mentioned inputting the spatial information and time series sequence number information of each frame in the multi-frame image into the time series information extraction sub-model to obtain the time series information of the multi-frame image includes: marking the frame number corresponding to each frame in the multi-frame image according to the preset time series sequence number information; inputting the spatial information of the multi-frame image after marking the frame number into the spatial information pooling layer in the time series information extraction sub-model to obtain the spatial information of the multi-frame image after pooling; inputting the spatial information of the multi-frame image after pooling into the third convolutional layer and the fourth convolutional layer in sequence, and then establishing the dependency relationship between each frame in the multi-frame image through the third convolutional layer and the fourth convolutional layer to obtain the spatial information of the multi-frame image after establishing the dependency relationship; inputting the spatial information of the multi-frame image after establishing the dependency relationship into the time series information extraction sub-model to obtain the time series information of the multi-frame image.

[0012] In some embodiments, the acquisition of hand keypoint feature information of each frame in a multi-frame image includes: for each frame in a multi-frame image with a preset number of frames, inputting the frame image into a regression model to obtain the hand keypoint coordinates of the frame image; if the frame image is not the first frame in the multi-frame image, then for any hand keypoint coordinate of the frame image, determining the keypoint motion information corresponding to the hand keypoint of the frame image based on the change in the hand keypoint coordinate of the frame image compared to the hand keypoint coordinate of the previous frame image; or, if the frame image is the first frame in the multi-frame image, then the keypoint motion information corresponding to the frame image is the preset keypoint motion information.

[0013] In some embodiments, the method further includes: performing multiple gesture recognition operations on the video to be detected within a preset time period; the gesture recognition operations are used to obtain gesture recognition results in multiple frames of images in the video to be detected; if the multiple gesture recognition results obtained by performing multiple gesture recognition operations within the preset time period are inconsistent, the gesture recognition result that appears most frequently among the multiple gesture recognition results is determined as the final gesture recognition result.

[0014] It should be understood that the more times a gesture recognition result appears within a preset time period, the more likely it is to correspond to the actual dynamic gesture within that preset time period. Therefore, when multiple gesture recognition results are inconsistent, the gesture recognition result that appears most frequently within the preset time period can be determined as the final gesture recognition result, thereby improving the accuracy of gesture recognition.

[0015] Secondly, a gesture recognition device is provided, comprising: an acquisition module for acquiring hand key point feature information of each frame of a multi-frame image; the multi-frame images are obtained by capturing images of a moving hand; the hand key point feature information of a frame image includes the coordinates of the hand key points in that frame image and the motion information of the hand key points, the motion information of the hand key points in that frame image being used to characterize the change of the coordinates of the hand key points in that frame image relative to the coordinates of the corresponding hand key points in the previous frame image; and a processing module for inputting the hand key point feature information of each frame of the multi-frame image, a preset spatial sequence number, and a preset time sequence number into a gesture recognition model based on a graph convolutional neural network to obtain a gesture recognition result; wherein, the spatial sequence number is used to mark the sequence number of the hand key points in each frame of the multi-frame image, and the time sequence number is used to mark the frame number corresponding to the spatial information of each frame of the multi-frame image.

[0016] In some embodiments, the above processing module is specifically used to input the hand keypoint feature information and preset spatial sequence number information of each frame of the multi-frame images into the spatial information extraction sub-model to obtain the hand keypoint adjacency matrix of each frame image; the hand keypoint adjacency matrix of each frame image is used to represent the connection relationship between any two hand keypoints in each frame image; the graph convolutional neural network of the spatial information extraction sub-model is used to process the hand keypoint adjacency matrix of each frame image to obtain the spatial information of each frame image in the multi-frame images; the spatial information and time sequence number information of each frame image in the multi-frame images are input into the temporal information extraction sub-model to obtain the temporal information of the multi-frame images; the temporal information of the multi-frame images is used to represent the spatial information change of the multi-frame images from the first frame image to the last frame image; the temporal information of the multi-frame images is input into the classification sub-model to obtain the gesture recognition result.

[0017] In some embodiments, the above-described processing module is specifically used to, for each frame of a multi-frame image, fuse the hand keypoint feature information and spatial sequence number information of the frame image through a spatial information extraction sub-model to obtain the fused information of the frame image; input the fused information of the frame image into the first convolutional layer and the second convolutional layer of the spatial information extraction sub-model respectively to obtain the matrix of the fused information of the frame image output by the first convolutional layer and the matrix of the fused information of the frame image output by the second convolutional layer; transpose the matrix of the fused information of the frame image output by the first convolutional layer and multiply it with the matrix of the fused information of the frame image output by the second convolutional layer to obtain the adjacency matrix of the hand keypoints of the frame image.

[0018] In some embodiments, the above processing module is specifically used to mark the frame number corresponding to each frame of the multi-frame images according to the preset time series sequence information; input the spatial information of the multi-frame images after marking the frame number to the spatial information pooling layer in the temporal information extraction sub-model to obtain the spatial information of the multi-frame images after pooling; input the spatial information of the multi-frame images after pooling to the third convolutional layer and the fourth convolutional layer in sequence, and then establish the dependency relationship between each frame of the multi-frame images through the third convolutional layer and the fourth convolutional layer to obtain the spatial information of the multi-frame images after establishing the dependency relationship; input the spatial information of the multi-frame images after establishing the dependency relationship to the temporal information pooling layer in the temporal information extraction sub-model to obtain the temporal information of the multi-frame images.

[0019] In some embodiments, the acquisition module is specifically used to input each frame of a multi-frame image with a preset number of frames into a regression model to obtain the hand keypoint coordinates of the frame image; if the frame image is not the first frame of the multi-frame image, then for any hand keypoint coordinate of the frame image, the keypoint motion information corresponding to the hand keypoint of the frame image is determined based on the change in the hand keypoint coordinate of the frame image compared to the hand keypoint coordinate of the previous frame image; or, if the frame image is the first frame of the multi-frame image, then the keypoint motion information corresponding to the frame image is the preset keypoint motion information.

[0020] In some embodiments, the above-mentioned processing module is further configured to perform multiple gesture recognition operations on the video to be detected within a preset time period; the gesture recognition operation is used to obtain gesture recognition results in multiple frames of images in the video to be detected; if the multiple gesture recognition results obtained by performing multiple gesture recognition operations within the preset time period are inconsistent, the gesture recognition result that appears most frequently among the multiple gesture recognition results is determined as the final gesture recognition result.

[0021] Thirdly, embodiments of this application provide a gesture recognition device, including: a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code including computer instructions; wherein, when the processor executes the computer instructions, it causes the gesture recognition device to perform a gesture recognition method as described in the first aspect and any of its possible design schemes.

[0022] Fourthly, embodiments of this application provide a computer-readable storage medium including computer instructions that, when executed on a computer, cause the computer to perform the methods provided in the first aspect and possible implementations.

[0023] Fifthly, embodiments of this application provide a computer program product containing computer instructions that, when executed on a computer, cause the computer to perform the methods provided in the first aspect and possible implementations described above.

[0024] It should be noted that the aforementioned computer instructions may be stored, in whole or in part, on a computer-readable storage medium. This computer-readable storage medium may be packaged together with the controller's processor, or it may be packaged separately from the controller's processor; this application does not impose any limitations on this.

[0025] For a detailed description of aspects two through five and their various implementations in this application, please refer to the detailed description in aspect one and its various implementations. The beneficial effects of aspects two through five and their various implementations can be found in the analysis of the beneficial effects of aspect one and its various implementations; they will not be repeated here. Attached Figure Description

[0026] Figure 1 This is a flowchart illustrating a gesture recognition method according to some embodiments. Figure 1 ;

[0027] Figure 2 This is a schematic diagram of a key hand point according to some embodiments;

[0028] Figure 3 This is a schematic diagram of the results of a gesture recognition model according to some embodiments;

[0029] Figure 4 This is a flowchart illustrating a gesture recognition method according to some embodiments. Figure 2 ;

[0030] Figure 5 This is a schematic diagram of a hand key point topology according to some embodiments;

[0031] Figure 6This is a flowchart illustrating a training method for a gesture recognition model according to some embodiments;

[0032] Figure 7 This is a schematic diagram of the structure of a gesture recognition device according to some embodiments. Figure 1 ;

[0033] Figure 8 This is a schematic diagram of the structure of a gesture recognition device according to some embodiments. Figure 2 . Detailed Implementation

[0034] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0035] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner. The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0036] To facilitate understanding, a brief introduction to the relevant concepts involved in this application will be provided first.

[0037] RGB: The RGB color mode is a commonly used color standard. It obtains various colors by changing the three color channels of red, green and blue and superimposing them. RGB represents the colors of the three channels of red, green and blue.

[0038] YUV: In the YUV image color space, Y represents the luminance component, U represents the hue component, and V represents the saturation component.

[0039] In the HIS (Hue-Saturation-Intensity) image color space, H represents the hue component, S represents the lightness component, and I represents the saturation component.

[0040] RGBD: An image containing RGB color space information and corresponding depth map information. It records the color imaging and depth information of the photographed area.

[0041] Graph Convolutional Networks (GCNs): GCNs extend the convolution operation from traditional data to graph data and play a central role in building many other complex graph neural network models.

[0042] The above is an introduction to some of the concepts involved in the embodiments of this application, which will not be repeated below.

[0043] As described in the background section, gesture recognition has become a research hotspot in the field of human-computer interaction. Typically, gesture recognition is achieved by processing and recognizing gestures within images. One approach to gesture recognition involves constructing key hand features and directly inputting these features into a gesture classification model to obtain the recognition result. However, this method uses only a single input parameter, resulting in low accuracy. Another approach employs a dual-branch network to obtain the relationships between hand joints for gesture recognition. However, this method requires a significant amount of computation due to its dual-branch network architecture.

[0044] To address this, this application provides a gesture recognition method, which includes: acquiring hand key point feature information of each frame in a multi-frame image; the multi-frame images are obtained by capturing images of a moving hand; the hand key point feature information of a frame image includes the coordinates of the hand key points in that frame image and the motion information of the hand key points, wherein the motion information of the hand key points in that frame image is used to characterize the change of the coordinates of the hand key points in that frame image relative to the coordinates of the corresponding hand key points in the previous frame image; inputting the hand key point feature information of each frame in the multi-frame image, a preset spatial sequence number, and a preset time sequence number into a gesture recognition model based on a graph convolutional neural network to obtain a gesture recognition result; wherein the spatial sequence number is used to mark the sequence number of the hand key points in each frame image in the multi-frame image, and the time sequence number is used to mark the frame number corresponding to the spatial information of each frame image in the multi-frame image.

[0045] As can be seen, the gesture recognition method proposed in this application uses hand key point feature information, spatial sequence number information, and temporal sequence number information from multiple frames of images as input to the gesture recognition model. Compared to inputting only hand key point features, this multi-information input method incorporates more source data information from multiple frames of images, thus improving the accuracy of gesture recognition. Furthermore, this application does not require the construction of a complex dual-branch network when building the gesture recognition model, thereby reducing the computational load during the gesture recognition process.

[0046] It should be noted that the gesture recognition method provided in this application can be applied to any gesture recognition scenario. For example, in the transportation field, gesture recognition can be used to obtain corresponding traffic instructions from traffic police officers' hand signals; in the field of home appliance technology or vehicle-machine interaction, gesture recognition can be used to set instructions for home appliances or vehicles based on user gestures; and in social situations, gesture recognition can be used to translate sign language for special groups. It should be understood that the embodiments of this application do not limit the specific application scenarios of the gesture recognition method.

[0047] The gesture recognition method provided in this application can be executed by a gesture recognition device. For example, the gesture recognition device can be a server. Alternatively, the gesture recognition device can be an electronic chip with image processing capabilities, such as a GPU (Graphics Processing Unit); or it can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), augmented reality (AR) / virtual reality (VR) device, etc. Optionally, the gesture recognition device can have a shooting function to acquire video or images to be processed in real time. Optionally, the gesture recognition device can also have a storage function to store images or images that need to be recognized by gestures, or to store gesture recognition results. Optionally, the gesture recognition device can also connect to the Internet to realize cloud gesture recognition. This application does not impose special limitations on the specific form of the gesture recognition device. The following description uses the gesture recognition device as a server as an example.

[0048] For ease of understanding, the gesture recognition method provided in this application will be described in detail below with reference to the accompanying drawings.

[0049] Figure 1 This application presents a gesture recognition method, which is applied in the field of image processing technology. For example... Figure 1 As shown, the gesture recognition method includes the following steps S101 to S102:

[0050] S101. Obtain the key feature information of the hand in each frame of the multi-frame image.

[0051] The hand keypoint feature information of a single frame includes the coordinates of the hand keypoints in that frame and the motion information of the hand keypoints. The motion information of the hand keypoints in a single frame is used to characterize the change in the coordinates of the hand keypoints in that frame relative to the corresponding coordinates of the hand keypoints in the previous frame.

[0052] In some embodiments, the aforementioned multi-frame images can be multiple frames from a video captured in real time by a camera device, or multiple frames from a video to be identified stored in a storage device. For example, in the field of home appliance technology, a camera device can capture a user's actions in real time, thereby performing gesture recognition based on multiple frames from the captured video. Alternatively, when watching videos of sign language communication by special groups, to facilitate the recognition of sign language meanings, relevant videos can be read from a storage device, thereby performing gesture recognition on multiple frames from those videos.

[0053] In some embodiments, the aforementioned multi-frame images can be multi-frame images with a preset number of frames. It should be understood that a dynamic gesture is a continuous action in time, therefore, relying solely on a single frame image cannot accurately identify a dynamic gesture. Furthermore, since the duration of a gesture is short, using too many frames during gesture recognition would increase the computational load. Therefore, the number of frames in the multi-frame images can be set to a preset number. For example, the preset number of frames can be 10. It should be noted that the setting of the preset number of frames should be related to the specific application scenario of the gesture recognition method. If the gesture to be recognized has a long duration, the preset number of frames can be set larger; if the gesture to be recognized has a short duration, the preset number of frames can be set smaller. This application does not limit the preset number of frames for the aforementioned multi-frame images.

[0054] In some embodiments, the acquisition of multiple frames can be multiple frames within a continuous time period, or multiple frames extracted periodically from the video to be detected in chronological order.

[0055] For example, periodically extracting multiple frames from a video to be detected in chronological order can be implemented as follows: as the video to be detected is played or captured, it is detected whether the playback time or capture time has reached the starting point of the image frame extraction period. The image frame extraction period is determined based on the number of frames in the video to be detected, with a consecutive number of image frames constituting one image frame extraction period. If the playback time or capture time is detected to have reached the starting point of the image frame extraction period, the image frames in the video to be detected at the point when the playback time or capture time reaches the starting point of the image frame extraction period are extracted, until the number of extracted image frames reaches a preset number. Based on this, gesture recognition in the video to be detected does not need to perform gesture recognition on all images in the video; instead, it can skip frames according to the image frame extraction period, thus saving computational resources for gesture recognition.

[0056] In some embodiments, before acquiring hand key point feature information from multiple frames of images, image preprocessing is performed on the multiple frames of images, such as image brightness transformation, sharpening, moiré removal, or Gaussian noise removal. This application does not limit the specific operations of image preprocessing.

[0057] In some embodiments, hand keypoints are used to locate key feature points in the human hand. For example, such as... Figure 2 The image shown is a schematic diagram of key points of the hand. Figure 2 Image (a) shows a diagram of a human hand; correspondingly, Figure 2 (b) in the middle shows Figure 2 Multiple key hand points corresponding to the hand in (a) of the image. Figure 2 In (b) of the diagram, each black dot corresponds to a key hand feature. It can be seen that... Figure 2 The key points of the hand were located in the diagram, which identified the main joints of the hand.

[0058] In some embodiments, the hand keypoint coordinates can be two-dimensional spatial coordinates. For example, when the format of the multi-frame images is RGB, YUV, or HSI, each frame of the multi-frame images is input into a two-dimensional keypoint regression model to obtain the two-dimensional hand keypoint coordinates in each frame of the multi-frame images.

[0059] In other embodiments, the hand keypoint coordinates can be three-dimensional spatial coordinates. For example, when the format of the multi-frame images is RGBD or similar, each frame of the multi-frame images is input into a three-dimensional keypoint regression model to obtain the three-dimensional hand keypoint coordinates in each frame of the multi-frame images.

[0060] In some examples, the aforementioned two-dimensional and three-dimensional keypoint regression models can be keypoint regression models constructed based on convolutional neural networks. It should be understood that this application does not limit the specific construction method of the keypoint regression model.

[0061] In some embodiments, the keypoint motion information of each frame in a multi-frame image can be determined according to the following methods Sa1 to Sa2:

[0062] Sa1. For each frame in a multi-frame image, if the frame is not the first frame in the multi-frame image, then for any hand key point coordinate in the frame, the key point motion information corresponding to the hand key point in the frame is determined based on the change in the hand key point coordinate in the frame compared to the hand key point coordinate in the previous frame.

[0063] For example, the key point motion information of this frame image satisfies the following relationship:

[0064] V t =P t -P t-1

[0065] Among them, V t P represents the motion information of key points in the t-th frame of the image. t P represents the keypoint coordinates of the t-th frame of the image. t-1 This represents the coordinates of key points in the (t-1)th frame of the image, where t is a positive integer greater than 1.

[0066] Sa2. For each frame in a multi-frame image, if the frame is the first frame in the multi-frame image, then the key point motion information corresponding to the frame is the preset key point motion information.

[0067] For example, the key point motion information of the frame image can be an all-zero matrix of the same type as the key point coordinates P1 of the frame image.

[0068] S102. Input the hand key point feature information, preset spatial sequence number information and preset time sequence number information of each frame of the multi-frame images into the gesture recognition model based on graph convolutional neural network to obtain the gesture recognition result.

[0069] Among them, the spatial sequence number information is used to mark the sequence number of the hand key points in each frame of the multi-frame image, and the time sequence number information is used to mark the frame number corresponding to the spatial information of each frame of the multi-frame image.

[0070] In some embodiments, the spatial sequence number information is related to the total number of hand keypoints identified in the keypoint regression model. For example, when the total number of hand keypoints is set to 21, the spatial sequence number is a 21*21 matrix, and each row of the matrix corresponding to the spatial sequence number corresponds to one hand keypoint.

[0071] In some embodiments, the time series sequence number information is related to the total number of frames in the multi-frame images in step S102. For example, when the total number of frames in the multi-frame images is 10, the time series sequence number is a 10*10 matrix information, and each row of the matrix corresponding to the time series sequence number corresponds to one frame in the multi-frame images.

[0072] In some embodiments, such as Figure 3 As shown, the aforementioned gesture recognition model based on graph convolutional neural networks includes a spatial information extraction sub-model, a temporal information extraction sub-model, and a classification sub-model based on graph convolutional neural networks. Further, as... Figure 4 As shown, step S102 above can be specifically implemented as steps S1021 to S1023:

[0073] S1021. Input the hand key point feature information and the preset spatial sequence number information of each frame of the multi-frame image into the spatial information extraction sub-model to obtain the adjacency matrix of the hand key points of each frame of the image; process the adjacency matrix of the hand key points of each frame of the image through the graph convolutional neural network of the spatial information extraction sub-model to obtain the spatial information of each frame of the multi-frame image.

[0074] The adjacency matrix of hand keypoints in each frame is used to represent the connection relationship between any two hand keypoints in that frame, and can be used to construct the topological structure of hand keypoints.

[0075] For example, Figure 5 The diagram shows a schematic representation of the topology of key points in a hand. (Refer to...) Figure 5 Based on the adjacency matrix of hand key points, each hand key point ( Figure 5 The black dots in the diagram are connected to form the topological structure of hand keypoints. It can be seen that the adjacency matrix of hand keypoints contains morphological information of the hand, which can be used for further gesture recognition.

[0076] In some examples, it is still referred to Figure 3The above-mentioned method of inputting the hand keypoint feature information and the preset spatial sequence number information of each frame of the multi-frame image into the spatial information extraction sub-model to obtain the hand keypoint adjacency matrix of each frame can be specifically implemented in the following steps: For each frame of the multi-frame image, the hand keypoint feature information and the spatial sequence number information of the frame are fused by the spatial information extraction sub-model to obtain the fused information of the frame; the fused information of the frame is input into the first convolutional layer and the second convolutional layer of the spatial information extraction sub-model respectively to obtain the matrix of the fused information of the frame output by the first convolutional layer and the matrix of the fused information of the frame output by the second convolutional layer; the matrix of the fused information of the frame output by the first convolutional layer is transposed and multiplied by the matrix of the fused information of the frame output by the second convolutional layer to obtain the hand keypoint adjacency matrix G of the frame.

[0077] It should be understood that the above-mentioned adjacency matrix of hand key points is an adjacency matrix generated by a graph convolutional neural network. Therefore, compared with manually designed adjacency matrices, this adjacency matrix of hand key points generated by graph convolutional neural networks is more suitable for gesture recognition in different scenarios and has strong generalization.

[0078] S1022. Input the spatial information and time series sequence information of each frame in the multi-frame image into the time series information extraction sub-model to obtain the time series information of the multi-frame image.

[0079] The temporal information of the aforementioned multi-frame images is used to characterize the spatial information changes of the multi-frame images from the first frame to the last frame.

[0080] In some examples, it is still referred to Figure 3 The above step S1022 can be specifically implemented as follows: according to the preset spatial sequence number information, the frame number corresponding to each frame image in the multi-frame image is marked; the spatial information of the multi-frame image after marking the frame number is input into the spatial information pooling layer in the temporal information extraction sub-model to obtain the spatial information of the multi-frame image after pooling; the spatial information of the multi-frame image after pooling is input into the third convolutional layer and the fourth convolutional layer in sequence, and then the dependency relationship between each frame image in the multi-frame image is established through the third convolutional layer and the fourth convolutional layer to obtain the spatial information of the multi-frame image after establishing the dependency relationship; the spatial information of the multi-frame image after establishing the dependency relationship is input into the temporal information pooling layer in the temporal information extraction sub-model to obtain the temporal information of the multi-frame image.

[0081] S1023. Input the temporal information of multiple frames of images into the classification sub-model to obtain the gesture recognition result.

[0082] The aforementioned classification sub-models can be constructed based on machine learning methods such as decision trees, support vector machines, or random forests, or they can be constructed based on fully connected (FC) layers and SoftMax functions. This application does not impose any restrictions on this.

[0083] It should be understood that the spatial information of each of the aforementioned multi-frame images contains information related to the gesture to be recognized. However, since a gesture is a dynamic and continuous process, it is necessary to construct the dependency relationship between the multi-frame images through a temporal information extraction sub-model, thereby obtaining the temporal information of the multi-frame images. Furthermore, gesture recognition is performed based on the temporal information of the multi-frame images after establishing the dependency relationship, in order to improve the accuracy of gesture recognition.

[0084] Figure 1 The provided technical solution offers at least the following advantages: by using hand keypoint feature information, spatial sequence number information, and temporal sequence number information from multiple frames of images as input to the gesture recognition model, this multi-information input method incorporates more source data information from multiple frames of images compared to inputting only hand keypoint features, thus improving the accuracy of gesture recognition. Furthermore, this application does not require the construction of a complex dual-branch network when building the gesture recognition model, thereby reducing the computational load during the gesture recognition process.

[0085] In some embodiments, such as Figure 6 As shown, the above gesture recognition model is trained according to the following steps S1 to S5:

[0086] S1. Obtain the gesture recognition model and sample set to be trained.

[0087] The sample set includes multiple training samples, and each training sample includes multiple frames of images and corresponding gesture type labels for the multiple frames of images.

[0088] S2. For each training sample in the sample set, input the multi-frame images in the training sample into the gesture recognition model to be trained, and obtain the gesture type recognition result corresponding to the multi-frame images in the training sample.

[0089] S3. Determine the loss value of the training samples based on the gesture type recognition results corresponding to the multi-frame images in the training samples and the gesture type labels of the training samples.

[0090] In some examples, the loss value of the training samples can be determined according to the following loss function:

[0091]

[0092] Where x represents the classification confidence of the gesture type identified from multiple frames of images in the training samples.

[0093] S4. Based on the loss values ​​of the training samples, determine whether the gesture recognition model has converged.

[0094] S5. If the gesture recognition model does not converge, update the parameters of the gesture recognition model according to the gradient descent method and repeat steps S2 to S5 above; or, if the gesture recognition model converges, determine the current gesture recognition model as the trained gesture recognition model.

[0095] Based on this, the parameters of the gesture recognition model can be updated using the backpropagation algorithm, thereby completing the training of the gesture recognition model.

[0096] In some embodiments, multiple gestures can be preset, and then gesture recognition can be performed on multiple frames of images based on these preset gestures to obtain gesture recognition results. The gesture recognition results can have multiple confidence levels, each confidence level corresponding to a preset gesture. The preset gesture corresponding to the highest confidence level is selected from the multiple confidence levels as the final gesture recognition result.

[0097] In other embodiments, considering that the duration of detected gestures is not fixed and varies depending on the speed of individual movements and the type of gesture, a suitable preset time period can be set for gesture recognition according to user needs. For example, in the fields of smart homes or vehicle-to-everything (V2X) interaction, user gestures can be recognized to control smart home appliances or in-vehicle devices to execute commands corresponding to those gestures. In this case, the gestures associated with the commands are generally short in duration. Therefore, the preset time period can be set relatively short. For example, the preset time period can be set to 1 second. It should be understood that the specific value of the preset time period is only an example, and the setting of the preset time period should be based on the specific application scenario; this application does not impose specific limitations on it.

[0098] Furthermore, to improve the accuracy of gesture recognition, multiple gesture recognition operations are performed on the video to be detected within a preset time period. The gesture recognition operations are used to obtain gesture recognition results in multiple frames of images in the video to be detected. If the multiple gesture recognition results obtained from the multiple gesture recognition operations within the preset time period are inconsistent, the gesture recognition result that appears most frequently among the multiple gesture recognition results is determined as the final gesture recognition result.

[0099] In some embodiments, the server executes the gesture recognition method shown in this application by default to save the user's selection time.

[0100] In other embodiments, the server enables or disables the gesture recognition function based on user instructions. Specifically, if a user instruction to enable the gesture recognition function is detected, the server executes the gesture recognition method described in this application; if a user instruction to disable the gesture recognition function is detected, the server executes the gesture recognition method described in this application. This adaptably meets the user's needs in different scenarios.

[0101] In some embodiments, if the server enables the gesture recognition function and does not detect user activity characteristics within a preset time period, the gesture recognition function is disabled; or, if the server disables the gesture recognition function and detects user activity characteristics within a preset time period, the gesture recognition function is enabled.

[0102] For example, if the preset duration is half an hour, then if no user gesture operation is detected within half an hour after the server enables the gesture recognition function, the gesture recognition function will be turned off; or, if no user gesture operation is detected within half an hour after the server disables the gesture recognition function, the gesture recognition function will be turned off.

[0103] In some examples, the server includes an infrared detection module that detects user activity features within a preset time period. Specifically, the infrared detection module of the server detects the presence of human activity features within a preset time period.

[0104] As can be seen, the above mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, the embodiments of this application provide corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the modules and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0105] like Figure 7 As shown, this application provides a gesture recognition device for performing... Figure 7 The gesture recognition method shown is illustrated. The gesture recognition device 300 includes an acquisition module 301 and a processing module 302.

[0106] The acquisition module 301 is used to acquire the hand key point feature information of each frame of a multi-frame image; the multi-frame images are obtained by capturing the moving hand; the hand key point feature information of a frame image includes the coordinates of the hand key points in that frame image and the motion information of the hand key points, and the motion information of the hand key points in a frame image is used to characterize the change of the coordinates of the hand key points in that frame image relative to the corresponding coordinates of the hand key points in the previous frame image.

[0107] The processing module 302 is used to input the hand key point feature information, preset spatial sequence number information and preset time sequence number information of each frame of the multi-frame images into the gesture recognition model based on graph convolutional neural network to obtain the gesture recognition result; wherein, the spatial sequence number information is used to mark the sequence number of the hand key points of each frame of the multi-frame images, and the time sequence number information is used to mark the frame number corresponding to the spatial information of each frame of the multi-frame images.

[0108] In some embodiments, the processing module 302 is specifically used to input the hand keypoint feature information and preset spatial sequence number information of each frame of the multi-frame images into the spatial information extraction sub-model to obtain the hand keypoint adjacency matrix of each frame image; the hand keypoint adjacency matrix of each frame image is used to represent the connection relationship between any two hand keypoints in each frame image; the hand keypoint adjacency matrix of each frame image is processed by the graph convolutional neural network of the spatial information extraction sub-model to obtain the spatial information of each frame image in the multi-frame images; the spatial information and time sequence number information of each frame image in the multi-frame images are input into the temporal information extraction sub-model to obtain the temporal information of the multi-frame images; the temporal information of the multi-frame images is used to represent the spatial information change of the multi-frame images from the first frame image to the last frame image; the temporal information of the multi-frame images is input into the classification sub-model to obtain the gesture recognition result.

[0109] In some embodiments, the processing module 302 is specifically configured to, for each frame of a multi-frame image, fuse the hand keypoint feature information and spatial sequence number information of the frame image through a spatial information extraction sub-model to obtain the fused information of the frame image; input the fused information of the frame image into the first convolutional layer and the second convolutional layer of the spatial information extraction sub-model respectively to obtain the matrix of the fused information of the frame image output by the first convolutional layer and the matrix of the fused information of the frame image output by the second convolutional layer; transpose the matrix of the fused information of the frame image output by the first convolutional layer and multiply it with the matrix of the fused information of the frame image output by the second convolutional layer to obtain the adjacency matrix of the hand keypoints of the frame image.

[0110] In some embodiments, the processing module 302 is specifically used to mark the frame number corresponding to each frame of the multi-frame images according to the preset time series sequence number information; input the spatial information of the multi-frame images after marking the frame number to the spatial information pooling layer in the temporal information extraction sub-model to obtain the spatial information of the multi-frame images after pooling; input the spatial information of the multi-frame images after pooling to the third convolutional layer and the fourth convolutional layer in sequence, and then establish the dependency relationship between each frame of the multi-frame images through the third convolutional layer and the fourth convolutional layer to obtain the spatial information of the multi-frame images after establishing the dependency relationship; input the spatial information of the multi-frame images after establishing the dependency relationship to the temporal information pooling layer in the temporal information extraction sub-model to obtain the temporal information of the multi-frame images.

[0111] In some embodiments, the acquisition module 301 is specifically used to input each frame of a multi-frame image with a preset number of frames into a regression model to obtain the hand keypoint coordinates of the frame image; if the frame image is not the first frame of the multi-frame image, then for any hand keypoint coordinate of the frame image, the keypoint motion information corresponding to the hand keypoint of the frame image is determined based on the change in the hand keypoint coordinate of the frame image compared to the hand keypoint coordinate of the previous frame image; or, if the frame image is the first frame of the multi-frame image, then the keypoint motion information corresponding to the frame image is the preset keypoint motion information.

[0112] In some embodiments, the processing module 302 is further configured to perform multiple gesture recognition operations on the video to be detected within a preset time period; the gesture recognition operation is used to obtain gesture recognition results in multiple frames of images in the video to be detected; if the multiple gesture recognition results obtained by performing multiple gesture recognition operations within the preset time period are inconsistent, the gesture recognition result that appears most frequently among the multiple gesture recognition results is determined as the final gesture recognition result.

[0113] It should be noted that, Figure 7 The module division shown is illustrative and represents only one logical functional division; in actual implementation, other division methods are possible. For example, two or more functions can be integrated into a single processing module. These integrated modules can be implemented either in hardware or as software functional modules.

[0114] In implementing the functions of the integrated modules described above using hardware, this application provides another possible structural schematic diagram of the gesture recognition device involved in the above embodiments. For example... Figure 8As shown, the gesture recognition device 400 includes a processor 402 and a bus 404. Optionally, the gesture recognition device may also include a memory 401; alternatively, the gesture recognition device may also include a communication interface 403.

[0115] Processor 402 may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 402 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 402 may also be a combination that implements computing functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0116] Communication interface 403 is used to connect to other devices via a communication network. This communication network can be Ethernet, wireless access network, wireless local area network (WLAN), etc.

[0117] The memory 401 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0118] As one possible implementation, the memory 401 can exist independently of the processor 402. The memory 401 can be connected to the processor 402 via a bus 404 and is used to store instructions or program code. When the processor 402 calls and executes the instructions or program code stored in the memory 401, it can implement the gesture recognition method provided in the embodiments of this application.

[0119] In another possible implementation, the memory 401 can also be integrated with the processor 402.

[0120] Bus 404 can be an extended industry standard architecture (EISA) bus, etc. Bus 404 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0121] Through the above description of the implementation methods, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the gesture recognition device can be divided into different functional modules to complete all or part of the functions described above.

[0122] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be executed by computer instructions instructing related hardware. The program can be stored in the aforementioned computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be any of the foregoing embodiments or memory. The aforementioned computer-readable storage medium can also be an external storage device for the gesture recognition device, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the gesture recognition device. Further, the aforementioned computer-readable storage medium can include both internal storage units of the gesture recognition device and external storage devices. The aforementioned computer-readable storage medium is used to store the aforementioned computer program and other programs and data required by the gesture recognition device. The aforementioned computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0123] This application also provides a computer program product, which includes a computer program that, when run on a computer, causes the computer to execute any of the gesture recognition methods provided in the above embodiments.

[0124] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, the disclosure, and the appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0125] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.

[0126] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A gesture recognition method, characterized in that, The method includes: The key hand feature information of each frame in a multi-frame image is obtained; the multi-frame images are obtained by capturing the moving hand; the key hand feature information of a frame image includes the coordinates of the key hand points in that frame image and the motion information of the key hand points. The motion information of the key hand points in a frame image is used to characterize the change of the coordinates of the key hand points in that frame image relative to the coordinates of the corresponding key hand points in the previous frame image. The hand keypoint feature information, preset spatial sequence number information, and preset temporal sequence number information of each frame in the multi-frame images are input into a gesture recognition model based on a graph convolutional neural network. The gesture recognition model generates a hand keypoint adjacency matrix based on the hand keypoint feature information and the spatial sequence number information, including: For each of the multiple frames, the key hand feature information and the spatial sequence number information of the frame are fused to obtain the fused information of the frame. The fusion information of the frame image is input into the first convolutional layer and the second convolutional layer respectively to obtain the matrix of fusion information of the frame image output by the first convolutional layer and the matrix of fusion information of the frame image output by the second convolutional layer. The matrix of fusion information of the frame image output by the first convolutional layer is transposed and multiplied with the matrix of fusion information of the frame image output by the second convolutional layer to obtain the adjacency matrix of the hand key points of the frame image. The gesture recognition model uses the time series sequence information to mark the frame number and establishes inter-frame dependencies through pooling and double convolution to obtain the gesture recognition result; wherein, the spatial sequence sequence information is used to mark the sequence number of the hand key points in each frame of the multi-frame images, and the time series sequence information is used to mark the frame number corresponding to the spatial information of each frame of the multi-frame images.

2. The method according to claim 1, characterized in that, The gesture recognition model based on graph convolutional neural networks includes a spatial information extraction sub-model, a temporal information extraction sub-model, and a classification sub-model. The model inputs the hand keypoint feature information, preset spatial sequence number information, and preset temporal sequence number information of each frame in the multi-frame images into the gesture recognition model based on graph convolutional neural networks. The gesture recognition model generates a hand keypoint adjacency matrix based on the hand keypoint feature information and the spatial sequence number information, uses the temporal sequence number information to label the frame number, and establishes inter-frame dependencies through pooling and double convolution to obtain the gesture recognition result, including: The hand keypoint feature information and preset spatial sequence number information of each frame in the multi-frame image are input into the spatial information extraction sub-model to obtain the hand keypoint adjacency matrix of each frame image. The hand keypoint adjacency matrix of a frame image is used to represent the connection relationship between any two hand keypoints in that frame image. The graph convolutional neural network of the spatial information extraction sub-model is used to process the hand keypoint adjacency matrix of each frame image to obtain the spatial information of each frame image in the multi-frame image. The spatial information of each frame in the multi-frame image and the preset time series sequence number information are input into the time series information extraction sub-model to obtain the time series information of the multi-frame image; the time series information of the multi-frame image is used to characterize the spatial information change of the multi-frame image from the first frame image to the last frame image; The temporal information of the multi-frame images is input into the classification sub-model to obtain the gesture recognition result.

3. The method according to claim 2, characterized in that, The step of inputting the spatial information of each frame in the multi-frame images and the preset time series sequence number information into the time series information extraction sub-model to obtain the time series information of the multi-frame images includes: Based on the preset time series sequence information, the frame number corresponding to each frame in the multi-frame images is marked; The spatial information of the multi-frame images after the frame number is marked is input into the spatial information pooling layer in the temporal information extraction sub-model to obtain the spatial information of the multi-frame images after pooling. The spatial information of the multi-frame images after pooling is sequentially input into the third and fourth convolutional layers. The dependency relationship between each frame of the multi-frame images is established through the third and fourth convolutional layers to obtain the spatial information of the multi-frame images after the dependency relationship is established. The spatial information of the multi-frame images after establishing the dependency relationship is input into the temporal information pooling layer in the temporal information extraction sub-model to obtain the temporal information of the multi-frame images.

4. The method according to claim 1, characterized in that, The step of obtaining the hand key point feature information of each frame in multiple frames of images includes: For each frame of a multi-frame image with a preset number of frames, the frame image is input into the regression model to obtain the coordinates of the hand key points in that frame image; If the frame image is not the first frame image in the multi-frame image, then for any hand key point coordinate in the frame image, the key point motion information corresponding to the hand key point in the frame image is determined based on the change in the hand key point coordinates in the frame image compared to the hand key point coordinates in the previous frame image; or, if the frame image is the first frame image in the multi-frame image, then the key point motion information corresponding to the frame image is the preset key point motion information.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: Within a preset time period, multiple gesture recognition operations are performed on the video to be detected within the preset time period; the gesture recognition operations are used to obtain the gesture recognition results of multiple frames in the video to be detected. If the multiple gesture recognition results obtained from the multiple gesture recognition operations within the preset time period are inconsistent, the gesture recognition result that appears most frequently among the multiple gesture recognition results shall be determined as the final gesture recognition result.

6. A gesture recognition device, characterized in that, include: The acquisition module is used to acquire the key hand feature information of each frame in a multi-frame image. The multi-frame images are obtained by capturing images of a moving hand; The hand key point feature information of a frame image includes the coordinates of the hand key points in the frame image and the motion information of the hand key points. The motion information of the hand key points in a frame image is used to characterize the change of the coordinates of the hand key points in the frame image relative to the corresponding coordinates of the hand key points in the previous frame image. The processing module is used to input the hand keypoint feature information, preset spatial sequence number information, and preset temporal sequence number information of each frame in the multi-frame images into a gesture recognition model based on a graph convolutional neural network. The gesture recognition model generates a hand keypoint adjacency matrix based on the hand keypoint feature information and the spatial sequence number information, including: For each of the multiple frames, the key hand feature information and the spatial sequence number information of the frame are fused to obtain the fused information of the frame. The fusion information of the frame image is input into the first convolutional layer and the second convolutional layer respectively to obtain the matrix of fusion information of the frame image output by the first convolutional layer and the matrix of fusion information of the frame image output by the second convolutional layer. The matrix of fusion information of the frame image output by the first convolutional layer is transposed and multiplied with the matrix of fusion information of the frame image output by the second convolutional layer to obtain the adjacency matrix of the hand key points of the frame image. The gesture recognition model uses the time series sequence information to mark the frame number and establishes inter-frame dependencies through pooling and double convolution to obtain the gesture recognition result; wherein, the spatial sequence sequence information is used to mark the sequence number of the hand key points in each frame of the multi-frame images, and the time series sequence information is used to mark the frame number corresponding to the spatial information of each frame of the multi-frame images.

7. The apparatus according to claim 6, characterized in that, The gesture recognition model based on graph convolutional neural networks includes a spatial information extraction sub-model, a temporal information extraction sub-model, and a classification sub-model based on graph convolutional neural networks. The processing module is specifically used to input the hand keypoint feature information and the preset spatial sequence number information of each frame of the multi-frame images into the spatial information extraction sub-model to obtain the hand keypoint adjacency matrix of each frame image; the hand keypoint adjacency matrix of each frame image is used to represent the connection relationship between any two hand keypoints in each frame image; the graph convolutional neural network of the spatial information extraction sub-model is used to process the hand keypoint adjacency matrix of each frame image to obtain the spatial information of each frame of the multi-frame images; The spatial information of each frame in the multi-frame images and the preset time series sequence information are input into the time series information extraction sub-model to obtain the time series information of the multi-frame images; The temporal information of the multi-frame images is used to characterize the spatial information changes of the multi-frame images from the first frame image to the last frame image; the temporal information of the multi-frame images is input into the classification sub-model to obtain the gesture recognition result; The processing module is specifically used to mark the frame number corresponding to each frame in the multi-frame images according to the preset time series sequence information; input the spatial information of the multi-frame images after marking the frame number to the spatial information pooling layer in the temporal information extraction sub-model to obtain the spatial information of the multi-frame images after pooling; input the spatial information of the multi-frame images after pooling to the third convolutional layer and the fourth convolutional layer in sequence, and then establish the dependency relationship between each frame in the multi-frame images through the third convolutional layer and the fourth convolutional layer to obtain the spatial information of the multi-frame images after establishing the dependency relationship; input the spatial information of the multi-frame images after establishing the dependency relationship to the temporal information pooling layer in the temporal information extraction sub-model to obtain the temporal information of the multi-frame images; The acquisition module is specifically used to input each frame of a multi-frame image with a preset number of frames into a regression model to obtain the coordinates of the hand key points in that frame; if the frame is not the first frame of the multi-frame image, then for any hand key point coordinate in the frame, the key point motion information corresponding to the hand key point in the frame is determined based on the change in the hand key point coordinates of the frame compared to the hand key point coordinates of the previous frame; or, if the frame is the first frame of the multi-frame image, then the key point motion information corresponding to the frame is the preset key point motion information. The processing module is further configured to perform multiple gesture recognition operations on the video to be detected within a preset time period; the gesture recognition operations are used to obtain gesture recognition results of multiple frames in the video to be detected. If the multiple gesture recognition results obtained from the multiple gesture recognition operations within the preset time period are inconsistent, the gesture recognition result that appears most frequently among the multiple gesture recognition results shall be determined as the final gesture recognition result.

8. A gesture recognition device, characterized in that, include: A memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code including computer instructions; Wherein, when the processor executes the computer instructions, the gesture recognition device causes the gesture recognition device to perform the gesture recognition method as described in any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions; When the computer instructions are executed on the gesture recognition device, the gesture recognition device performs the gesture recognition method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Gesture recognition method and device based on space-time diagram convolutional neural network

    CN112329525A