Gesture recognition method and device, equipment and storage medium

By combining the detection results of the image sensor and the depth sensor, and using historical hand shape information for feature extraction, the problem of inaccurate estimation of the monocular camera scale is solved, and the accuracy of gesture recognition is improved.

CN120236320APending Publication Date: 2025-07-01BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311841103.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, monocular cameras have problems with inaccurate scale estimation when used for gesture recognition, resulting in low accuracy of recognition results.

Method used

During the recognition process, the detection results of the image sensor and the detection results of the depth sensor are combined, and feature extraction is performed using historical hand shape information, and a deep learning model is input to improve the accuracy of gesture recognition.

Benefits of technology

Through the combination of three dimensions: hand shape information, image information, and depth information, the accuracy of gesture recognition method is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236320A_ABST
    Figure CN120236320A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a gesture recognition method and device, equipment and a storage medium, and the method comprises the steps: obtaining historical hand shape information which corresponds to a to-be-recognized hand object and is recognized last time, and obtaining a first detection result of an image sensor and a second detection result of a depth sensor; performing feature extraction on the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information, selecting at least one target detection result from the first detection result and the second detection result, and performing feature extraction on each target detection result to obtain a second feature vector corresponding to the historical hand shape information; obtaining at least one second feature vector corresponding to the at least one target detection result; and inputting the first feature vector and the at least one second feature vector into a preset deep learning model, and outputting a gesture recognition result of the current recognition corresponding to the hand object. Through the hand shape information, the accuracy of the gesture recognition result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of image processing technology, and in particular, to a gesture recognition method, apparatus, device, and storage medium. Background Art

[0002] With the development of image processing technology, technicians have developed more and more image processing scenarios. For example, virtual reality scenarios, augmented reality scenarios, etc. In many interactive scenarios, it is necessary to recognize the gestures of a hand object to obtain the operation instructions corresponding to the gesture.

[0003] In the prior art, the gesture recognition method includes: first, collecting a hand image through a gray fisheye camera, and then inputting the hand image into a trained neural network model to obtain the gesture recognition result corresponding to the hand image.

[0004] However, the inventor found that the prior art has at least the following technical problems: the gray fisheye camera is a monocular camera, and the monocular camera has the problem of inaccurate scale estimation, so the accuracy of the recognition result obtained by the above method is low. Summary of the Invention

[0005] Embodiments of the present disclosure provide a gesture recognition method, apparatus, device, and storage medium to overcome the problem of low accuracy in gesture recognition of hand images.

[0006] In a first aspect, embodiments of the present disclosure provide a gesture recognition method applied to an electronic device, where an image sensor and a depth sensor are installed in the electronic device, and the method includes:

[0007] Obtaining the historical hand shape information of the last recognition corresponding to the hand object to be recognized, and obtaining the first detection result of the image sensor and the second detection result of the depth sensor;

[0008] Performing feature extraction on the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information, and selecting at least one target detection result from the first detection result and the second detection result, and performing feature extraction on each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result;

[0009] Inputting the first feature vector and the at least one second feature vector into a preset deep learning model, and outputting the gesture recognition result of the current recognition corresponding to the hand object.

[0010] In a second aspect, embodiments of the present disclosure provide a gesture recognition apparatus applied to an electronic device, where an image sensor and a depth sensor are installed in the electronic device, and the apparatus includes:

[0011] An acquisition module, configured to acquire historical hand shape information of a hand object to be recognized in a previous recognition, and acquire a first detection result of the image sensor and a second detection result of the depth sensor;

[0012] A feature extraction module, configured to extract features from the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information, and select at least one target detection result from the first detection result and the second detection result, and extract features from each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result;

[0013] An identification module, configured to input the first feature vector and the at least one second feature vector into a preset deep learning model, and output a gesture recognition result of the current recognition corresponding to the hand object.

[0014] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:

[0015] A processor and a memory communicatively connected to the processor;

[0016] The memory stores computer-executable instructions;

[0017] The processor executes the computer-executable instructions stored in the memory to implement the gesture recognition method as described in the first aspect and various possible designs of the first aspect above.

[0018] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the gesture recognition method as described in the first aspect and various possible designs of the first aspect above is implemented.

[0019] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the gesture recognition method as described in the first aspect and various possible designs of the first aspect above is implemented.

[0020] The gesture recognition method, device, equipment and storage medium provided in this embodiment include: obtaining the historical hand shape information of the hand object to be recognized in the previous recognition, and obtaining the first detection result of the image sensor and the second detection result of the depth sensor; extracting features from the historical hand shape information to obtain the first feature vector corresponding to the historical hand shape information, and selecting at least one target detection result from the first detection result and the second detection result, and extracting features from each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result; inputting the first feature vector and the at least one second feature vector into a preset deep learning model, and outputting the gesture recognition result of the current recognition of the hand object. In the embodiments of the present disclosure, since during the recognition process, on the basis of the first detection result of the image sensor and the second detection result of the depth sensor, the hand shape information is also added, that is, the gesture of the hand object is recognized through three dimensions: hand shape information, image information, and depth information, so the accuracy of the gesture recognition method for recognizing the hand object is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic diagram of the application scenario of a gesture recognition method provided in an embodiment of the present disclosure;

[0023] Figure 2 It is a flowchart of a gesture recognition method provided in an embodiment of the present disclosure;

[0024] Figure 3 It is a schematic diagram of a gesture recognition method provided in an embodiment of the present disclosure;

[0025] Figure 4 It is a schematic diagram of another gesture recognition method provided in an embodiment of the present disclosure;

[0026] Figure 5 It is a flowchart of another gesture recognition method provided in an embodiment of the present disclosure;

[0027] Figure 6 It is a structural block diagram of a gesture recognition device provided in an embodiment of the present disclosure;

[0028] Figure 7 It is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. Detailed implementation manners

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0030] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to select authorization or rejection.

[0031] With the development of image processing technology, technicians have developed more and more image processing scenarios. For example, virtual reality scenarios, augmented reality scenarios, etc. In many interactive scenarios, it is necessary to recognize the gestures of a hand object and obtain the operation instructions corresponding to the gestures.

[0032] In the prior art, the gesture recognition method includes: first, collecting a hand image through a gray fisheye camera, and then inputting the hand image into a trained neural network model to obtain the gesture recognition result corresponding to the hand image. However, the gray fisheye camera is a monocular camera, and the monocular camera has the problem of inaccurate scale estimation, so the accuracy of the recognition result obtained by the above method is relatively low.

[0033] Currently, to solve the problem of inaccurate scale estimation of monocular cameras, depth data can be collected through a depth sensor, and a trained neural network model can be used to recognize the gestures of a hand object. However, because the depth camera itself can filter the background at a long distance through depth information and only retain the hand information at a short distance. Therefore, the recognition robustness is higher, but the accuracy of the depth camera itself is generally low (the error is generally close to 1 cm), and the information discrimination of only the depth map is insufficient, so the accuracy of the recognition result is also relatively low.

[0034] Therefore, it can be seen that how to improve the accuracy of the gesture recognition method for gesture recognition of a hand object is a technical problem that needs to be solved urgently at present.

[0035] To solve the above problems, the following technical concept is provided in this embodiment: obtaining the hand shape information of a hand object and inputting the hand shape information into a gesture recognition model to improve the gesture recognition accuracy of the model.

[0036] Correspondingly, the specific steps may include: First, obtain the historical hand shape information of the last recognition corresponding to the hand object to be recognized, and obtain the first detection result of the image sensor and the second detection result of the depth sensor; then, extract features from the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information, and select at least one target detection result from the first detection result and the second detection result, and extract features from each target detection result to obtain at least one second feature vector corresponding to at least one target detection result; finally, input the first feature vector and at least one second feature vector into a preset deep learning model, and output the gesture recognition result of the current recognition corresponding to the hand object.

[0037] In this case, since in the recognition process, based on the first detection result of the image sensor and the second detection result of the depth sensor, the hand shape information is also added, that is, the gesture of the hand object is recognized through these three dimensions of hand shape information, image information, and depth information, so the accuracy of the gesture recognition method for recognizing the hand object is improved.

[0038] The application scenarios of the embodiments of the present disclosure are explained below: The gesture recognition method provided by the embodiments of the present disclosure can be applied to multiple scenarios. For example, it can be applied to an AR (Augmented Reality) game scenario. Figure 1 It is a diagram of an application scenario of the gesture recognition method provided by the embodiments of the present disclosure. Specifically, as Figure 1 shown, Figure 1 it includes a hand object, and the gesture of the hand object can be recognized by using the gesture recognition method provided by the embodiments of the present disclosure. Then, determine the operation instruction corresponding to the gesture (for example, pick up a medical kit, open a medical kit, etc.). Finally, perform the corresponding operation on the medical kit according to the operation instruction corresponding to the gesture.

[0039] The gesture recognition method provided by the embodiments of the present disclosure is described in detail below with detailed embodiments.

[0040] Figure 2 It is a flowchart of the gesture recognition method provided by the embodiments of the present disclosure Figure 1 . The method of this embodiment can be applied to an electronic device, and an image sensor and a depth sensor are installed in the electronic device. Referring to Figure 2 , the method includes:

[0041] S201. Obtain the historical hand shape information of the last recognition corresponding to the hand object to be recognized, and obtain the first detection result of the image sensor and the second detection result of the depth sensor.

[0042] In this embodiment, the hand shape information includes the shape and size of the hand object. For example, information such as the length and thickness of each finger. Among them, the image sensor can be any model of camera module. For example, it can be a grayscale camera, a grayscale fisheye camera, a color camera, etc. Among them, the depth sensor can be any model of depth camera. For example, a TOF (Time of flight) camera, etc.

[0043] Optionally, if the image sensor is a grayscale camera, the first detection result of the image sensor includes a grayscale image. Optionally, if the depth sensor is a TOF camera, the second detection result of the depth sensor includes a depth image.

[0044] S202. Extract features from the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information, and select at least one target detection result from the first detection result and the second detection result, and extract features from each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result.

[0045] In this step, if the angle of the first field of view corresponding to the image sensor is greater than the angle of the second field of view corresponding to the depth sensor, it is possible that both the image sensor and the depth sensor can detect the hand object, or it is possible that the image sensor can detect the hand object, but the depth sensor cannot detect the hand object. The following describes the two cases separately.

[0046] In the first case, the image sensor can detect the hand object, and the depth sensor can also detect the hand object. Correspondingly, selecting at least one target detection result from the first detection result and the second detection result, and extracting features from each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result includes: if the first detection result includes the hand object and the second detection result includes the hand object, then determining the first detection result and the second detection result as the target detection results; extracting features from the first detection result to obtain a second feature vector corresponding to the first detection result, and extracting features from the second detection result to obtain a second feature vector corresponding to the second detection result.

[0047] Optionally, the first detection result may include a grayscale image corresponding to the hand object, and the second detection result includes a depth image corresponding to the hand object.

[0048] In some embodiments, extracting features from the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information includes: extracting features from the historical hand shape information through a trained first feature extraction model to obtain a first feature vector corresponding to the historical hand shape information.

[0049] Exemplarily, as Figure 3 shown, the first feature extraction model can be: an encoder model.

[0050] In some embodiments, performing feature extraction on the first detection result to obtain a second feature vector corresponding to the first detection result, and performing feature extraction on the second detection result to obtain a second feature vector corresponding to the second detection result, includes: performing feature extraction on the first detection result through a trained second feature extraction model to obtain a second feature vector corresponding to the first detection result, and performing feature extraction on the second detection result through a trained third feature extraction model to obtain a second feature vector corresponding to the second detection result.

[0051] Exemplarily, as Figure 3 shown, the second feature extraction model can be: a backbone - gray (bone and joint grayscale) model. The third feature extraction model can be: a backbone - depth (bone and joint depth) model.

[0052] In the second case, as Figure 4 shown, the image sensor can detect a hand object, but the depth sensor cannot detect the hand object. Correspondingly, selecting at least one target detection result from the first detection result and the second detection result, and performing feature extraction on each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result, includes: if the first detection result includes a hand object and the second detection result does not include a hand object, then determining the first detection result as the target detection result; performing feature extraction on the first detection result to obtain a second feature vector corresponding to the first detection result.

[0053] Optionally, the first detection result includes a grayscale image corresponding to the hand object.

[0054] S203. Input the first feature vector and at least one second feature vector into a preset deep learning model, and output the gesture recognition result of the current recognition corresponding to the hand object.

[0055] In this step, the preset deep learning model can be a trained deep learning model. During the training process, training samples and validation samples can be input to train the initial deep learning model to obtain a preset deep learning model with an accuracy greater than a preset value. Exemplarily, as Figure 3 shown, the preset deep learning model can be a Fuse (multimodal) model.

[0056] Among them, the input information of the preset deep learning model includes: the first feature information corresponding to the hand shape information, the second feature information corresponding to the grayscale image, and the third feature information corresponding to the depth image. The output information of the preset deep learning model includes: the gesture recognition result corresponding to the hand object. Among them, the gesture recognition result may include hand pose information and hand shape information.

[0057] Optionally, the hand pose information may include the coordinates of multiple key points and the rotation angle of each key point. Among them, the key points may be joints in the hand object. In some embodiments, the hand pose information may further include the hand pose. For example, a fist pose, a grasping pose, etc. Among them, the coordinates of the key points corresponding to each pose and the rotation angle of the key points are different.

[0058] It should be noted that, continuing to refer to Figure 4 , in the case where the image sensor can detect the hand object, but the depth sensor cannot detect the hand object, only the grayscale image and the optimized hand shape information (shape parameter) of the previous frame are used as the input of the Fuse (multi-modal) model to predict the gesture recognition result. And, because the pure grayscale image lacks scale information, the hand pose information in the gesture recognition result can be discarded, and the optimized hand shape information of the previous frame is used subsequently. Correspondingly, the method further includes: discarding the hand shape information recognized this time and using the historical hand shape information as the hand shape information recognized this time.

[0059] The gesture recognition method provided in this embodiment includes: obtaining the historical hand shape information recognized last time corresponding to the hand object to be recognized, and obtaining the first detection result of the image sensor and the second detection result of the depth sensor; extracting features from the historical hand shape information to obtain the first feature vector corresponding to the historical hand shape information, and selecting at least one target detection result from the first detection result and the second detection result, and extracting features from each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result; inputting the first feature vector and the at least one second feature vector into a preset deep learning model to output the gesture recognition result of the hand object recognized this time. In the embodiments of the present disclosure, since during the recognition process, based on the first detection result of the image sensor and the second detection result of the depth sensor, the hand shape information is also added, that is, through the three dimensions of hand shape information, image information, and depth information, the gesture of the hand object is recognized, so the accuracy of the gesture recognition method for recognizing the hand object is improved.

[0060] It should be noted that the gesture recognition result includes hand pose information and hand shape information. The second detection result includes the depth image corresponding to the hand object. The gesture recognition provided in this embodiment can further refine the obtained hand shape information according to the three-dimensional point cloud data corresponding to the depth image to improve the accuracy of the hand shape information. Correspondingly, as Figure 5 shown, the method further includes:

[0061] S501. Convert the depth image into three-dimensional point cloud data, where the three-dimensional point cloud data includes the first coordinate information of multiple first vertices corresponding to the hand object.

[0062] Exemplarily, the first coordinate information corresponding to the three-dimensional point cloud data can be represented as Pcd. Among them, the first coordinate information can be three-dimensional coordinate points.

[0063] S502. Input the hand pose information and the hand shape information into a preset hand model, and output a mesh graph corresponding to the hand object, where the mesh graph includes the second coordinate information of multiple second vertices corresponding to the hand object.

[0064] Exemplarily, the preset hand model can be the MANO hand model. Among them, the second coordinate information can be three-dimensional coordinate points.

[0065] S503. Adjust the hand shape information according to the first coordinate information of multiple first vertices and the second coordinate information of multiple second vertices to obtain new hand shape information, and update the hand shape information in the gesture recognition result to the new hand shape information.

[0066] Optionally, the number of multiple first vertices is greater than the number of multiple second vertices; correspondingly, this step is: for each second vertex, select the target vertex closest to the second vertex from multiple first vertices; determine the distance between the second vertex and the target vertex according to the second coordinate information of the second vertex and the first coordinate information of the target vertex; adjust the hand shape information until the distance between the second vertex and the target vertex is the smallest, and determine the target hand shape information corresponding to the smallest distance as the new hand shape information.

[0067] Exemplarily, the minimum distance between the second vertex and the target vertex can be determined according to a preset fitting equation. The fitting equation is:

[0068] β=argmin θ,β ||FK(θ,β)-Pcd||

[0069] where θ represents the hand pose information, β represents the hand shape information, FK(θ,β) represents the hand model, and Pcd represents the three-dimensional point cloud data.

[0070] In the embodiments of the present disclosure, the obtained hand shape information can be further refined according to the three-dimensional point cloud data corresponding to the depth image, so as to improve the accuracy of the hand shape information, and further improve the accuracy of recognizing the gesture of the hand object through the three dimensions of hand shape information, image information, and depth information.

[0071] Figure 6 FIG. 4 is a structural block diagram of a gesture recognition device provided by an embodiment of the present disclosure, which is applied to an electronic device. An image sensor and a depth sensor are installed in the electronic device. The device includes: an acquisition module 601, a feature extraction module 602, and an identification module 603.

[0072] The acquisition module 601 is configured to acquire the historical hand shape information of the hand object to be recognized in the previous recognition, and acquire the first detection result of the image sensor and the second detection result of the depth sensor.

[0073] The feature extraction module 602 is configured to extract features from the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information, and select at least one target detection result from the first detection result and the second detection result, and extract features from each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result.

[0074] The identification module 603 is configured to input the first feature vector and the at least one second feature vector into a preset deep learning model, and output the gesture recognition result of the current recognition corresponding to the hand object.

[0075] According to one or more embodiments of the present disclosure, the feature extraction module 602 selects at least one target detection result from the first detection result and the second detection result, and extracts features from each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result, which specifically includes: if the first detection result includes the hand object and the second detection result includes the hand object, determining the first detection result and the second detection result as target detection results; extracting features from the first detection result to obtain a second feature vector corresponding to the first detection result, and extracting features from the second detection result to obtain a second feature vector corresponding to the second detection result.

[0076] According to one or more embodiments of the present disclosure, the angle of the first field of view corresponding to the image sensor is greater than the angle of the second field of view corresponding to the depth sensor; correspondingly, the feature extraction module 602 selects at least one target detection result from the first detection result and the second detection result, and performs feature extraction on each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result. Specifically, if the first detection result includes the hand object and the second detection result does not include the hand object, the first detection result is determined as the target detection result; feature extraction is performed on the first detection result to obtain a second feature vector corresponding to the first detection result.

[0077] According to one or more embodiments of the present disclosure, the first detection result includes a grayscale image corresponding to the hand object, and the second detection result includes a depth image corresponding to the hand object.

[0078] According to one or more embodiments of the present disclosure, the gesture recognition result includes hand pose information and hand shape information, and the second detection result includes a depth image corresponding to the hand object; the device further includes: an adjustment module; the adjustment module is configured to convert the depth image into three-dimensional point cloud data, where the three-dimensional point cloud data includes first coordinate information of a plurality of first vertices corresponding to the hand object; input the hand pose information and the hand shape information into a preset hand model, and output a mesh graph corresponding to the hand object, where the mesh graph includes second coordinate information of a plurality of second vertices corresponding to the hand object; adjust the hand shape information according to the first coordinate information of the plurality of first vertices and the second coordinate information of the plurality of second vertices to obtain new hand shape information, and update the hand shape information in the gesture recognition result to the new hand shape information.

[0079] According to one or more embodiments of the present disclosure, the number of the plurality of first vertices is greater than the number of the plurality of second vertices; correspondingly, the adjustment module adjusts the hand shape information according to the first coordinate information of the plurality of first vertices and the second coordinate information of the plurality of second vertices to obtain new hand shape information. Specifically, for each second vertex, a target vertex closest to the second vertex is selected from the plurality of first vertices; the distance between the second vertex and the target vertex is determined according to the second coordinate information of the second vertex and the first coordinate information of the target vertex; the hand shape information is adjusted until the distance between the second vertex and the target vertex is the smallest, and the target hand shape information corresponding to the smallest distance is determined as the new hand shape information.

[0080] According to one or more embodiments of the present disclosure, the gesture recognition result includes hand pose information and hand shape information, the first detection result includes the hand object and the second detection result does not include the hand object; the apparatus further includes: a discarding module; the discarding module is configured to discard the hand shape information recognized this time and use the historical hand shape information as the hand shape information recognized this time.

[0081] Among them, the acquisition module 601, the feature extraction module 602, and the recognition module 603 are connected in sequence. The gesture recognition apparatus provided in this embodiment can execute the technical solutions of the above method embodiments, and the implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.

[0082] Figure 7 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present disclosure. Refer to Figure 7 , the electronic device 700 can be a terminal device or a server. Among them, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (Personal Digital Assistant, abbreviated as PDA), tablet computers (Portable Android Device, abbreviated as PAD), portable multimedia players (Portable Media Player, abbreviated as PMP), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0083] As Figure 7 shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (Read Only Memory, abbreviated as ROM) 702 or the program loaded from the storage device 708 into the random access memory (Random Access Memory, abbreviated as RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0084] Typically, the following devices can be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and a communication device 709. The communication device 709 can allow the electronic device 700 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 the electronic device 700 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0085] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above functions defined in the methods of the embodiments of the present disclosure are performed.

[0086] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0087] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately without being assembled into the electronic device.

[0088] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above-mentioned embodiments.

[0089] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0091] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0092] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0093] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0094] In a first aspect, according to one or more embodiments of the present disclosure, there is provided a gesture recognition method applied to an electronic device, in which an image sensor and a depth sensor are installed, and the method includes:

[0095] Obtaining historical hand shape information of the last recognition corresponding to a hand object to be recognized, and obtaining a first detection result of the image sensor and a second detection result of the depth sensor;

[0096] Performing feature extraction on the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information, and selecting at least one target detection result from the first detection result and the second detection result, and performing feature extraction on each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result;

[0097] Inputting the first feature vector and the at least one second feature vector into a preset deep learning model, and outputting a gesture recognition result of the current recognition corresponding to the hand object.

[0098] According to one or more embodiments of the present disclosure, selecting at least one target detection result from the first detection result and the second detection result, and performing feature extraction on each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result includes: if the first detection result includes the hand object and the second detection result includes the hand object, determining the first detection result and the second detection result as target detection results; performing feature extraction on the first detection result to obtain a second feature vector corresponding to the first detection result, and performing feature extraction on the second detection result to obtain a second feature vector corresponding to the second detection result.

[0099] According to one or more embodiments of the present disclosure, the angle of the first field of view corresponding to the image sensor is greater than the angle of the second field of view corresponding to the depth sensor; correspondingly, selecting at least one target detection result from the first detection result and the second detection result, and performing feature extraction on each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result, including: if the first detection result includes the hand object and the second detection result does not include the hand object, determining the first detection result as the target detection result; performing feature extraction on the first detection result to obtain a second feature vector corresponding to the first detection result.

[0100] According to one or more embodiments of the present disclosure, the first detection result includes a grayscale image corresponding to the hand object, and the second detection result includes a depth image corresponding to the hand object.

[0101] According to one or more embodiments of the present disclosure, the gesture recognition result includes hand pose information and hand shape information, and the second detection result includes a depth image corresponding to the hand object; the method further includes: converting the depth image into three-dimensional point cloud data, where the three-dimensional point cloud data includes first coordinate information of a plurality of first vertices corresponding to the hand object; inputting the hand pose information and the hand shape information into a preset hand model, and outputting a mesh graph corresponding to the hand object, where the mesh graph includes second coordinate information of a plurality of second vertices corresponding to the hand object; adjusting the hand shape information according to the first coordinate information of the plurality of first vertices and the second coordinate information of the plurality of second vertices to obtain new hand shape information, and updating the hand shape information in the gesture recognition result to the new hand shape information.

[0102] According to one or more embodiments of the present disclosure, the number of the plurality of first vertices is greater than the number of the plurality of second vertices; correspondingly, the adjusting the hand shape information according to the first coordinate information of the plurality of first vertices and the second coordinate information of the plurality of second vertices to obtain new hand shape information includes: for each second vertex, selecting a target vertex closest to the second vertex from the plurality of first vertices; determining a distance between the second vertex and the target vertex according to the second coordinate information of the second vertex and the first coordinate information of the target vertex; adjusting the hand shape information until the distance between the second vertex and the target vertex is the smallest, and determining the target hand shape information corresponding to the smallest distance as the new hand shape information.

[0103] According to one or more embodiments of the present disclosure, the gesture recognition result includes hand pose information and hand shape information, the first detection result includes the hand object and the second detection result does not include the hand object; the method further includes: discarding the hand shape information recognized this time, and using the historical hand shape information as the hand shape information recognized this time.

[0104] In a second aspect, according to one or more embodiments of the present disclosure, there is provided a gesture recognition device applied to an electronic device, where an image sensor and a depth sensor are installed in the electronic device, and the device includes:

[0105] An acquisition module, configured to acquire the historical hand shape information recognized last time corresponding to the hand object to be recognized, and acquire the first detection result of the image sensor and the second detection result of the depth sensor;

[0106] A feature extraction module, configured to extract features from the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information, and select at least one target detection result from the first detection result and the second detection result, and extract features from each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result;

[0107] A recognition module, configured to input the first feature vector and the at least one second feature vector into a preset deep learning model, and output the gesture recognition result of the hand object recognized this time.

[0108] According to one or more embodiments of the present disclosure, the feature extraction module selects at least one target detection result from the first detection result and the second detection result, and extracts features from each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result, which specifically includes: if the first detection result includes the hand object and the second detection result includes the hand object, determining the first detection result and the second detection result as target detection results; extracting features from the first detection result to obtain a second feature vector corresponding to the first detection result, and extracting features from the second detection result to obtain a second feature vector corresponding to the second detection result.

[0109] According to one or more embodiments of the present disclosure, the angle of the first field of view corresponding to the image sensor is greater than the angle of the second field of view corresponding to the depth sensor; correspondingly, the feature extraction module selects at least one target detection result from the first detection result and the second detection result, and performs feature extraction on each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result, which specifically includes: if the first detection result includes the hand object and the second detection result does not include the hand object, then determining the first detection result as the target detection result; performing feature extraction on the first detection result to obtain a second feature vector corresponding to the first detection result.

[0110] According to one or more embodiments of the present disclosure, the first detection result includes a grayscale image corresponding to the hand object, and the second detection result includes a depth image corresponding to the hand object.

[0111] According to one or more embodiments of the present disclosure, the gesture recognition result includes hand pose information and hand shape information, and the second detection result includes a depth image corresponding to the hand object; the apparatus further includes: an adjustment module; the adjustment module is configured to convert the depth image into three-dimensional point cloud data, where the three-dimensional point cloud data includes first coordinate information of a plurality of first vertices corresponding to the hand object; input the hand pose information and the hand shape information into a preset hand model, and output a mesh graph corresponding to the hand object, where the mesh graph includes second coordinate information of a plurality of second vertices corresponding to the hand object; adjust the hand shape information according to the first coordinate information of the plurality of first vertices and the second coordinate information of the plurality of second vertices to obtain new hand shape information, and update the hand shape information in the gesture recognition result to the new hand shape information.

[0112] According to one or more embodiments of the present disclosure, the number of the plurality of first vertices is greater than the number of the plurality of second vertices; correspondingly, the adjustment module adjusts the hand shape information according to the first coordinate information of the plurality of first vertices and the second coordinate information of the plurality of second vertices to obtain new hand shape information, which specifically includes: for each second vertex, selecting a target vertex closest to the second vertex from the plurality of first vertices; determining the distance between the second vertex and the target vertex according to the second coordinate information of the second vertex and the first coordinate information of the target vertex; adjusting the hand shape information until the distance between the second vertex and the target vertex is the smallest, and determining the target hand shape information corresponding to the smallest distance as the new hand shape information.

[0113] According to one or more embodiments of the present disclosure, the gesture recognition result includes hand pose information and hand shape information, the first detection result includes the hand object and the second detection result does not include the hand object; the apparatus further includes: a discarding module; the discarding module is configured to discard the hand shape information recognized this time and use the historical hand shape information as the hand shape information recognized this time.

[0114] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, including: a processor, and a memory communicatively connected to the processor;

[0115] The memory stores computer-executable instructions;

[0116] The processor executes the computer-executable instructions stored in the memory to implement the gesture recognition method as described in the first aspect above and various possible designs of the first aspect.

[0117] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the gesture recognition method as described in the first aspect above and various possible designs of the first aspect is implemented.

[0118] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the gesture recognition method as described in the first aspect above and various possible designs of the first aspect is implemented.

[0119] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0120] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable subcombination in multiple embodiments.

[0121] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A gesture recognition method, characterized in that, Applied to an electronic device, an image sensor and a depth sensor are installed in the electronic device, and the method includes: Obtaining the historical hand shape information of the last recognition corresponding to the hand object to be recognized, and obtaining a first detection result of the image sensor and a second detection result of the depth sensor; Performing feature extraction on the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information, and selecting at least one target detection result from the first detection result and the second detection result, and performing feature extraction on each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result; Inputting the first feature vector and the at least one second feature vector into a preset deep learning model, and outputting a gesture recognition result of the current recognition corresponding to the hand object.

2. The method according to claim 1, wherein Selecting at least one target detection result from the first detection result and the second detection result, and performing feature extraction on each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result, including: If the first detection result includes the hand object and the second detection result includes the hand object, determining the first detection result and the second detection result as the target detection results; Performing feature extraction on the first detection result to obtain a second feature vector corresponding to the first detection result, and performing feature extraction on the second detection result to obtain a second feature vector corresponding to the second detection result.

3. The method according to claim 1, characterized in that An angle of a first field of view angle corresponding to the image sensor is greater than an angle of a second field of view angle corresponding to the depth sensor; Correspondingly, selecting at least one target detection result from the first detection result and the second detection result, and performing feature extraction on each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result, including: If the first detection result includes the hand object and the second detection result does not include the hand object, determining the first detection result as the target detection result; Performing feature extraction on the first detection result to obtain a second feature vector corresponding to the first detection result.

4. The method according to claim 2 or 3, characterized in that, The first detection result includes a grayscale image corresponding to the hand object, and the second detection result includes a depth image corresponding to the hand object.

5. The method according to claim 1, characterized in that The gesture recognition result includes hand pose information and hand shape information, and the second detection result includes a depth image corresponding to the hand object; the method further includes: Converting the depth image into three-dimensional point cloud data, where the three-dimensional point cloud data includes first coordinate information of a plurality of first vertices corresponding to the hand object; Inputting the hand pose information and the hand shape information into a preset hand model, and outputting a mesh graph corresponding to the hand object, where the mesh graph includes second coordinate information of a plurality of second vertices corresponding to the hand object; Adjust the hand shape information according to the first coordinate information of the multiple first vertices and the second coordinate information of the multiple second vertices to obtain new hand shape information, and update the hand shape information in the gesture recognition result to the new hand shape information.

6. The method according to claim 5, characterized in that, The number of the multiple first vertices is greater than the number of the multiple second vertices; Accordingly, the adjusting the hand shape information according to the first coordinate information of the multiple first vertices and the second coordinate information of the multiple second vertices to obtain new hand shape information includes: For each second vertex, select a target vertex closest to the second vertex from the multiple first vertices; Determine the distance between the second vertex and the target vertex according to the second coordinate information of the second vertex and the first coordinate information of the target vertex; Adjust the hand shape information until the distance between the second vertex and the target vertex is minimized, and determine the target hand shape information corresponding to the minimum distance as the new hand shape information.

7. The method according to claim 1, characterized in that, The gesture recognition result includes hand pose information and hand shape information, the first detection result includes the hand object and the second detection result does not include the hand object; the method further includes: Discard the hand shape information recognized this time, and use the historical hand shape information as the hand shape information recognized this time.

8. A gesture recognition device, characterized in that, Applied to an electronic device, an image sensor and a depth sensor are installed in the electronic device, and the device includes: An acquisition module, configured to acquire the historical hand shape information recognized last time corresponding to the hand object to be recognized, and acquire a first detection result of the image sensor and a second detection result of the depth sensor; A feature extraction module, configured to extract features from the historical hand shape information to obtain a first feature vector corresponding to the historical hand shape information, and select at least one target detection result from the first detection result and the second detection result, and extract features from each target detection result to obtain at least one second feature vector corresponding to the at least one target detection result; An identification module, configured to input the first feature vector and the at least one second feature vector into a preset deep learning model, and output a gesture recognition result recognized this time corresponding to the hand object.

9. An electronic device, characterized in that, Includes: A processor and a memory communicatively connected to the processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the gesture recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the gesture recognition method according to any one of claims 1 to 7 is implemented.

11. A computer program product, characterized in that, Includes a computer program, and when the computer program is executed by the processor, the gesture recognition method according to any one of claims 1 to 7 is implemented.