Gesture recognition method, system and equipment and storage medium
By using gloves to recognize gestures in VR devices and using deep learning models and light source feature tables to recognize gestures, the problem of low accuracy of camera recognition for naked hands in VR devices is solved, and higher recognition accuracy and interactive comfort are achieved.
Patent Information
- Application Number
- CN202311635252.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
The cameras in existing VR devices have low accuracy in user naked hand gesture recognition, resulting in limited interaction accuracy.
The glove is used as a gesture recognition device. By obtaining the light source information of the light source point on the glove, using a preset deep learning model and light source feature table, the light source position and pixel coordinates are determined, thereby identifying the target gesture.
Improves the accuracy of gesture recognition, provides a non-naked hand recognition method, reduces costs, while enhancing the accuracy and comfort of the interaction.
Smart Images

Figure CN120066245A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gesture recognition, and in particular to a gesture recognition method, system, device and storage medium. Background Art
[0002] In recent years, with the development of VR (Virtual Reality) technology, people's requirements for the interaction of head-mounted VR devices have become higher and higher, and higher requirements have been put forward for battery life, high resolution, low latency, etc. Gesture recognition, as a natural and efficient interaction method for head-mounted VR devices, also poses higher requirements for gesture recognition.
[0003] The traditional gesture recognition method is to directly recognize the bare hand gestures of users through the built-in camera in the VR device, and then realize information interaction. This gesture recognition method has great defects, and there will be a problem that the accuracy of the built-in camera in the VR device for recognizing the bare hands of users is not high. Therefore, there is an urgent need for a gesture recognition method to improve the accuracy of gesture recognition. Summary of the Invention
[0004] The main purpose of the present invention is to propose a gesture recognition method, system, device and storage medium, aiming to improve the accuracy of gesture recognition.
[0005] To achieve the above object, the present invention provides a gesture recognition method, which is applied to a gesture recognition system. The gesture recognition system includes a glove. The gesture recognition method includes the following steps:
[0006] Obtain the light source information of the light source points that can be collected on the glove; wherein, the light source information includes the light source pixel coordinates and the light source features;
[0007] Determine the light source position on the glove according to the light source features and a preset light source feature table;
[0008] Determine the target gesture in a preset deep learning model according to the light source position and the light source pixel coordinates.
[0009] Optionally, when the light source points that can be collected are all the light source points on the glove, the step of determining the target gesture in a preset deep learning model according to the light source position and the light source pixel coordinates includes:
[0010] Based on the light source position, determine the target light source pixel coordinates corresponding to each light source point in the light source pixel coordinates; wherein, the target light source pixel coordinates include the light source point and the light source pixel coordinates corresponding to the light source point;
[0011] Determine the target light source pixel coordinates in a preset deep learning model as the input gesture output information, and use the gesture output information as the target gesture; wherein, the preset deep learning model includes a pre-trained gesture recognition model.
[0012] Optionally, when the collectable light source points are partial light source points on the glove, the step of determining the target gesture in the preset deep learning model according to the light source position and the light source pixel coordinates further includes:
[0013] Based on the light source position, determine the target light source pixel coordinates corresponding to each of the light source points in the light source pixel coordinates; wherein, the target light source pixel coordinates include the light source point and the light source pixel coordinates corresponding to the light source point.
[0014] Based on the collectable light source points and all the light source points on the glove, determine the uncollected light source points, and determine the filled light source pixel coordinates of the uncollected light source points; wherein, the filled light source pixel coordinates include filling the light source pixel coordinates of the uncollected light source points according to a preset rule.
[0015] Determine the target light source pixel coordinates and the filled light source pixel coordinates in a preset deep learning model as the input gesture output information, and use the gesture output information as the target gesture; wherein, the preset deep learning model includes a pre-trained gesture recognition model.
[0016] Optionally, the light source feature includes the light source frequency and the light source brightness. The step of determining the light source position on the glove according to the light source feature and a preset light source feature table includes:
[0017] Determine the target light source frequency matching the light source frequency in the preset light source feature table, and determine the installation position corresponding to the glove of the target light source frequency, and use the installation position as the light source position.
[0018] Or,
[0019] Determine the target light source brightness matching the light source brightness in the preset light source feature table, and determine the installation position corresponding to the glove of the target light source brightness, and use the installation position as the light source position.
[0020] Optionally, the glove is provided with light source points with different features at a plurality of different setting positions. The gesture recognition method further includes:
[0021] Determine the light source features of each of the light source points, and determine the setting positions of each of the light source points on the glove.
[0022] Construct a preset light source feature table of each of the light source points on the glove based on the setting positions and the light source features.
[0023] Optionally, the setting positions at least include the fingertip positions and joint positions of the glove. The step of constructing a preset light source feature table for each light source point on the glove based on the setting positions and the light source features includes:
[0024] Constructing a first preset light source feature table for the light source points at the fingertip positions on the glove based on the fingertip positions and the light source features at the fingertip positions;
[0025] Constructing a second preset light source feature table for the light source points at the joint positions on the glove based on the joint positions and the light source features at the joint positions, and summarizing the first preset light source feature table and the second preset light source feature table to obtain a preset light source feature table.
[0026] Optionally, the gesture recognition system further includes an event camera. The step of obtaining the light source information of the light source points that can be collected on the glove includes:
[0027] Responding to a collection instruction to obtain a status image of the glove collected by the event camera, and using the brightness information in the status image as the light source information of the light source points that can be collected; wherein, the brightness information includes the light source points that can be collected on the glove.
[0028] In addition, to achieve the above object, the present invention further provides a gesture recognition system including a glove. The gesture recognition system further includes:
[0029] An information acquisition module, configured to acquire the light source information of the light source points that can be collected on the glove; wherein, the light source information includes light source pixel coordinates and light source features;
[0030] A position determination module, configured to determine the light source positions on the glove according to the light source features and a preset light source feature table;
[0031] A gesture determination module, configured to determine a target gesture in a preset deep learning model according to the light source positions and the light source pixel coordinates.
[0032] In addition, to achieve the above object, the present invention further provides a gesture recognition device, including: a memory, a processor, and a gesture recognition program stored on the memory and executable on the processor. When the gesture recognition program is executed by the processor, the steps of the above-mentioned gesture recognition method are implemented.
[0033] In addition, to achieve the above object, the present invention further provides a gesture recognition storage medium, on which a gesture recognition program is stored. When the gesture recognition program is executed by a processor, the steps of the above-mentioned gesture recognition method are implemented.
[0034] The gesture recognition method of the present invention is applied to a gesture recognition system, which includes a glove. By acquiring the light source information of the light source points that can be collected on the glove; wherein, the light source information includes the light source pixel coordinates and the light source characteristics; determining the light source position on the glove according to the light source characteristics and a preset light source characteristic table; determining the target gesture in a preset deep learning model according to the light source position and the light source pixel coordinates. By collecting the light source information of the light source points that can be collected on the glove, and then recognizing the gesture in the preset deep learning model based on the light source information and its corresponding light source position and light source pixel coordinates, it avoids the phenomenon of low accuracy in the bare - hand recognition of users by the built - in camera in the existing VR devices. This gesture recognition method not only provides a non - bare - hand recognition method based on the glove for gesture recognition, but also recognizes the gesture in the preset deep learning model through the light source position and light source pixel coordinates on the glove, thereby providing a basis for gesture recognition and improving the accuracy of gesture recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a schematic structural diagram of a gesture recognition device in the hardware operating environment related to the embodiment solution of the present invention;
[0036] Figure 2 is a schematic flowchart of the first embodiment of the gesture recognition method of the present invention;
[0037] Figure 3 is a schematic module diagram of the gesture recognition system of the present invention;
[0038] Figure 4 is a flowchart of the gesture recognition of the present invention;
[0039] Figure 5 is a physical diagram of a glove in the gesture recognition system of the present invention;
[0040] Figure 6 is another physical diagram of a glove in the gesture recognition system of the present invention.
[0041] The realization of the object, functional characteristics and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0043] Refer to Figure 1 , Figure 1 is a schematic structural diagram of a gesture recognition device in the hardware operating environment related to the embodiment solution of the present invention.
[0044] As Figure 1As shown, the gesture recognition device may include: a processor 0003, such as a Central Processing Unit (CPU), a communication bus 0001, an acquisition interface 0002, a processing interface 0004, and a memory 0005. Among them, the communication bus 0001 is used to implement connection communication between these components. The acquisition interface 0002 may include an information collection system and an acquisition unit such as a computer. Optionally, the acquisition interface 0002 may further include a standard wired interface and a wireless interface. The processing interface 0004 may optionally include a standard wired interface and a wireless interface. The memory 0005 may be a high-speed Random Access Memory (RAM) or a stable Non-Volatile Memory (NVM), such as a disk memory. Optionally, the memory 0005 may also be a storage system independent of the aforementioned processor 0003.
[0045] Those skilled in the art can understand that Figure 1 the structure shown in does not constitute a limitation on the gesture recognition device, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0046] As Figure 1 shown, in the memory 0005 as a storage medium, there may be included an operating system, an acquisition interface module, a processing interface module, and a gesture recognition program.
[0047] In Figure 1 the gesture recognition device shown, the communication bus 0001 is mainly used to implement connection communication between components; the acquisition interface 0002 is mainly used to connect to a background server and perform data communication with the background server; the processing interface 0004 is mainly used to connect to a deployment end (user end) and perform data communication with the deployment end; the processor 0003 and the memory 0005 in the gesture recognition device of the present invention may be arranged in the gesture recognition device, and the gesture recognition device calls the gesture recognition program stored in the memory 0005 through the processor 0003 and executes the gesture recognition method provided by the embodiments of the present invention.
[0048] Based on the above hardware structure, embodiments of the gesture recognition method of the present invention are proposed.
[0049] Embodiments of the present invention provide a gesture recognition method. Referring to Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of the gesture recognition method of the present invention.
[0050] In this embodiment, the gesture recognition method is applied to a gesture recognition system, and the gesture recognition system includes a glove. The gesture recognition method includes:
[0051] Step S10, obtain the light source information of the light source points that can be collected on the glove; wherein, the light source information includes the light source pixel coordinates and the light source characteristics;
[0052] Existing gesture recognition technologies can be divided into two categories: bare - hand recognition and non - bare - hand recognition. The advantage of bare - hand recognition is that it does not need to rely on additional external devices and can achieve interaction through the built - in camera in the AR device, which has great advantages in portability. At the same time, there are problems of low accuracy and lack of interaction feedback between devices when using the camera to recognize bare - hand gestures. Non - bare - hand recognition has high accuracy and can increase feedback through additional sensors, bringing a better interaction comfort. The glove - type peripheral device greatly improves portability due to its advantage of being wearable. However, the existing non - bare - hand recognition requires the use of sensors, resulting in a high cost. The technical problem to be solved by the present invention is: aiming at the phenomenon of high cost of non - bare - hand recognition, a new non - bare - hand recognition method is proposed. The core lies in: encoding the light source by combining the signal characteristics of the event camera, embedding LED light sources at the joint points of the glove, and encoding different light sources according to frequency, so as to realize the extraction of the gesture joint point positions by the event camera and further recognition. Accurately identify the image coordinates of each joint through the optical communication principle, and draw the joint gesture image according to the coordinates to achieve gesture recognition.
[0053] In this embodiment, the gesture recognition system at least includes a glove and an event camera, and the glove and the event camera are communicatively connected (collecting LED light). Inside the event camera, a controller for the entire control can be set to execute the gesture recognition method, or a controller can be set in the gesture recognition system to be connected to the event camera to execute the gesture recognition method. Before gesture recognition, the LED lights are installed on the glove. On the one hand, the entire glove can be installed at a fixed distance. On the other hand, with reference to Figure 5 , Figure 5It is a physical diagram of a glove in a gesture recognition system. LED lights are set at the joints and fingertips of each finger on the glove to detect light source information. At the same time, the LED lights at each position can be encoded in advance. As shown in the figure, the fingertip of the ring finger is encoded as LED1, the first finger joint is encoded as LED2, the second finger joint is encoded as LED3, and the wrist joint is encoded as LED20. Furthermore, all fingertips and joints can be encoded, that is, it can be determined which position the LED light is at, and then gesture recognition can be performed. When recognizing a gesture, the light source information of the light source points that can be collected on the glove will be obtained. Among them, the light source information includes the light source pixel coordinates and the light source characteristics. That is, the event camera is used to determine the light source points that can be collected at each position on the glove, that is, the LEDs that can collect light source information on the glove. At this time, there may be a phenomenon that some LEDs are blocked by the gesture, and then the light source information of the light source points that can be collected will be collected. The light source pixel coordinates refer to the coordinates of the pixel where the light source point is located during collection, and the light source characteristics refer to the characteristics of the light source of the collected light source point, such as frequency, brightness, etc. Furthermore, the position and posture of the entire glove can be determined based on the light source information, so as to judge the posture based on the light source information.
[0054] Step S20, determine the light source position on the glove according to the light source characteristics and a preset light source characteristic table;
[0055] In this embodiment, after determining the light source information, the target light source characteristic corresponding to the light source characteristic in the preset light source characteristic table will be determined. Furthermore, the light source position of the light source point corresponding to the light source characteristic in the preset light source characteristic table can be determined on the glove. Among them, the target light source characteristic refers to the light source characteristic that is the same as the collected light source characteristic in the preset light source characteristic table, the preset light source characteristic table refers to a table recording each light source characteristic and position on the glove, and the light source position refers to the position of the light source point on the glove. For example, after designing the glove, during use, the light sources of different positions of the LEDs are first encoded and marked. The marking method can use frequency encoding, brightness encoding, etc. Taking frequency encoding as an example, the glove has 20 groups of light sources (assuming that the marking starts from the fingertip of the little finger, and the marking method is from top to bottom, from left to right, and finally the wrist joint is marked). According to the light source numbers 1, 2, 3,..., 20, they respectively correspond to different frequencies, such as 0.1 kHz, 0.15 kHz, 0.2 kHz,..., 1 kHz, and the above information is written into the preset light source characteristic table. By using the event camera to collect the light source signal, if the light source characteristic of a certain light source point is 1 kHz, then in the preset light source characteristic table, it is determined that the wrist joint corresponding to label 20 is the light source position of the light source characteristic 1 kHz. Furthermore, the light source position of each light source information collected can be determined, and then the gesture of the glove can be determined based on the light source position.
[0056] Step S30, determine the target gesture in a preset deep learning model according to the light source position and the light source pixel coordinates.
[0057] In this embodiment, by adding LED light sources at the fingertips, knuckles, and wrist joints of the glove, different blinking frequencies are used for different LED light sources to achieve encoding of the light sources and distinguish the light source information. The glove is photographed by an event camera to obtain the light source information, and the joint positions corresponding to the light sources are distinguished according to the light source frequency, thereby realizing the extraction of gesture skeletons. The target gesture is determined in a preset deep learning model according to the light source position and the light source pixel coordinates. The preset deep learning model can be a neural network for gesture recognition or a learning model for gesture recognition. The LED light source here can emit infrared light, which can avoid interfering with other devices and disturbing the user. As shown in Figure 6 , Figure 6 is another physical diagram of the glove in the gesture recognition system. The light source pixel coordinates of 20 points are determined through the light source information, that is, the light source pixel coordinates of 20 points are determined based on the light source position. Then, the respective light source pixel coordinates can be directly connected to form a glove gesture diagram. The connected lines in the figure are theoretical bone lines. Then, the gesture is recognized based on the bone lines. The direct connection method is only applicable to the case where the light source pixel coordinates of all light source points are determined. Then, the gesture can be recognized in a preset deep learning model through the light source position and the light source pixel coordinates on the glove, which provides a basis for gesture recognition and improves the accuracy of gesture recognition.
[0058] Furthermore, this embodiment also provides a schematic diagram of a gesture recognition technical solution. Refer to Figure 4 , Figure 4It is a flowchart for gesture recognition. In this embodiment, event camera data acquisition is performed, where data acquisition refers to obtaining the light source information of the light source points that can be collected on the glove. Then, by identifying the LED light source highlights and coding information, that is, by identifying the light sources of LEDs at different positions to determine which position on the glove the light source is, the LEDs at each position can be encoded in advance before the entire recognition process. Then, a joint diagram can be drawn based on the pixel position and coding of the light source. Among them, the pixel position refers to the light source pixel coordinates included in the light source information, that is, to determine the position of each light source on the glove, and then the entire gesture can be determined. If all are LEDs on the entire glove, directly determining the position of each LED will determine the entire gesture. That is, the most ideal situation is that the information of all the LED lights can be determined, and then the user's gesture can be directly determined. When the light source information collected by the event camera cannot collect the light source information of all the LEDs, the data missing situation is judged, and then the data is supplemented to unify the data format. Finally, a neural network is used to recognize the gesture information, that is, based on the supplemented data and the collected data, the trained neural network or learning model is input to determine the gesture. For example, if about half of the light source information is collected, the other half that has not been collected is supplemented in a preset manner. Finally, the light source information at each position in the light source information and the supplemented light source information are input into the neural network for recognition, thereby realizing gesture recognition. It not only provides a non-barehand recognition method for gesture recognition based on the glove, but also recognizes the gesture in the neural network through the light source position and pixel position on the glove, thereby providing a basis for gesture recognition and improving the accuracy of gesture recognition.
[0059] The gesture recognition method in this embodiment is applied to a gesture recognition system. The gesture recognition system includes a glove, and the light source information of the light source points that can be collected on the glove is obtained. Among them, the light source information includes light source pixel coordinates and light source features. The light source position on the glove is determined according to the light source features and a preset light source feature table. The target gesture is determined in a preset deep learning model according to the light source position and the light source pixel coordinates. By collecting the light source information of the light source points that can be collected on the glove, and then recognizing the gesture in the preset deep learning model based on the light source information and its corresponding light source position and light source pixel coordinates, the phenomenon that the accuracy of the barehand recognition of the user by the built-in camera in the existing VR device is not high is avoided. This gesture recognition method not only provides a non-barehand recognition method for gesture recognition based on the glove, but also recognizes the gesture in the preset deep learning model through the light source position and light source pixel coordinates on the glove, thereby providing a basis for gesture recognition and improving the accuracy of gesture recognition.
[0060] Further, based on the first embodiment of the gesture recognition method of the present invention, a second embodiment of the gesture recognition method of the present invention is proposed. When the collectable light source points are all the light source points on the glove, the step of determining the target gesture in the preset deep learning model according to the light source position and the light source pixel coordinates includes:
[0061] Step a, based on the light source position, determine the target light source pixel coordinates corresponding to each light source point in the light source pixel coordinates; wherein, the target light source pixel coordinates include the light source point and the light source pixel coordinates corresponding to the light source point;
[0062] Step b, determine the gesture output information with the target light source pixel coordinates as the input in the preset deep learning model, and use the gesture output information as the target gesture; wherein, the preset deep learning model includes a pre-trained gesture recognition model.
[0063] In this embodiment, when determining the target gesture, it will be divided into two cases. One is when the collectable light source points are all the light source points on the glove, that is, the light source information of all the light source points on the glove is collected. Based on this, the target light source pixel coordinates corresponding to each light source point will be determined in the light source pixel coordinates based on the light source position. That is to say, the light source position has been determined, and then the light source pixel coordinates corresponding to the light source point corresponding to the light source position are determined. The target light source pixel coordinates refer to the light source pixel coordinates of each light source point. Taking the light source point of the wrist joint as an example, the light source information of the light source point of the wrist joint (not knowing it is the wrist joint, described with the actual light source information) is obtained through the event camera, and then it is determined that this light source information is of the wrist joint based on the light source characteristics in the actual light source information. Furthermore, the light source pixel coordinates in the actual light source information can be used as the target light source pixel coordinates of the wrist joint. That is, the target light source pixel coordinates include the light source point and the light source pixel coordinates corresponding to the light source point, and then the light source pixel coordinates of each light source point can be determined. Finally, determine the gesture output information with the target light source pixel coordinates as the input in the preset deep learning model, and use the gesture output information as the target gesture; wherein, the preset deep learning model includes a pre-trained gesture recognition model. That is, the target light source pixel coordinates of each light source point are used as the input of the entire pre-trained gesture recognition model or pre-trained gesture recognition neural network, and then the actually output gesture can be determined as the target gesture. That is, the function of the above model or neural network is to determine the gesture through the input target light source pixel coordinates of each light source point. Furthermore, the gesture can be determined by the target light source pixel coordinates of the light source point, which improves the accuracy of gesture recognition.
[0064] Further, when the collectable light source points are partial light source points on the glove, the step of determining the target gesture in the preset deep learning model according to the light source position and the light source pixel coordinates further includes:
[0065] Step c: Based on the light source position, determine the target light source pixel coordinates corresponding to each of the light source points in the light source pixel coordinates; wherein, the target light source pixel coordinates include the light source point and the light source pixel coordinates corresponding to the light source point.
[0066] Step d: Based on the collectable light source points and all the light source points on the glove, determine the uncollected light source points, and determine the supplemented light source pixel coordinates of the uncollected light source points; wherein, the supplemented light source pixel coordinates include supplementing the light source pixel coordinates of the uncollected light source points according to a preset rule.
[0067] Step e: In a preset deep learning model, determine the target light source pixel coordinates and the supplemented light source pixel coordinates as the gesture output information for input, and use the gesture output information as the target gesture; wherein, the preset deep learning model includes a pre-trained gesture recognition model.
[0068] In this embodiment, when the collectable light source points are partial light source points on the glove, that is, the light source information of all the light source points on the glove is not collected. Based on this, based on the light source position, the target light source pixel coordinates corresponding to each light source point are determined in the light source pixel coordinates, that is, the light source position has been determined, and then the light source pixel coordinates corresponding to the light source points corresponding to the light source position are determined. The target light source pixel coordinates refer to the light source pixel coordinates of each light source point, that is, the target light source pixel coordinates include the light source point and the light source pixel coordinates corresponding to the light source point, and then the light source pixel coordinates of each light source point (collectable light source points) can be determined. At the same time, there will still be a part of uncollected light source points, which will be determined based on the collectable light source points and all the light source points on the glove, and then the supplemented light source pixel coordinates of the uncollected light source points are determined; wherein, the supplemented light source pixel coordinates include supplementing the light source pixel coordinates of the uncollected light source points according to a preset rule, and the preset rule is a user-defined rule, such as zero coordinates. Finally, in the preset deep learning model, the target light source pixel coordinates and the supplemented light source pixel coordinates are determined as the gesture output information for input, and the gesture output information is used as the target gesture; wherein, the preset deep learning model includes a pre-trained gesture recognition model. That is, the target light source pixel coordinates of each light source point and the supplemented light source pixel coordinates of the uncollected light source points are used as the input of the entire pre-trained gesture recognition model or pre-trained gesture recognition neural network, and then the actually output gesture can be determined as the target gesture, that is, the function of the above model or neural network is to determine the gesture through the input target light source pixel coordinates of each light source point. Furthermore, the gesture can be determined through the target light source pixel coordinates of the light source points, improving the accuracy of gesture recognition.
[0069] Further, based on the first embodiment and / or the second embodiment of the gesture recognition method of the present invention, a third embodiment of the gesture recognition method of the present invention is proposed. The glove is provided with light source points with different characteristics at multiple different setting positions. The gesture recognition method further includes:
[0070] Step f, determining the light source characteristics of each of the light source points and determining the setting positions of each of the light source points on the glove;
[0071] Step g, constructing a preset light source characteristic table for each of the light source points on the glove based on the setting positions and the light source characteristics.
[0072] In this embodiment, the characteristics of the event camera: The information output by the event camera is different from that of traditional sensors. The event camera captures "events", which can be simply understood as "changes in pixel brightness" in this embodiment, that is, the event camera outputs the change situation of pixel brightness. By encoding the frequency of the LED light source, the event camera can obtain the brightness information of the LED and extract the joints based on the light source frequency information. Different LED light sources use different flashing frequencies to achieve encoding of the light source and distinguish the light source information. The glove is photographed by the event camera to obtain the light source information, and the joint positions corresponding to the light sources are distinguished according to the light source frequency, so as to extract the gesture skeleton. Before the entire gesture recognition, light source points with different characteristics are marked at multiple different setting positions on the glove. By determining the light source characteristics of each light source point and simultaneously determining the setting positions of the light source points on the glove, a preset light source characteristic table for each light source point on the glove is constructed based on the setting positions and the light source characteristics. Among them, the light source characteristic refers to the characteristic information of the light source, such as frequency, brightness, etc., and the setting position refers to the installation position on the glove. Finally, a preset light source characteristic table of the characteristics and positions of each light source point on the glove can be determined. That is, determine Figure 5 the light source characteristics emitted by each of the light source points numbered 1-20, and then the light source at which position can be determined during gesture recognition based on the light source characteristics, that is, the pixel position of the light source at that position can be determined, and thus gesture recognition can be achieved.
[0073] Further, the setting positions at least include the fingertip positions and joint positions of the glove. The step of constructing a preset light source characteristic table for each of the light source points on the glove based on the setting positions and the light source characteristics includes:
[0074] Step h, constructing a first preset light source characteristic table for the light source points at the fingertip positions on the glove based on the fingertip positions and the light source characteristics at the fingertip positions;
[0075] Step i, construct a second preset light source feature table for the light source points at the joint positions on the glove based on the joint positions and the light source features of the joint positions, and summarize the first preset light source feature table and the second preset light source feature table to obtain a preset light source feature table.
[0076] In this embodiment, the set positions at least include the fingertip positions and joint positions of the glove, and reference can be made to Figure 5 , the set positions at least include the two finger joints of the thumb, the three finger joints of the other four fingers, as well as each fingertip and the wrist joint. Furthermore, a first preset light source feature table for the light source points at the fingertip positions on the glove can be constructed based on the fingertip positions and the light source features of the fingertip positions. At the same time, a second preset light source feature table for the light source points at the joint positions on the glove can be constructed based on the joint positions and the light source features of the joint positions. Finally, the first preset light source feature table and the second preset light source feature table are summarized to obtain a preset light source feature table. The advantage of constructing two preset light source feature tables here is that since the fingertips are common occluded positions, the fingertips are placed in the first preset light source feature table, and the second preset light source feature table is preferentially judged each time, thereby greatly reducing the time for determining the light source features. Secondly, a low-temperature chip or an easily detectable light source can be set at occluded positions such as the fingertips, thereby greatly improving the detection accuracy.
[0077] Further, based on the first embodiment, the second embodiment, and / or the third embodiment of the gesture recognition method of the present invention, a fourth embodiment of the gesture recognition method of the present invention is proposed. The light source features include light source frequency and light source brightness. The step of determining the light source position on the glove according to the light source features and the preset light source feature table includes:
[0078] Step S21, determine the target light source frequency that matches the light source frequency in the preset light source feature table, and determine the installation position corresponding to the glove of the target light source frequency, and use the installation position as the light source position;
[0079] Step S22, determine the target light source brightness that matches the light source brightness in the preset light source feature table, and determine the installation position corresponding to the glove of the target light source brightness, and use the installation position as the light source position.
[0080] In this embodiment, the light source features are exemplified by including the light source frequency and the light source brightness. When the light source feature is the light source frequency, the target light source frequency that matches the light source frequency in the preset light source feature table is determined, and the installation position corresponding to the glove for the target light source frequency is determined, and the installation position is used as the light source position. Conversely, when the light source feature is the light source brightness, the target light source brightness that matches the light source brightness in the preset light source feature table is determined, and the installation position corresponding to the glove for the target light source brightness is determined, and the installation position is used as the light source position. Among them, the target light source frequency refers to the light source frequency equal to the light source frequency in the preset light source feature table, the installation position refers to the setting position of the target light source frequency defined in the preset light source feature table on the glove, and the target light source brightness refers to the light source brightness equal to the light source brightness in the preset light source feature table. Furthermore, the position of the light source information obtained at this time on the glove can be determined based on the frequency or brightness. At this time, the characteristics of other light sources can also be detected, and thus a basis for determining gestures can be provided based on the determined light source position.
[0081] Furthermore, based on the first embodiment, the second embodiment, the third embodiment, and / or the fourth embodiment of the gesture recognition method of the present invention, a fifth embodiment of the gesture recognition method of the present invention is proposed. Before the step of obtaining the light source information of the light source points that can be collected on the glove, it includes:
[0082] Step S00, obtaining the input training light source information, and training the initial deep learning model based on the training light source information to obtain a preset deep learning model.
[0083] In this embodiment, in the preset deep learning model, the model will be trained in advance. Because in actual use, there will be a phenomenon of blocked light sources. Therefore, by obtaining the input training light source information, and then training the initial deep learning model based on the training light source information to obtain a preset deep learning model, where the training light source information refers to the light source information in the case where there are blocked light source points and the full light source pixel coordinates cannot be determined, and it can also be the case of full light source point light source information in other different situations. The initial deep learning model refers to an untrained model. Furthermore, the model can be trained based on various situations, and then the collected light source information in different situations can be recognized during subsequent gesture recognition, ensuring the accuracy of recognition.
[0084] Furthermore, the gesture recognition system further includes an event camera. The step of obtaining the light source information of the light source points that can be collected on the glove includes:
[0085] Responding to the acquisition instruction to obtain the status image of the glove collected by the event camera, and using the brightness information in the status image as the light source information of the light source points that can be collected; wherein, the brightness information includes the light source points that can be collected on the glove.
[0086] In this embodiment, the event camera is a component of the gesture recognition system, and can also be other instruments for collecting light source points, which is not limited herein. By controlling the event camera to respond to the collection instruction to collect the state image of the glove, that is, the collection instruction is an instruction initiated by the user or internally customized to collect gestures, and the state image refers to the state of the glove at this time, including at least the glove posture, the situation and position of the light source emitting light. Furthermore, the brightness information in the state image is used as the light source information of the collectable light source points; wherein, the brightness information includes the collectable light source points on the glove, that is, the light emitting points of the light sources that can be collected, and subsequent judgments can be made based on the light source information.
[0087] The present invention also provides a schematic diagram of the modules of a gesture recognition system, referring to Figure 3 , the gesture recognition system includes:
[0088] An information acquisition module A01, configured to acquire the light source information of the collectable light source points on the glove; wherein, the light source information includes the light source pixel coordinates and the light source features.
[0089] A position determination module A02, configured to determine the light source position on the glove according to the light source features and a preset light source feature table.
[0090] A gesture determination module A03, configured to determine the target gesture in a preset deep learning model according to the light source position and the light source pixel coordinates.
[0091] Optionally, the gesture determination module A03 includes:
[0092] Based on the light source position, determine the target light source pixel coordinates corresponding to each light source point in the light source pixel coordinates; wherein, the target light source pixel coordinates include the light source point and the light source pixel coordinates corresponding to the light source point.
[0093] Determine the target light source pixel coordinates as the gesture output information input in a preset deep learning model, and use the gesture output information as the target gesture; wherein, the preset deep learning model includes a pre-trained gesture recognition model.
[0094] Optionally, the gesture determination module A03 includes:
[0095] Based on the light source position, determine the target light source pixel coordinates corresponding to each light source point in the light source pixel coordinates; wherein, the target light source pixel coordinates include the light source point and the light source pixel coordinates corresponding to the light source point.
[0096] Determine the uncollected light source points based on the collectable light source points and all the light source points on the glove, and determine the supplementary light source pixel coordinates of the uncollected light source points; wherein the supplementary light source pixel coordinates include supplementing the light source pixel coordinates of the uncollected light source points according to a preset rule;
[0097] The target light source pixel coordinates and the padded light source pixel coordinates are determined in a preset deep learning model as input gesture output information, and the gesture output information is used as a target gesture; wherein the preset deep learning model includes a pre-trained gesture recognition model.
[0098] Optionally, the location determination module A02 includes:
[0099] Determine a target light source frequency that matches the light source frequency in a preset light source characteristic table, and determine an installation position corresponding to the glove of the target light source frequency, and use the installation position as the light source position;
[0100] or,
[0101] Determine a target light source brightness that matches the light source brightness in a preset light source characteristic table, and determine an installation position of the target light source brightness corresponding to the glove, and use the installation position as the light source position.
[0102] Optionally, the information acquisition module A01 includes:
[0103] Determining the light source characteristics of each of the light source points, and determining the setting position of each of the light source points on the glove;
[0104] A preset light source characteristic table of each light source point on the glove is constructed based on the setting position and the light source characteristic.
[0105] Optionally, the information acquisition module A01 includes:
[0106] Constructing a first preset light source feature table of the light source point at the fingertip position on the glove based on the fingertip position and the light source feature at the fingertip position;
[0107] A second preset light source feature table of the light source points at the joint positions on the glove is constructed based on the joint positions and the light source features at the joint positions, and the first preset light source feature table and the second preset light source feature table are summarized to obtain a preset light source feature table.
[0108] Optionally, the information acquisition module A01 includes:
[0109] In response to the acquisition instruction, obtain the status image of the glove acquired by the event camera, and use the luminance information in the status image as the light source information of the collectible light source points; wherein, the luminance information includes the collectible light source points on the glove.
[0110] The methods executed by the above program modules can refer to the various embodiments of the gesture recognition method of the present invention, which will not be elaborated here.
[0111] The present invention also provides a gesture recognition device.
[0112] The device of the present invention includes: a memory, a processor, and a gesture recognition program stored on the memory and executable on the processor. When the gesture recognition program is executed by the processor, the steps of the gesture recognition method as described above are implemented.
[0113] The present invention also provides a storage medium.
[0114] The storage medium of the present invention stores a gesture recognition program. When the gesture recognition program is executed by the processor, the steps of the gesture recognition method as described above are implemented.
[0115] Among them, the method implemented when the gesture recognition program running on the processor is executed can refer to the various embodiments of the gesture recognition method of the present invention, which will not be elaborated here.
[0116] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0117] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0119] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A gesture recognition method, characterized in that, the gesture recognition method is applied to a gesture recognition system, the gesture recognition system includes a glove, and the gesture recognition method includes the following steps: Obtain the light source information of the light source points that can be collected on the glove; wherein, the light source information includes the light source pixel coordinates and the light source characteristics; Determine the light source position on the glove according to the light source characteristics and a preset light source characteristic table; Determine the target gesture in a preset deep learning model according to the light source position and the light source pixel coordinates.
2. The gesture recognition method according to claim 1, characterized in that, when the light source points that can be collected are all the light source points on the glove, the step of determining the target gesture in a preset deep learning model according to the light source position and the light source pixel coordinates includes: Based on the light source position, determine the target light source pixel coordinates corresponding to each light source point in the light source pixel coordinates; wherein, the target light source pixel coordinates include the light source point and the light source pixel coordinates corresponding to the light source point; Determine the target light source pixel coordinates as the input gesture output information in a preset deep learning model, and use the gesture output information as the target gesture; wherein, the preset deep learning model includes a pre-trained gesture recognition model.
3. The gesture recognition method according to claim 1, characterized in that, when the light source points that can be collected are partial light source points on the glove, the step of determining the target gesture in a preset deep learning model according to the light source position and the light source pixel coordinates further includes: Based on the light source position, determine the target light source pixel coordinates corresponding to each light source point in the light source pixel coordinates; wherein, the target light source pixel coordinates include the light source point and the light source pixel coordinates corresponding to the light source point; Based on the light source points that can be collected and all the light source points on the glove, determine the uncollected light source points, and determine the supplemented light source pixel coordinates of the uncollected light source points; wherein, the supplemented light source pixel coordinates include supplementing the light source pixel coordinates of the uncollected light source points according to a preset rule; Determine the target light source pixel coordinates and the supplemented light source pixel coordinates as the input gesture output information in a preset deep learning model, and use the gesture output information as the target gesture; wherein, the preset deep learning model includes a pre-trained gesture recognition model.
4. The gesture recognition method according to claim 1, characterized in that, the light source characteristics include the light source frequency and the light source brightness, and the step of determining the light source position on the glove according to the light source characteristics and a preset light source characteristic table includes: Determine the target light source frequency that matches the light source frequency in the preset light source characteristic table, and determine the installation position corresponding to the glove of the target light source frequency, and use the installation position as the light source position; Or, Determine the target light source brightness that matches the light source brightness in the preset light source characteristic table, and determine the installation position corresponding to the glove of the target light source brightness, and use the installation position as the light source position.
5. The gesture recognition method according to claim 1, characterized in that, The glove is provided with light source points with different characteristics at multiple different setting positions. The gesture recognition method further includes: Determining the light source characteristics of each of the light source points and determining the setting positions of each of the light source points on the glove; Constructing a preset light source characteristic table of each of the light source points on the glove based on the setting positions and the light source characteristics.
6. The gesture recognition method according to claim 5, wherein, The setting positions at least include the fingertip positions and joint positions of the glove. The step of constructing a preset light source characteristic table of each of the light source points on the glove based on the setting positions and the light source characteristics includes: Constructing a first preset light source characteristic table of the light source points at the fingertip positions on the glove based on the fingertip positions and the light source characteristics of the fingertip positions; Constructing a second preset light source characteristic table of the light source points at the joint positions on the glove based on the joint positions and the light source characteristics of the joint positions, and summarizing the first preset light source characteristic table and the second preset light source characteristic table to obtain a preset light source characteristic table.
7. The gesture recognition method according to any one of claims 1 to 6, wherein, The gesture recognition system further includes an event camera. The step of obtaining the light source information of the light source points that can be collected on the glove includes: Responding to a collection instruction to obtain a status image of the glove collected by the event camera, and using the brightness information in the status image as the light source information of the light source points that can be collected; wherein, the brightness information includes the light source points that can be collected on the glove.
8. A gesture recognition system, wherein, The gesture recognition system includes a glove. The gesture recognition system further includes: An information acquisition module for acquiring the light source information of the light source points that can be collected on the glove; wherein, the light source information includes light source pixel coordinates and light source characteristics; A position determination module for determining the light source positions on the glove according to the light source characteristics and a preset light source characteristic table; A gesture determination module for determining a target gesture in a preset deep learning model according to the light source positions and the light source pixel coordinates.
9. A gesture recognition device, wherein, The gesture recognition device includes: a memory, a processor, and a gesture recognition program stored on the memory and executable on the processor. When the gesture recognition program is executed by the processor, the steps of the gesture recognition method according to any one of claims 1 to 7 are implemented.
10. A storage medium, wherein, A gesture recognition program is stored on the storage medium. When the gesture recognition program is executed by a processor, the steps of the gesture recognition method according to any one of claims 1 to 7 are implemented.