3D gesture recognition method, system, electronic device, storage medium and program product

By dividing the 3D gesture recognition scene into near-field and far-field interaction modes, and using 3D and 2D cameras to obtain high-precision and high-resolution information, the problem of poor stability in 3D gesture recognition is solved, and the recognition rate and immersion are improved.

CN122135422APending Publication Date: 2026-06-02BEIJING SHIYAN TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING SHIYAN TECH CO LTD
Filing Date
2024-12-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing 3D gesture recognition technology suffers from poor stability in 3D interaction, affecting the sense of immersion. This is mainly due to problems such as reflection, blind spots, occlusion, and difficulty in background segmentation caused by the free distance between the human hand and the camera.

Method used

The 3D gesture recognition scenario is divided into near-field interaction mode and far-field interaction mode. In the near-field interaction mode, a 3D camera is used to obtain high-precision, low-latency 3D position and key point information. In the far-field interaction mode, a 3D camera is used to obtain 3D position information and a 2D camera is used to obtain high-resolution key point information, and coordinate fusion is performed.

Benefits of technology

It improves gesture recognition rate and efficiency, enhances the immersive experience of 3D interaction, and optimizes the user interaction experience through gesture classification and pattern differentiation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135422A_ABST
    Figure CN122135422A_ABST
Patent Text Reader

Abstract

This invention provides a 3D gesture recognition method, system, electronic device, storage medium, and program product. The method includes: in a near-field interaction mode, acquiring a first 3D image containing a hand captured by a 3D camera; obtaining first 3D position information of the hand based on the first 3D image; obtaining a first 2D image based on the first 3D image; obtaining hand joint information based on the first 2D image; and determining 3D position information of the hand joints based on the first 3D position information and the hand joint information to determine gesture information; in a far-field interaction mode, acquiring a second 3D image containing a hand captured by a 3D camera; acquiring a second 2D image containing a hand captured by a 2D camera; obtaining second 3D position information of the hand based on the second 3D image; obtaining hand joint information based on the second 2D image; and performing coordinate fusion on the second 3D position information and the hand joint information to determine the 3D position information of the hand joints to determine gesture information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of machine vision technology, and in particular to a 3D gesture recognition method, system, electronic device, storage medium and program product. Background Technology

[0002] Currently, gesture recognition methods for interactive humans in three-dimensional (3D) space are a hot research topic. The mainstream approach in this field is to perform gesture recognition based on two-dimensional (2D) images, particularly using 2D cameras to extract gesture information from interactive humans at all distances. However, in 3D interaction, because the distance between the human hand and the camera is relatively flexible, various problems such as reflections, blind spots, occlusion, and difficulties in background segmentation are encountered during actual image processing. These issues result in poor stability of gesture recognition and seriously affect the immersive experience of 3D interaction. Summary of the Invention

[0003] This invention provides a 3D gesture recognition method, system, electronic device, storage medium, and program product to solve the problem of poor stability of 3D gesture recognition in related solutions, which seriously affects the immersive experience of 3D interaction.

[0004] To solve the above-mentioned technical problems, the present invention is implemented as follows:

[0005] In a first aspect, embodiments of the present invention provide a 3D gesture recognition method, including:

[0006] Determine whether to enter near-field interaction mode or far-field interaction mode;

[0007] When entering the near-field interaction mode, a first 3D image containing the interactive person's hand is acquired by the first 3D camera. The first 3D position information of the interactive person's hand is obtained based on the first 3D image. A first 2D image is obtained based on the first 3D image. First hand joint information is obtained based on the first 2D image. The 3D position information of the hand joint is determined based on the first 3D position information and the first hand joint information. The 3D position of the hand joint can be used for gesture recognition.

[0008] When entering the far-field interaction mode, a second 3D image containing the interactive person's hand is acquired by a second 3D camera, and a second 2D image containing the interactive person's hand is acquired by a 2D camera. The second 3D position information of the interactive person's hand is obtained based on the second 3D image, and the second hand joint information is obtained based on the second 2D image. The coordinates of the second 3D position information and the second hand joint information are fused to determine the 3D position information of the hand joints. The 3D position of the hand joints can be used for gesture recognition.

[0009] Optionally, determining whether to enter near-field interaction mode or far-field interaction mode includes:

[0010] Acquire a third 3D image, including the hands of the interacting person, captured by a second 3D camera;

[0011] The distance between the interactive user's hand and the display device is determined based on the third 3D image;

[0012] The near-field interaction mode or the far-field interaction mode is determined based on the distance.

[0013] Optionally, determining whether to enter near-field interaction mode or far-field interaction mode based on the distance includes:

[0014] If the distance between the user's hand and the display device is less than or equal to a first distance, it is determined that the user's hand is located within the near-field interaction space, and the near-field interaction mode is entered.

[0015] If the distance between the user's hand and the display device is greater than a first distance and less than or equal to a second distance, it is determined that the user's hand is located in the far-field interaction space, and the user is determined to enter the far-field interaction mode.

[0016] Optional, also includes:

[0017] In the far-field interaction mode, if the hand of the interacting person is detected to enter the buffer, a preparation operation for entering the near-field interaction mode is performed. The buffer belongs to the far-field interaction space and is at a distance from the display device that is greater than the first distance and less than or equal to the third distance. The third distance is less than the second distance. The preparation operation includes: activating the first 3D camera.

[0018] Optionally, the first 3D image is a phase image; obtaining the first 3D position information of the interactive person's hand based on the first 3D image includes: obtaining a depth image based on the phase image, the depth image containing the depth information of the interactive person; performing threshold segmentation on the depth image, extracting the interactive person's hand, and obtaining the first 3D position information of the interactive person's hand;

[0019] The step of obtaining the first hand joint information based on the first 2D image includes: analyzing the first 2D image using a neural network model to obtain the first hand joint information.

[0020] Optionally, the second 3D image is a phase image; obtaining the second 3D position information of the interactive person's hand based on the second 3D image includes: obtaining a depth image based on the phase image, the depth image containing the depth information of the interactive person; performing threshold segmentation on the depth image, extracting the interactive person's hand, and obtaining the second 3D position information of the interactive person's hand;

[0021] The step of obtaining the second hand joint information based on the second 2D image includes: analyzing the second 2D image using a neural network model to obtain the second hand joint information.

[0022] Optional, also includes:

[0023] In the functional gesture mode, gestures are recognized based on the 3D position information of the hand joints, and the recognized gestures are converted into functional commands for the display device.

[0024] In the operation gesture mode, the corresponding action is performed on the object displayed on the display device based on the 3D position information of the hand, or the gesture is recognized based on the 3D position information of the hand joints, and the corresponding action is performed on the object displayed on the display device based on the gesture.

[0025] Optionally, the functional gestures include at least one of the following: a gesture to activate the gesture function, a gesture to switch between 2D / 3D display modes, a gesture to switch the signal source, and a gesture to deactivate the gesture function;

[0026] and / or

[0027] The operational gestures include at least one of the following: zoom in gesture, zoom out gesture, move gesture, and rotate gesture.

[0028] Optionally, the operation gesture mode includes at least one of the following: high degree of freedom mode, button mode, and symbol gesture mode;

[0029] In the operation-type gesture mode, the method of performing corresponding actions on objects displayed on the display device based on the 3D position information of the hand, or recognizing gestures based on the 3D position information of the hand joints and performing corresponding actions on objects displayed on the display device based on the gestures, includes:

[0030] In the high degree of freedom mode, the relationship between the interactive person's hand and torso is determined based on the 3D position information of the interactive person's hand, and corresponding actions are performed on the objects displayed by the display device based on the relationship.

[0031] and / or

[0032] In the button mode, the interaction state between the hand and the virtual button is determined based on the 3D position information of the hand and the 3D position information of the virtual button displayed on the display device. Based on the interaction state between the hand and the virtual button, the corresponding action is performed on the object displayed on the display device.

[0033] and / or

[0034] In the symbol gesture mode, the gesture is recognized based on the 3D position information of the hand joints, and the corresponding action is performed on the object displayed on the display device based on the recognized gesture.

[0035] Optionally, in the functional gesture mode, recognizing gestures based on the 3D position information of the hand joints and converting the determined gestures into functional commands includes:

[0036] Obtain the duration of the gesture;

[0037] If the duration of the gesture exceeds a preset time, the determined gesture will be converted into a function command.

[0038] Secondly, embodiments of the present invention provide a 3D gesture recognition system, comprising:

[0039] The determination module is used to determine whether to enter the near-field interaction mode or the far-field interaction mode;

[0040] The first execution module is configured to, when entering the near-field interaction mode, acquire a first 3D image of the interactive person's hand captured by a first 3D camera, obtain first 3D position information of the interactive person's hand based on the first 3D image, obtain a first 2D image based on the first 3D image, obtain first hand joint information based on the first 2D image, and determine 3D position information of the hand joint based on the first 3D position information and the first hand joint information, wherein the 3D position of the hand joint can be used for gesture recognition;

[0041] The second execution module is configured to, when entering the far-field interaction mode, acquire a second 3D image containing the interactive person's hand captured by a second 3D camera, and acquire a second 2D image containing the interactive person's hand captured by a 2D camera; obtain second 3D position information of the interactive person's hand based on the second 3D image; obtain second hand joint information based on the second 2D image; and perform coordinate fusion on the second 3D position information and the second hand joint information to determine the 3D position information of the hand joints, wherein the 3D position of the hand joints can be used for gesture recognition.

[0042] Thirdly, embodiments of the present invention provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the 3D gesture recognition method as described in the first aspect above.

[0043] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the 3D gesture recognition method described in the first aspect above.

[0044] Fifthly, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the 3D gesture recognition method as described in the first aspect.

[0045] In this embodiment of the invention, the 3D gesture recognition scenario is divided into near-field interaction mode and far-field interaction mode. Based on the requirement of high precision and low latency in near-field interaction mode, the 3D position information of the interacting person's hand can be determined based on the 3D image captured by the 3D camera, and a 2D image can be obtained based on the 3D image captured by the 3D camera. Hand joint information is obtained based on the 2D image, and the 3D position information of the hand joint is determined based on the 3D position information of the hand and the hand joint information. The 3D position of the hand joint can be used for gesture recognition. Since the 2D image is obtained from the 3D image captured by the 3D camera, there is no need for coordinate transformation between the two, thus the latency is low. Based on the characteristics of far-field interaction mode, which requires high resolution and low fluctuation but has low latency requirements, the 3D position information of the interacting person's hand is obtained from 3D images captured by a 3D camera, and the hand joint information is obtained from 2D images captured by a 2D camera. Coordinate fusion is performed on the 3D position information of the hand and the hand joint information to determine the 3D position information of the hand joints. Gesture information is then determined based on the 3D position information of the hand joints. Because the high-resolution 2D images captured by the 2D camera are used to obtain the hand joint information, high-resolution hand joint information can be obtained, avoiding inaccurate 3D position information of the hand joints. By distinguishing between near-field and far-field gesture recognition, the gesture recognition rate and efficiency can be significantly improved. Attached Figure Description

[0046] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0047] Figure 1 This is a flowchart illustrating the 3D gesture recognition method according to an embodiment of the present invention;

[0048] Figure 2 This is a schematic diagram illustrating the division of the interaction space according to an embodiment of the present invention;

[0049] Figure 3 This is one of the schematic diagrams of hand joint information according to an embodiment of the present invention;

[0050] Figure 4 This is a second schematic diagram of hand joint information according to an embodiment of the present invention;

[0051] Figure 5 This is a schematic diagram of the data flow in the near-field interaction mode according to an embodiment of the present invention;

[0052] Figure 6 This is a schematic diagram illustrating the driving method and data processing method of the depth TOF camera according to an embodiment of the present invention;

[0053] Figure 7 This is a schematic diagram of the information acquisition method under the far-field interaction mode of this invention.

[0054] Figure 8 This is a flowchart illustrating a 3D gesture recognition method according to another embodiment of the present invention;

[0055] Figure 9 This is a flowchart illustrating another embodiment of the 3D gesture recognition method of the present invention;

[0056] Figure 10 This is a schematic diagram illustrating gesture classification according to an embodiment of the present invention;

[0057] Figure 11 This is a schematic diagram of the activation gesture of the functional gesture mode according to an embodiment of the present invention;

[0058] Figure 12 This is a schematic diagram of the operation gesture mode according to an embodiment of the present invention;

[0059] Figure 13 This is a schematic diagram of the structure of the 3D gesture recognition system according to an embodiment of the present invention;

[0060] Figure 14 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] Please refer to Figure 1 This invention provides a 3D gesture recognition method, comprising:

[0063] Step S11: Determine whether to enter near-field interaction mode or far-field interaction mode;

[0064] In this embodiment of the invention, the near-field interaction mode refers to the mode in which the distance between the user's hand and the display device is within a first distance range. At this time, the distance between the user's hand and the display device is relatively close, and it can also be called the near-field mode or the first mode, etc. The far-field interaction mode refers to the mode in which the distance between the user's hand and the display device is within the range of the first distance to the second distance, where the second distance is greater than the first distance. At this time, the distance between the user's hand and the display device is relatively far, and it can also be called the far-field mode or the second mode, etc. The invention does not limit this.

[0065] Step S12: When entering the near-field interaction mode, acquire a first 3D image of the interactive person's hand captured by the first 3D camera, obtain the first 3D position information of the interactive person's hand based on the first 3D image, obtain a first 2D image based on the first 3D image, obtain the first hand joint information based on the first 2D image, determine the 3D position information of the hand joint based on the first 3D position information and the first hand joint information, and the 3D position of the hand joint can be used for gesture recognition;

[0066] In this embodiment of the invention, the first 3D position information of the interactive user's hand can be the preset 3D position information of hand joints, such as the 3D position information of 21 hand joints. However, in actual engineering, due to the influence of jitter noise, the jitter noise of a single hand joint is relatively large. Therefore, the 3D position information of the entire hand can also be obtained. The specific number can be related to the distance between the hand and the display screen. For example, when the distance between the hand and the display screen is closer, more points are collected, and when the distance between the hand and the display screen is farther, fewer points are collected. In this case, the accurate 3D position information of the hand joints can be obtained based on the 3D position information of multiple points around the hand joints. For example, the accurate 3D position information of the hand joints can be obtained by performing median filtering on the 3D position information of multiple points around the hand joints.

[0067] Step S13: When entering the far-field interaction mode, acquire a second 3D image containing the interactive person's hand captured by the second 3D camera, and acquire a second 2D image containing the interactive person's hand captured by the 2D camera. Obtain the second 3D position information of the interactive person's hand based on the second 3D image, obtain the second hand joint information based on the second 2D image, and perform coordinate fusion on the second 3D position information and the second hand joint information to determine the 3D position information of the hand joints. The 3D position of the hand joints can be used for gesture recognition.

[0068] In this embodiment of the invention, optionally, gesture information can be analyzed based on the interrelationship of the 3D position information of hand joints. For example, if the analysis indicates that the distance between all five fingers is greater than the width of two fingers, it is considered to be an action of spreading the five fingers; otherwise, it is considered to be an action of clenching a fist.

[0069] In this embodiment of the invention, the 3D camera can be a depth camera, such as a time-of-flight (TOF) camera or a structured light camera, and the 2D camera can be a grayscale camera or an RGB camera.

[0070] In this embodiment of the invention, the number of the first 3D cameras can be one or more. The number of the second 3D cameras can also be one or more. The number of the 2D cameras can also be one or more.

[0071] In this embodiment of the invention, the first 3D camera and the second 3D camera can be different 3D cameras. Different cameras can have different parameters and / or different placement positions. For example, the first 3D camera used for near-field interaction mode has a large angle and can be placed on the left and right sides of the display screen, such as the lower left and / or lower right corners of the screen. The second 3D camera used for far-field interaction mode has a small angle and can be placed in the middle of the upper and / or lower side of the display screen. 2D cameras are also typically placed in the middle of the upper and / or lower side of the display screen.

[0072] In this embodiment of the invention, 3D position information can also be referred to as 3D coordinate information, or 3D spatial information, or 3D spatial coordinate information, etc.

[0073] In this embodiment of the invention, the hand joint information may include information on 21 hand joints, or information on a subset of the 21 joints. For example, it may include coordinate information of fingertips and joints. The hand joint information may also include the relative positional relationships between joints in a 2D coordinate system.

[0074] In this embodiment of the invention, coordinate fusion refers to coordinate system transformation, which means finding the depth information of the hand's joints in the 3D coordinate system based on the coordinates of the joints in the 2D coordinate system.

[0075] In the above steps, steps S12 and S13 are executed selectively. If the near-field interaction mode is entered, step S12 is executed; if the far-field interaction mode is entered, step S13 is executed.

[0076] In this embodiment of the invention, the 3D gesture recognition scenario is divided into near-field interaction mode and far-field interaction mode. Based on the requirement of high precision and low latency in near-field interaction mode, the 3D position information of the interacting person's hand can be determined based on the 3D image captured by the 3D camera, and a 2D image can be obtained based on the 3D image captured by the 3D camera. Hand joint information is obtained based on the 2D image, and the 3D position information of the hand joint is determined based on the 3D position information of the hand and the hand joint information. The 3D position of the hand joint can be used for gesture recognition. Since the 2D image is obtained from the 3D image captured by the 3D camera, there is no need for coordinate transformation between the two, thus the latency is low. Based on the characteristics of far-field interaction mode, which requires high resolution and low fluctuation but has low latency requirements, the 3D position information of the interacting person's hand is obtained from 3D images captured by a 3D camera, and the hand joint information is obtained from 2D images captured by a 2D camera. Coordinate fusion is performed on the 3D position information of the hand and the hand joint information to determine the 3D position information of the hand joints. Gesture information is then determined based on the 3D position information of the hand joints. Because the high-resolution 2D images captured by the 2D camera are used to obtain the hand joint information, high-resolution hand joint information can be obtained, avoiding inaccurate 3D position information of the hand joints. By distinguishing between near-field and far-field gesture recognition, the gesture recognition rate and efficiency can be significantly improved.

[0077] In this embodiment of the invention, optionally, determining whether to enter the near-field interaction mode or the far-field interaction mode includes:

[0078] Step S111: Acquire a third 3D image containing the hands of the interacting person, captured by the second 3D camera;

[0079] Step S112: Determine the distance between the interactive person's hand and the display device based on the third 3D image;

[0080] In this embodiment of the invention, the distance can also be the distance between the interactive person's hand and the second 3D camera.

[0081] Step S113: Determine whether to enter near-field interaction mode or far-field interaction mode based on the distance.

[0082] In this embodiment of the invention, 3D images captured by a second 3D camera for far-field interaction mode are used to determine the distance between the interactive person's hand and the display device, thereby determining whether to enter near-field interaction mode or far-field interaction mode.

[0083] In other embodiments of the present invention, images captured by a first 3D camera or multiple 2D cameras may be used, or images captured by other cameras different from the first 3D camera, the second 3D camera and the 2D camera may be used to determine the distance between the interactive person's hand and the display device. The present invention does not impose any limitations.

[0084] In other embodiments of the present invention, it is not limited to using images captured by a camera to determine the distance between the interactive person's hand and the display device. It is also possible to determine whether to enter the near-field interaction mode or the far-field interaction mode by means of remote control or voice control.

[0085] In this embodiment of the invention, optionally, determining whether to enter the near-field interaction mode or the far-field interaction mode based on the distance includes:

[0086] Step S1131: If the distance between the interactive person's hand and the display device is less than or equal to the first distance, it is determined that the interactive person's hand is located in the near-field interaction space, and it is determined that the near-field interaction mode is entered.

[0087] Step S1132: If the distance between the interactive person's hand and the display device is greater than the first distance and less than or equal to the second distance, it is determined that the interactive person's hand is located in the far-field interaction space, and the far-field interaction mode is entered.

[0088] The interactive space can also be called the response space, the interactive area, the screen-out interactive range, or the effective display range of the screen, etc.

[0089] For example, the first distance is 50cm, meaning the near-field interaction space is within 0 to 50cm of the display screen. The second distance is 150cm, meaning the far-field interaction space is within 50 to 150cm of the display screen.

[0090] In an embodiment of the present invention, optionally, the 3D gesture recognition method further includes: in the far-field interaction mode, if the hand of the interacting person is detected to enter the buffer, a preparation operation for entering the near-field interaction mode is performed, wherein the buffer belongs to the far-field interaction space and the distance between it and the display device is greater than the first distance and less than or equal to the third distance, wherein the third distance is less than the second distance, and the preparation operation includes: activating the first 3D camera.

[0091] Please refer to Figure 2 , Figure 2This diagram illustrates the division of the interaction space according to an embodiment of the present invention. The near-field precision space is the near-field interaction space, and the far-field response space is the far-field interaction space. The dark blue display screen in the diagram contains 3D cameras located in the lower left and lower right corners for the near-field interaction mode; cameras for the far-field interaction mode are not shown. The yellow cuboids located within the near-field and far-field interaction spaces represent virtual objects displayed on the display screen.

[0092] In this embodiment of the invention, by setting a buffer, preparation operations for entering the near-field interaction mode can be performed in advance when the user is about to enter the near-field interaction space, such as waking up the first 3D camera in advance, thereby improving the response speed.

[0093] In this embodiment of the invention, optionally, in near-field interaction mode, the second 3D camera and 2D camera used for far-field interaction mode can be turned off, and in far-field interaction mode, the first 3D camera used for near-field interaction mode can be turned off, thereby saving power consumption.

[0094] The following sections provide a detailed explanation of the 3D gesture recognition methods in near-field and far-field interaction modes, respectively.

[0095] 1. 3D gesture recognition in near-field interaction mode

[0096] The 3D gesture recognition scheme in near-field interaction mode can also be called the co-optical path scheme, meaning that both the image used to determine the 3D position information of the hand and the image used to determine the hand's joint information come from a 3D camera. In near-field interaction mode, high-speed (low latency) and accurate coordinate recognition of interactive gestures are required, but the resolution requirement is relatively low.

[0097] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the data flow in near-field interaction mode, from... Figure 5 As can be seen from the embodiments of the present invention, a phase image of the interactive person's hand (i.e., the aforementioned first 3D image, the phase of which is determined by the phase of the 3D camera) is acquired using a 3D camera in the near-field interaction space. A depth image can be parsed from the phase image, containing the interactive person's depth information. Threshold segmentation is performed on the depth image to extract the interactive person's hand, obtaining the first 3D position information of the interactive person's hand. Simultaneously, a grayscale image (i.e., the aforementioned first 2D image) is obtained based on the phase image for extracting hand joint information. The method used is to analyze the grayscale image using a neural network model to obtain the first hand joint information. The hand joint information may include the core information of the five fingers and the palm, such as... Figure 3 As shown and Figure 4 As shown, data outside the key points is invalid data, which can effectively reduce the amount of information processing.

[0098] In this scheme, both 2D and 3D images are obtained using a 3D camera. Therefore, the first hand joint information and the first 3D position information of the hand do not need to be fused with coordinates. The 3D position information of the hand joints can be directly obtained. Based on the 3D position information of the hand joints, the gesture information is determined, that is, the joint posture relationship of the gesture is identified. For example, when the five fingers are extended, all joints between the wrist and the five fingers are not folded, and the distance between the joints is greater than the width of two fingers. The gesture posture can be effectively obtained by defining the gesture posture. Then, the action can be judged by temporal multi-posture judgment.

[0099] Please refer to Figure 6 , Figure 6 This diagram illustrates the driving and data processing methods of the depth TOF camera according to an embodiment of the present invention. In this embodiment, the depth TOF camera is initialized and driven by a System on Chip (SOC). Simultaneously, the SOC monitors the working status of the depth TOF camera and controls its core parameters. The SOC drives the depth TOF camera to extract the phase image of the gesture. This image is transmitted through the TOF driver board and the SOC's Mobile Industry Processor Interface (MIPI) interface (or USB interface) to the Image Signal Processor (ISP) core within the SOC. The ISP core extracts depth and grayscale information from the phase image. Through this process, gesture information is obtained. After processing multiple frames of phase images, user behavior analysis results are obtained, and gesture commands are converted. These commands are then transmitted to a PC via an information interface (USB or Ethernet). The PC calls a physical model to perform collision detection between the gesture commands and the displayed 3D content model. The collision results are then incorporated into visual feedback to drive the display screen for 3D (or 2D) display.

[0100] 2. 3D gesture recognition in far-field interaction mode

[0101] The 3D gesture recognition scheme in far-field interaction mode can also be called the heterogeneous path scheme, where the 3D position information for determining the hand comes from a 3D camera, while the image used to determine the hand's joint information comes from a 2D camera. In near-field interaction mode, higher resolution is required, but the requirements for interaction accuracy and latency are reduced.

[0102] In the embodiments of this invention, please refer to Figure 7In the far-field interaction space, a 3D camera is used to acquire a phase image of the interactive person's hand (i.e., the second 3D image mentioned above, the phase of which is determined by the phase of the 3D camera). A depth image can be parsed from the phase image, containing the depth information of the interactive person. Threshold segmentation is performed on the depth image to extract the interactive person's hand, obtaining the second 3D position information of the interactive person's hand. Simultaneously, a 2D camera (such as an infrared (IR) camera) is used to acquire an RGB or grayscale image of the interactive person's hand (i.e., the second 2D image mentioned above) for extracting hand joint information. The method used is to analyze the grayscale image using a neural network model to obtain the second hand joint information, which may include core information of the five fingers and palm, such as... Figure 3 As shown and Figure 4 As shown.

[0103] In this scheme, 2D images are obtained using a 2D camera and 3D images are obtained using a 3D camera. Therefore, the obtained second hand joint information and second 3D position information of the hand need to be fused by coordinates to obtain the 3D position information of the hand joints. Based on the 3D position information of the hand joints, the gesture information is determined, that is, the joint posture relationship of the gesture is identified. For example, in the state of five fingers extended, all joints between the wrist and the five fingers are not folded, and the distance between the joints is greater than the width of two fingers. The gesture posture can be effectively obtained by defining the gesture posture, and the action can be judged by temporal multi-posture judgment.

[0104] In this solution, 2D and 3D images are not on the same optical path and can be customized according to requirements. It can achieve higher resolution and higher frame rate 2D image acquisition, and the 3D image is not affected by the requirements of the 2D image, so that more accurate 3D information can be obtained more quickly.

[0105] In related gesture recognition solutions, all gestures are specifically defined and the same recognition method is used for all gestures. The requirements for gestures during operation are quite strict. In order to improve the recognition rate, the gesture requirements are special and strict. Training is often required before interaction. The fatigue caused by special gestures over a long period of time reduces the user's interaction experience.

[0106] Furthermore, during the gesture collection process, human interaction behavior is unpredictable. After users perform various unpredictable behaviors such as touching their face, rubbing their hands, or moving their hands from outside the interaction area, gesture misdetection will occur. When the interactive user interacts with the information in the display space of the 3D display screen, due to gesture misdetection, the content displayed on the 3D display screen cannot correctly correspond to the interactive user's behavior, which can easily lead to chaotic interactive feedback. Users may become confused and repeat operations in a short period of time, which seriously affects the 3D interactive experience, damages the evaluation of the product, and affects the product value.

[0107] To address the aforementioned issues, this invention categorizes gestures based on their frequency of use and real-time requirements, classifying them into two main categories: functional gestures and operational gestures. Functional gestures (such as switching between 2D and 3D screen displays) require high precision but are not sensitive to real-time performance. Operational gestures (such as gestures for moving or scaling displayed objects) have low tolerance for interaction latency but do not require high precision in gesture recognition. By classifying gestures, corresponding solutions are proposed for each category. For example, for operational gestures, in some cases, only the 3D position information of the hand needs to be recognized, without needing to recognize accurate gesture information. This can effectively improve interaction efficiency, reduce restrictions on user interaction, and enhance the immersive experience. This has significant market application value for glasses-free 3D display products.

[0108] Please refer to Figure 8 To achieve free and efficient gesture recognition, in this embodiment of the invention, optionally, the 3D gesture recognition method further includes:

[0109] Step S14: In the functional gesture mode, the gesture is recognized based on the 3D position information of the hand joints, and the recognized gesture is converted into a function command of the display device;

[0110] In this embodiment of the invention, the functional gesture mode can also be understood as a strong command scenario, such as switching the 2D / 3D display of the monitor, switching from HDMI to DP channel, etc., requiring precise switching commands. The functional gesture mode can also be called a functional mode or a third mode, etc., and this invention does not limit it.

[0111] Functional gestures are used infrequently and have low time response requirements, but high-precision gesture recognition can be achieved by using the 3D position of hand joints to obtain accurate gestures.

[0112] Step S15: In the operation gesture mode, perform a corresponding action on the object displayed on the display device according to the 3D position information of the hand, or recognize the gesture according to the 3D position information of the hand joints and perform a corresponding action on the object displayed on the display device according to the gesture.

[0113] In this embodiment of the invention, the operation gesture mode is a mode for operating virtual objects displayed on the display device, and may also be called operation mode or fourth mode, etc., which are not limited in this invention.

[0114] In this embodiment of the invention, optionally, in the operation gesture mode, a virtual hand can also be displayed through a display device to improve the visual experience.

[0115] In some embodiments, gesture recognition based on the 3D position information of the hand can also be used as an aid in the step of performing corresponding actions on the object displayed by the display device based on the 3D position information of the hand.

[0116] Operational gestures are used frequently and require fast response times, while also needing to prevent hand fatigue from prolonged interaction. The recognition scheme for operational gestures must balance simplicity and speed. In this embodiment of the invention, precise gesture recognition can be omitted; instead, the corresponding action is performed on the object displayed on the display device based directly on the 3D position information of the hand. This effectively reduces the difficulty of operation and improves the degree of freedom and recognition speed. Furthermore, since gestures are not recognized, any loss of recognition would affect the interaction effect, effectively preventing interaction confusion caused by recognition loss.

[0117] In this embodiment of the invention, either step S14 or step S15 may be performed.

[0118] In this embodiment of the invention, by classifying gestures and performing different operations, the contradiction between the two requirements of free operation and precise control of gestures can be resolved, ensuring a good interactive experience for users.

[0119] Please refer to Figure 9 In this embodiment of the invention, a functional gesture mode or an operational gesture mode can be determined by detecting a predetermined gesture. For example, if it is detected that the five fingers are spread open for more than a preset time, it is determined that the functional gesture mode is entered; if it is detected that the hand is raised, it is determined that the operational mode is entered.

[0120] The determination of exceeding the preset time is to prevent false detections.

[0121] The following sections will provide a detailed introduction to functional gestures and operational gestures.

[0122] 1. Functional gestures

[0123] In this embodiment of the invention, the functional gesture can also be called a command gesture. Optionally, in this embodiment of the invention, identifying gestures can also be classified as functional gestures. Identifying gestures and command gestures have similar characteristics.

[0124] In this embodiment of the invention, optionally, in the functional gesture mode, recognizing the gesture based on the 3D position information of the hand joints and converting the determined gesture into a functional command includes: obtaining the duration of the gesture; if the duration of the gesture exceeds a preset time, converting the determined gesture into a functional command.

[0125] Based on the high precision requirements of functional gestures, in this embodiment of the invention, a high-recognition-rate gesture can be used as the activation interface for the functional gesture mode. A high-recognition-rate gesture is, for example, five fingers spread wide; please refer to [reference needed]. Figure 11 When a user makes a specific gesture, such as spreading their five fingers, the entry command for the functional gesture mode is activated. The display device can show a visual feedback mechanism indicating "Gesture initiation" and simultaneously begin a countdown. Examples of countdown implementations include: 1) Displaying a numerical countdown, for example, numbers can be displayed from high to low. If the gesture is held, the countdown continues. If the countdown is complete, the user enters the functional gesture mode; otherwise, they do not. 2) Displaying a progress bar countdown. If the gesture is held, the countdown continues. If the countdown is complete, the user enters the functional gesture mode; otherwise, they do not. 3) Using sound to provide a reminder and assist the user in completing the countdown. When the countdown ends, the functional gesture mode is activated, and the display screen can show a prompt such as "Functional Gesture Mode."

[0126] Please refer to Figure 10 In this embodiment of the invention, optionally, the functional gesture includes at least one of the following: a gesture to activate the gesture function, a gesture to switch between 2D / 3D display modes, a gesture to switch the signal source, and a gesture to deactivate the gesture function. The gesture function is the function of recognizing gestures; if activated, gesture recognition is performed; if deactivated, gesture recognition is not performed.

[0127] In one approach, a single gesture can be used in multiple directions to achieve different functions. In this approach, if the five fingers are again recognized and held for a preset time, an operable interface is displayed. If the hand is moved to the left and held for a preset time, the system enters 2D (or 3D) mode, and a countdown bar appears. Once the countdown ends, the SOC control command sends a command to the screen's main control field-programmable gate array (FPGA) chip to adjust the screen to enter 2D mode. Similarly, gesture-based functions such as closing the gesture and switching the signal source can also be performed.

[0128] In another approach, various gestures can be used to achieve different functions. For example, a two-finger upward gesture can be used to switch between 2D and 3D. When the two fingers are detected to be pointing upward, a reminder of "Switching between 2D and 3D" appears. To prevent accidental operation, a countdown "3..2..1.." is added. If the countdown has not ended, the two fingers disappear and the functional gesture mode is exited. If the countdown ends but the two fingers have not disappeared, the 2D / 3D switching is completed.

[0129] In the above embodiments, the preset time used in different scenarios can be the same or different.

[0130] 2. Operational gestures

[0131] In this embodiment of the invention, optionally, the operational gesture includes at least one of the following: a zoom-in gesture, a zoom-out gesture, a movement gesture, and a rotation gesture. The zoom-in gesture is a gesture that zooms in on the object displayed on the display device; the zoom-out gesture is a gesture that zooms out on the object displayed on the display device; the movement gesture is a gesture that moves the object displayed on the display device; and the rotation gesture is a gesture that rotates the object displayed on the display device.

[0132] Please refer to Figure 12 In this embodiment of the invention, optionally, the operational gesture mode includes at least one of the following: a high degree of freedom mode, a button mode, and a marker gesture mode. The high degree of freedom mode does not recognize gestures, but only determines the relationship between the hand and torso. The button mode combines virtual buttons with the 3D position information of the hand and does not rely on gesture recognition. The marker gesture mode utilizes the 3D position information of joint points to obtain accurate gesture operation recognition results.

[0133] That is, in the operation-type gesture mode, performing corresponding actions on the object displayed on the display device based on the 3D position information of the hand, or recognizing gestures based on the 3D position information of the hand joints and performing corresponding actions on the object displayed on the display device based on the gestures, includes:

[0134] In the high degree of freedom mode, the relationship between the interactive person's hand and torso is determined based on the 3D position information of the interactive person's hand, and corresponding actions are performed on the objects displayed by the display device based on the relationship.

[0135] and / or

[0136] In the button mode, the interaction state between the hand and the virtual button is determined based on the 3D position information of the hand and the 3D position information of the virtual button displayed on the display device. Based on the interaction state between the hand and the virtual button, the corresponding action is performed on the object displayed on the display device.

[0137] and / or

[0138] In the symbol gesture mode, the gesture is recognized based on the 3D position information of the hand joints, and the corresponding action is performed on the object displayed on the display device based on the recognized gesture.

[0139] The high degree of freedom mode, button mode, and symbol gesture mode are explained in detail below.

[0140] The high-degree-of-freedom mode primarily involves body posture recognition, utilizing the relationship between the hand and torso as the entry point for interaction. In one implementation, when the hand and torso enter a linear state and the hand pauses for a preset duration, it is determined that the high-degree-of-freedom mode has been entered, allowing operations such as scaling, rotation, and movement. If the hand and torso are detected to enter a linear state again and the hand pauses for the preset duration, the high-degree-of-freedom mode can be exited via a preset gesture. This approach may not require gesture recognition; however, in some embodiments, gesture recognition may be performed as a supplement.

[0141] The button mode primarily combines virtual 3D buttons with body posture recognition. It uses virtual 3D buttons as the entry point for interaction, determining the interaction state between the hand and the virtual buttons using the 3D position information of the hand and the 3D position information of the virtual buttons displayed on the screen. Based on this interaction state, corresponding actions are performed on the objects displayed on the screen. In some embodiments, when a hand or virtual hand (telescopic operation) presses a virtual 3D button, interactive operations can be performed. However, immediate operation at this time is prone to misoperation. Therefore, the system can pause the hand for a preset time, return the hand to the center of the screen, or a combination of both, or use special gestures such as five fingers to confirm entry into button mode. Once the system determines that button mode has been entered, zooming, rotating, and moving operations can continue. If button mode is entered again, a preset gesture can be used to exit. In this solution, gesture recognition can also be used as a supplement.

[0142] The signature gesture mode primarily involves gesture posture recognition, utilizing special gestures as the entry point for interactive behavior. These special gestures are highly recognizable and offer a high degree of hand freedom. In some embodiments: when a single finger (or two or five fingers) is detected to be open, the system determines that it has entered the operational gesture mode. This is combined with changes in the 3D position information of the special gesture; for example, moving a single finger forward or backward represents scaling, while moving a single finger up or down, left or right allows for rotation, and movement is performed when the virtual hand mapped to a single finger comes into contact with an object. Upon re-entering the operational gesture mode, if the single finger disappears or multi-finger gesture is detected, the system enters a pause state. In other embodiments: when a single finger is detected pointing upward (or two fingers pointing left, five fingers pointing downward, etc.), the system determines that it has entered the operational gesture mode. When the system recognizes the upward finger movement, it activates the interaction. At this time, the hand can perform any gesture for interaction. Moving the hand forward or backward in any posture (except upward finger) represents scaling, while moving it up or down, left or right allows for rotation, and movement is performed when the virtual hand mapped to a single finger comes into contact with an object. Entering the operation gesture mode again allows for exit, such as by detecting a single finger pointing upwards or defining other exit gestures, such as a single finger pointing downwards to enter a pause state. This solution has high requirements for gesture recognition and also has high resource requirements.

[0143] Please refer to Figure 13 This invention also provides a 3D gesture recognition system, comprising:

[0144] The determination module 131 is used to determine whether to enter the near-field interaction mode or the far-field interaction mode.

[0145] The first execution module 132 is configured to, when entering the near-field interaction mode, acquire a first 3D image of the interactive person's hand captured by a first 3D camera, obtain first 3D position information of the interactive person's hand based on the first 3D image, obtain a first 2D image based on the first 3D image, obtain first hand joint information based on the first 2D image, and determine 3D position information of the hand joint based on the first 3D position information and the first hand joint information, wherein the 3D position of the hand joint can be used for gesture recognition;

[0146] The second execution module 133 is configured to, when entering the far-field interaction mode, acquire a second 3D image containing the interactive person's hand captured by a second 3D camera, and acquire a second 2D image containing the interactive person's hand captured by a 2D camera; obtain second 3D position information of the interactive person's hand based on the second 3D image; obtain second hand joint information based on the second 2D image; and perform coordinate fusion on the second 3D position information and the second hand joint information to determine the 3D position information of the hand joints, wherein the 3D position of the hand joints can be used for gesture recognition.

[0147] Optionally, the determining module 131 is used to acquire a third 3D image containing the hands of the interactive person captured by the second 3D camera; determine the distance between the hands of the interactive person and the display device based on the third 3D image; and determine whether to enter the near-field interaction mode or the far-field interaction mode based on the distance.

[0148] Optionally, determining whether to enter the near-field interaction mode or the far-field interaction mode based on the distance includes: if the distance between the user's hand and the display device is less than or equal to a first distance, determining that the user's hand is located within the near-field interaction space, and determining to enter the near-field interaction mode; if the distance between the user's hand and the display device is greater than the first distance and less than or equal to a second distance, determining that the user's hand is located within the far-field interaction space, and determining to enter the far-field interaction mode.

[0149] Optionally, the 3D gesture recognition system further includes:

[0150] The third execution module is used to perform a preparation operation to enter the near-field interaction mode if the hand of the interacting person is detected to enter the buffer in the far-field interaction mode. The buffer belongs to the far-field interaction space and is at a distance greater than the first distance and less than or equal to the third distance from the display device. The third distance is less than the second distance. The preparation operation includes: starting the first 3D camera.

[0151] Optionally, the first 3D image is a phase image; the first execution module 132 is used to obtain a depth image based on the phase image, the depth image containing the depth information of the interactive person; perform threshold segmentation on the depth image to extract the hand of the interactive person and obtain the first 3D position information of the hand of the interactive person; and use a neural network model to analyze the first 2D image to obtain the first hand joint information.

[0152] Optionally, the second 3D image is a phase image; the second execution module 133 is used to obtain a depth image based on the phase image, the depth image containing the depth information of the interactive person; perform threshold segmentation on the depth image to extract the hand of the interactive person and obtain the second 3D position information of the hand of the interactive person; and use a neural network model to analyze the second 2D image to obtain the second hand joint information.

[0153] Optionally, the 3D gesture recognition system further includes:

[0154] The third execution module is used to recognize gestures based on the 3D position information of the hand joints in the functional gesture mode, and convert the recognized gestures into functional instructions for the display device.

[0155] The fourth execution module is used to perform corresponding actions on objects displayed on the display device based on the 3D position information of the hand in the operation gesture mode, or to recognize gestures based on the 3D position information of the hand joints and perform corresponding actions on objects displayed on the display device based on the gestures.

[0156] Optionally, the functional gestures include at least one of the following: a gesture to activate the gesture function, a gesture to switch between 2D / 3D display modes, a gesture to switch the signal source, and a gesture to deactivate the gesture function;

[0157] and / or

[0158] The operational gestures include at least one of the following: zoom in gesture, zoom out gesture, move gesture, and rotate gesture.

[0159] Optionally, the operation gesture mode includes at least one of the following: high degree of freedom mode, button mode, and symbol gesture mode;

[0160] In the operation-type gesture mode, the method of performing corresponding actions on objects displayed on the display device based on the 3D position information of the hand, or recognizing gestures based on the 3D position information of the hand joints and performing corresponding actions on objects displayed on the display device based on the gestures, includes:

[0161] In the high degree of freedom mode, the relationship between the interactive person's hand and torso is determined based on the 3D position information of the interactive person's hand, and corresponding actions are performed on the objects displayed by the display device based on the relationship.

[0162] and / or

[0163] In the button mode, the interaction state between the hand and the virtual button is determined based on the 3D position information of the hand and the 3D position information of the virtual button displayed on the display device. Based on the interaction state between the hand and the virtual button, the corresponding action is performed on the object displayed on the display device.

[0164] and / or

[0165] In the symbol gesture mode, the gesture is recognized based on the 3D position information of the hand joints, and the corresponding action is performed on the object displayed on the display device based on the recognized gesture.

[0166] Optionally, in the functional gesture mode, recognizing the gesture based on the 3D position information of the hand joints and converting the determined gesture into a function command includes: obtaining the duration of the gesture; if the duration of the gesture exceeds a preset time, converting the determined gesture into a function command.

[0167] Please refer to Figure 14 The present invention also provides an electronic device 140, including a processor 141, a memory 142, and a computer program stored in the memory 142 and executable on the processor 141. When the computer program is executed by the processor 141, it implements the various processes of the above-described 3D gesture recognition method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0168] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described 3D gesture recognition method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0169] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0170] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0171] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0172] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of the present invention.

Claims

1. A 3D gesture recognition method, characterized in that, include: Determine whether to enter near-field interaction mode or far-field interaction mode; When entering the near-field interaction mode, a first 3D image containing the interactive person's hand is acquired by the first 3D camera. The first 3D position information of the interactive person's hand is obtained based on the first 3D image. A first 2D image is obtained based on the first 3D image. First hand joint information is obtained based on the first 2D image. The 3D position information of the hand joint is determined based on the first 3D position information and the first hand joint information. The 3D position of the hand joint can be used for gesture recognition. When entering the far-field interaction mode, a second 3D image containing the interactive person's hand is acquired by a second 3D camera, and a second 2D image containing the interactive person's hand is acquired by a 2D camera. The second 3D position information of the interactive person's hand is obtained based on the second 3D image, and the second hand joint information is obtained based on the second 2D image. The coordinates of the second 3D position information and the second hand joint information are fused to determine the 3D position information of the hand joints. The 3D position of the hand joints can be used for gesture recognition.

2. The method according to claim 1, characterized in that, The determination of entering near-field interaction mode or far-field interaction mode includes: Acquire a third 3D image, including the hands of the interacting person, captured by a second 3D camera; The distance between the interactive user's hand and the display device is determined based on the third 3D image; The near-field interaction mode or the far-field interaction mode is determined based on the distance.

3. The method according to claim 2, characterized in that, The step of determining whether to enter near-field interaction mode or far-field interaction mode based on the distance includes: If the distance between the user's hand and the display device is less than or equal to a first distance, it is determined that the user's hand is located within the near-field interaction space, and the near-field interaction mode is entered. If the distance between the user's hand and the display device is greater than a first distance and less than or equal to a second distance, it is determined that the user's hand is located in the far-field interaction space, and the user is determined to enter the far-field interaction mode.

4. The method according to claim 3, characterized in that, Also includes: In the far-field interaction mode, if the hand of the interacting person is detected to enter the buffer, a preparation operation for entering the near-field interaction mode is performed. The buffer belongs to the far-field interaction space and is at a distance from the display device that is greater than the first distance and less than or equal to the third distance. The third distance is less than the second distance. The preparation operation includes: activating the first 3D camera.

5. The method according to claim 1, characterized in that: The first 3D image is a phase image; obtaining the first 3D position information of the interactive person's hand based on the first 3D image includes: obtaining a depth image based on the phase image, the depth image containing the depth information of the interactive person; performing threshold segmentation on the depth image, extracting the interactive person's hand, and obtaining the first 3D position information of the interactive person's hand; The step of obtaining the first hand joint information based on the first 2D image includes: analyzing the first 2D image using a neural network model to obtain the first hand joint information.

6. The method according to claim 1, characterized in that, The second 3D image is a phase image; obtaining the second 3D position information of the interactive person's hand based on the second 3D image includes: obtaining a depth image based on the phase image, the depth image containing the depth information of the interactive person; performing threshold segmentation on the depth image, extracting the interactive person's hand, and obtaining the second 3D position information of the interactive person's hand; The step of obtaining the second hand joint information based on the second 2D image includes: analyzing the second 2D image using a neural network model to obtain the second hand joint information.

7. The method according to claim 1, characterized in that, Also includes: In the functional gesture mode, gestures are recognized based on the 3D position information of the hand joints, and the recognized gestures are converted into functional commands for the display device. In the operation gesture mode, the corresponding action is performed on the object displayed on the display device based on the 3D position information of the hand, or the gesture is recognized based on the 3D position information of the hand joints, and the corresponding action is performed on the object displayed on the display device based on the gesture.

8. The method according to claim 7, characterized in that: The functional gestures include at least one of the following: a gesture to activate the gesture function, a gesture to switch between 2D / 3D display modes, a gesture to switch the signal source, and a gesture to deactivate the gesture function; and / or The operational gestures include at least one of the following: zoom in gesture, zoom out gesture, move gesture, and rotate gesture.

9. The method according to claim 8, characterized in that, The operational gesture mode includes at least one of the following: high degree of freedom mode, button mode, and symbol gesture mode; In the operation-type gesture mode, the method of performing corresponding actions on objects displayed on the display device based on the 3D position information of the hand, or recognizing gestures based on the 3D position information of the hand joints and performing corresponding actions on objects displayed on the display device based on the gestures, includes: In the high degree of freedom mode, the relationship between the interactive person's hand and torso is determined based on the 3D position information of the interactive person's hand, and corresponding actions are performed on the objects displayed by the display device based on the relationship. and / or In the button mode, the interaction state between the hand and the virtual button is determined based on the 3D position information of the hand and the 3D position information of the virtual button displayed on the display device. Based on the interaction state between the hand and the virtual button, the corresponding action is performed on the object displayed on the display device. and / or In the symbol gesture mode, the gesture is recognized based on the 3D position information of the hand joints, and the corresponding action is performed on the object displayed on the display device based on the recognized gesture.

10. The method according to claim 7, characterized in that, In the functional gesture mode, recognizing gestures based on the 3D position information of the hand joints and converting the determined gestures into functional commands includes: Obtain the duration of the gesture; If the duration of the gesture exceeds a preset time, the determined gesture will be converted into a function command.

11. A 3D gesture recognition system, characterized in that, include: The determination module is used to determine whether to enter the near-field interaction mode or the far-field interaction mode; The first execution module is configured to, when entering the near-field interaction mode, acquire a first 3D image of the interactive person's hand captured by a first 3D camera, obtain first 3D position information of the interactive person's hand based on the first 3D image, obtain a first 2D image based on the first 3D image, obtain first hand joint information based on the first 2D image, and determine 3D position information of the hand joint based on the first 3D position information and the first hand joint information, wherein the 3D position of the hand joint can be used for gesture recognition; The second execution module is configured to, when entering the far-field interaction mode, acquire a second 3D image containing the interactive person's hand captured by a second 3D camera, and acquire a second 2D image containing the interactive person's hand captured by a 2D camera; obtain second 3D position information of the interactive person's hand based on the second 3D image; obtain second hand joint information based on the second 2D image; and perform coordinate fusion on the second 3D position information and the second hand joint information to determine the 3D position information of the hand joints, wherein the 3D position of the hand joints can be used for gesture recognition.

12. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the 3D gesture recognition method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the 3D gesture recognition method as described in any one of claims 1 to 10.

14. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the 3D gesture recognition method as described in any one of claims 1 to 10.