A high-precision three-dimensional gesture recognition method and system based on a single light field camera

By using a single light field camera for light field imaging and 3D reconstruction, combined with support vector machines for gesture recognition, the problem of insufficient accuracy and speed in gesture recognition in traditional methods is solved, and high-precision gesture recognition is achieved.

CN117292405BActive Publication Date: 2025-11-25NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311336799.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-13
Publication Date
2025-11-25
Estimated Expiration
2043-10-13

AI Technical Summary

Technical Problem

Existing gesture recognition methods based on RGB or depth cameras are insufficient to meet high requirements in terms of accuracy and speed. Traditional cameras have poor imaging quality and insufficient depth of field in low-light environments, resulting in insufficient accuracy in gesture recognition.

Method used

A single light field camera is used for light field imaging, a refocusing algorithm is used to reconstruct the gesture in three dimensions, and a support vector machine is used for feature extraction and recognition, which simplifies the imaging system and improves the recognition accuracy.

Benefits of technology

It achieves high-precision gesture recognition, reduces the error rate, and improves the accuracy and robustness of recognition, showing good commercial prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117292405B_ABST
    Figure CN117292405B_ABST
Patent Text Reader

Abstract

The application discloses a high-precision gesture recognition method and system based on a light field camera, and the method comprises the following steps: imaging a to-be-detected gesture by using the light field camera; based on a light field reconstruction principle, performing three-dimensional reconstruction on a light field image to obtain a three-dimensional model of the gesture; and performing feature extraction and recognition on the three-dimensional gesture image to recognize the gesture. The system can effectively improve the accuracy of gesture recognition detection by using the light field camera to perform three-dimensional imaging on the gesture. The application can obtain an accurate gesture light field image by using the light field camera, and compared with the existing structured light system and line laser system test method, only one light field camera is needed to perform three-dimensional gesture imaging, the system is simple, and the three-dimensional gesture can effectively improve the accuracy of gesture recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a high-precision three-dimensional gesture recognition method and system based on a single light field camera. Background Technology

[0002] With the rapid development of science and technology, gesture recognition technology has been widely applied in people's daily lives and industrial and commercial applications. In particular, gesture recognition methods based on computer vision can recognize gestures using only RGB or depth cameras, and the accuracy and speed of recognition are quite ideal. However, with the widespread application of gesture recognition technology, people have higher requirements for recognition accuracy and speed.

[0003] The concept of light field was proposed by Michael Faraday in 1846. Light field cameras record the direction information of the light field during the imaging process, and can fuse push and push gestures from different focal points into a single all-focus image. The acquired image undergoes a series of complete algorithmic processing, including digital multi-view, digital refocusing, and 3D reconstruction, to obtain a clear 3D gesture.

[0004] Traditional cameras suffer from focus and defocusing issues when capturing images. When shooting a scene, focusing on nearby objects causes distant objects to go out of focus and become blurry. Furthermore, to ensure sufficient depth of field so that objects at different depths can be in focus, traditional cameras require a small aperture, reducing light efficiency and resulting in poor image quality in low-light conditions. In contrast, light field cameras use microlens arrays to acquire four-dimensional light field information, allowing for imaging with a large aperture while maintaining sufficient depth of field, enabling single-frame 3D imaging of objects.

[0005] Light field imaging is an emerging image acquisition technology capable of capturing depth information of 3D scenes with high precision. This technology has enormous application potential in the field of gesture recognition, enabling 3D gesture reconstruction and thus improving the accuracy of gesture recognition. Summary of the Invention

[0006] The purpose of this invention is to provide a method and system for high-precision gesture recognition based on single-field imaging. This system uses light field imaging as a carrier, employs a refocusing algorithm to reconstruct the gesture in three dimensions, extracts and recognizes its features, and finally outputs the gesture. By utilizing the single-frame three-dimensional acquisition capability of a light field camera, it solves the problem of requiring multiple cameras for three-dimensional imaging in traditional methods, simplifies the imaging system, and obtains a three-dimensional gesture model that accurately represents the real gesture, thereby reducing the error rate of the gesture recognition algorithm and improving the accuracy of recognition.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0008] A high-precision gesture recognition method based on a single-light-field camera includes the following steps:

[0009] 1) Place a gesture in the area to be tested, and use a single light field camera to acquire several light field gesture images containing gesture depth information;

[0010] 2) Refocus the light field gesture image from step 1) to obtain a reconstructed 3D gesture image;

[0011] 3) Match the reconstructed 3D gesture images with predefined gesture target images to filter out valid reconstructed 3D gesture images;

[0012] 4) Use support vector machines to classify and recognize the effective reconstructed 3D gesture images.

[0013] Furthermore, by moving or changing the gesture in the area to be measured, the gesture in the area to be measured is imaged using a single light field camera microlens array to obtain several light field gesture images containing gesture depth information.

[0014] The light field gesture image simultaneously records the information of the gesture light rays on the microlens plane (s,t) and the angle information of the sensor plane (u,v), constructing a four-dimensional light field (u,v,s,t) dual-plane model, whose functional form is L=l(u,v,s,t).

[0015] Further, step 2) specifically includes:

[0016] The light field gesture image from step 1) is refocused using the following formula:

[0017]

[0018] Where f is the focal length, L(u,v,s,t) is the four-dimensional light field, α is the focal plane adjustment coefficient, E(s′,t′) is the intensity value at the position (s,t) of the microlens plane after refocusing, and (u,v) is the sensor plane coordinate.

[0019] Furthermore, after step 2) and before step 3), the process also includes denoising and filtering the reconstructed 3D gesture image.

[0020] Furthermore, step 3) specifically includes:

[0021] The corner points of the reconstructed 3D gesture image and the predefined gesture target image are extracted as the corresponding feature points;

[0022] Calculate the difference in the number of corner points between the reconstructed 3D gesture image and the predefined gesture target image. If the difference does not exceed a set threshold, the corresponding reconstructed 3D gesture image is considered a valid reconstructed 3D gesture image; otherwise, the corresponding reconstructed 3D gesture image is discarded.

[0023] Furthermore, corner points in the image are extracted using the Harris corner detection method.

[0024] On the other hand, the present invention also provides a high-precision gesture recognition system based on a single light field camera, comprising:

[0025] A single light field camera is used to acquire several light field gesture images containing gesture depth information in the area to be measured.

[0026] The reconstruction unit is used to refocus the light field gesture image acquired by the single light field camera to obtain a reconstructed 3D gesture image.

[0027] The feature matching unit is used to match the reconstructed 3D gesture image with a predefined gesture target image and filter out the valid reconstructed 3D gesture images.

[0028] The classification and recognition unit is used to classify and recognize valid reconstructed 3D gesture images using a support vector machine.

[0029] On the other hand, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described above.

[0030] On the other hand, the present invention also provides a high-precision gesture recognition device based on a single light field camera, including one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the methods described above.

[0031] Compared with existing technologies, the technical solution of the present invention has the following advantages:

[0032] First, light field imaging can acquire three-dimensional information in a single frame image. This information can be used to reconstruct three-dimensional gestures, thus providing accurate three-dimensional object information for gesture recognition algorithms. Compared with traditional two-dimensional images, it has higher recognition accuracy.

[0033] Secondly, light field imaging can acquire object information from different perspectives, and has better robustness and stability;

[0034] In addition, light field camera technology has been widely used in industrial and medical fields and has good commercial prospects. Attached Figure Description

[0035] Figure 1 This is a flowchart illustrating the high-precision gesture recognition method using a single-light-field camera as described in this application.

[0036] Figure 2 This is a schematic diagram of the three-dimensional light field acquisition principle in this application;

[0037] Figure 3 This is a flowchart illustrating the principle of light field gesture acquisition and data processing in this application, as well as the output of a three-dimensional gesture after light field data processing. Detailed Implementation

[0038] The following will be combined with the appendix Figure 1 To be continued Figure 3 This application provides a clear and complete description of the light field technology solution, the goal of which is to perform gesture recognition using images captured by a single light field camera.

[0039] like Figure 1 As shown, the method of the present invention includes the following steps:

[0040] Step 1: Light field gesture acquisition:

[0041] a. Place the single-light-field camera on a fixed stand and make a gesture over the area of ​​the object to be measured;

[0042] b. Use a light field camera to record the direction and intensity of light rays arriving at the camera from different angles, thereby obtaining a light field gesture image containing gesture depth information.

[0043] When acquiring light field gesture images with a single light field camera, such as Figure 2 As shown, light rays pass through the lens and microlens array to reach the camera sensor array, where they are recorded. The light field gesture image records the information of the gesture light rays on the microlens plane (s,t), and also records the angular information of the light rays on the sensor plane (u,v). The light field camera constructs a four-dimensional light field (u,v,s,t) biplane model using the microlens plane (s,t) and the sensor plane (u,v), meaning that a light ray passes through two planes and intersects them at (u,v) and (s,t) respectively. The four-dimensional light field can be represented as a function of the entire light field: L = l(u,v,s,t).

[0044] Step 2: Light field gesture data processing:

[0045] Based on the principle of light field digital refocusing, the light field gesture image is refocused using the refocusing method to obtain a reconstructed three-dimensional gesture image.

[0046] The following formula can be used to refocus any plane.

[0047]

[0048] Where L(u,v,s,t) is the four-dimensional light field, f is the focal length, α is the focal plane adjustment coefficient, and E(s′,t′) is the intensity value at the position (s,t) of the refocusing microlens plane.

[0049] Before proceeding with subsequent processing, the reconstructed 3D gesture image needs to be denoised and filtered to improve image quality, thereby improving data quality and accuracy.

[0050] Step 3: Feature Matching

[0051] a. Extract feature points from the reconstructed 3D gesture image;

[0052] b. Based on the extracted feature points, the reconstructed 3D gesture image is matched with the predefined gesture target image to select the effective reconstructed 3D gesture image.

[0053] Specifically, the corner points of the reconstructed 3D gesture image and the predefined gesture target image are extracted using the Harris corner detection method, and the corner points are used as feature points of the image. A corner point has a significant change in certain features relative to its neighboring pixels. When the window function moves in any direction, if the gray value within the window changes significantly, then a corner point is considered to have been detected.

[0054] Specifically, the difference in the number of corner points between the reconstructed 3D gesture image and the predefined gesture target image is calculated. If the difference does not exceed a set threshold, the corresponding reconstructed 3D gesture image is considered a valid reconstructed 3D gesture image; otherwise, the corresponding reconstructed 3D gesture image is discarded.

[0055] Step 4: Use support vector machines to classify and recognize the effective reconstructed 3D gesture images.

[0056] Support vectors satisfy the following conditions:

[0057] r i (w T f+b)≥1

[0058] Among them, w T f+b=0 is the hyperplane of the support vector machine classifier, where f represents the feature vector and w T Let f represent the normal vector corresponding to f, and b represent the linear offset, i = 1, 2, ..., m; divide the gesture into m samples, and find parameters w and b such that the sum of the distances between the support vectors of each category and the linear function of the hyperplane is maximized:

[0059]

[0060] The gesture dataset is divided into various types according to the gesture classification method, and the recognition results are output.

[0061] In summary, this application utilizes a light field camera for high-precision recognition of 3D gestures, enabling convenient acquisition of 3D images of gestures and providing valuable data for high-precision gesture recognition. This improves the accuracy of gesture recognition and opens up more possibilities for interaction between computers and users.

[0062] The present invention also provides a high-precision gesture recognition system based on a single light field camera, comprising:

[0063] A single light field camera is used to acquire several light field gesture images containing gesture depth information in the area to be measured; a single light field camera includes a macro lens, a microlens array, a main lens, and an industrial camera;

[0064] The reconstruction unit is used to refocus the light field gesture image acquired by the single light field camera to obtain a reconstructed 3D gesture image.

[0065] The feature matching unit is used to match the reconstructed 3D gesture image with a predefined gesture target image and filter out the valid reconstructed 3D gesture images.

[0066] The classification and recognition unit is used to classify and recognize valid reconstructed 3D gesture images using a support vector machine.

[0067] The high-precision gesture recognition system based on a single-light-field camera has the same technical solution as the aforementioned method, and will not be described again here.

[0068] Based on the same technical solution, the present invention also discloses a computer-readable storage medium for storing one or more programs, wherein the one or more programs include instructions that, when executed by a computing device, cause the computing device to perform the above-described high-precision gesture recognition method based on a single light field camera.

[0069] Based on the same technical solution, the present invention also discloses a computing device, including one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for executing the above-described high-precision gesture recognition method based on a single light field camera.

[0070] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0071] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0072] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0073] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0074] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art, under the guidance of the present invention, can apply these embodiments to the protection of the present invention without departing from the spirit and scope of the claims.

Claims

1. A high-precision gesture recognition method based on a single-light-field camera, characterized in that, Includes the following steps: 1) Place a gesture in the area to be tested, and use a single light field camera to acquire several light field gesture images containing gesture depth information; 2) Refocus the light field gesture image from step 1) to obtain a reconstructed 3D gesture image; 3) Match the reconstructed 3D gesture images with predefined gesture target images to filter out valid reconstructed 3D gesture images; 4) Use support vector machines to classify and recognize the effectively reconstructed 3D gesture images; The gestures that move or change the area to be measured are imaged using a single light field camera microlens array to obtain several light field gesture images containing gesture depth information. The light field gesture image simultaneously records the information of the gesture light rays in the microlens plane (s,t) and the angle information of the sensor plane (u,v), constructing a four-dimensional light field (u,v,s,t) dual-plane model, whose functional form is L=l(u,v,s,t); Step 2) specifically refers to: The light field gesture image from step 1) is refocused using the following formula: Where f is the focal length, L(u,v,s,t) is the four-dimensional light field, α is the focal plane adjustment coefficient, E(s′,t′) is the intensity value at the position (s,t) of the microlens plane after refocusing, and (u,v) is the sensor plane coordinate. Step 3) specifically includes: The corner points of the reconstructed 3D gesture image and the predefined gesture target image are extracted as corresponding feature points. The difference between the number of corner points of the reconstructed 3D gesture image and the predefined gesture target image is calculated. If the difference does not exceed the set threshold, the corresponding reconstructed 3D gesture image is a valid reconstructed 3D gesture image; otherwise, the corresponding reconstructed 3D gesture image is discarded.

2. The high-precision gesture recognition method based on a single-light-field camera according to claim 1, characterized in that, After step 2) and before step 3), the process also includes denoising and filtering the reconstructed 3D gesture image.

3. The high-precision gesture recognition method based on a single-light-field camera according to claim 1, characterized in that, Corner points in an image are extracted using the Harris corner detection method.

4. A system applying the high-precision gesture recognition method based on a single-light-field camera as described in any one of claims 1 to 3, characterized in that, include: A single light field camera is used to acquire several light field gesture images containing gesture depth information in the area to be measured. The reconstruction unit is used to refocus the light field gesture image acquired by the single light field camera to obtain a reconstructed 3D gesture image. The feature matching unit is used to match the reconstructed 3D gesture image with a predefined gesture target image and filter out the valid reconstructed 3D gesture images. The classification and recognition unit is used to classify and recognize valid reconstructed 3D gesture images using a support vector machine.

5. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 3.

6. A high-precision gesture recognition device based on a single-light-field camera, characterized in that, It includes one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Object three-dimensional reconstruction method based on single-optical-field camera

    CN106296811A

  • Method for realizing dynamic gesture recognition and control in integrated imaging display system

    CN111897433A