Information processing device, information processing system, information processing program, and information processing method

JP7927800B2Active Publication Date: 2026-10-01CANON KK
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2024144138
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-10-01
Estimated Expiration
2044-08-26

AI Technical Summary

Benefits of technology

【0007】 本発明によれば、ハンドジェスチャでCGを掴む際に、手の向きによってはジェスチャ姿勢を計算することが難しいような場合であっても、ジェスチャ姿勢を計算することが可能な技術を提供することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927800000001
    Figure 0007927800000001
  • Figure 0007927800000002
    Figure 0007927800000002
  • Figure 0007927800000003
    Figure 0007927800000003
Patent Text Reader

Abstract

Depending on the orientation of the hand, the fingertip may not appear in the image used for recognition of the hand gesture, and it may be difficult to accurately calculate the orientation in which the CG is grasped by the hand gesture.SOLUTION: An information processing apparatus includes a storage unit configured to store, in a first captured image, information based on a first posture in a case where a first hand included in the first captured image is in a first posture indicating a specific gesture, and an estimation unit configured to estimate, in a second captured image captured after the first captured image, a posture of the second hand indicating the specific gesture in a case where a specific part of the second hand is included in the second captured image even when the second hand included in the second captured image does not indicate the specific gesture, based on the information and the posture of the specific part.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus that performs hand gesture recognition.

Background Art

[0002] As technologies for fusing real-world objects and computer-generated CG (Computer Graphics) in real time, there exist technologies called Mixed Reality (MR) and Augmented Reality (AR). Mixed Reality and Augmented Reality use a device worn on the user's head called a head-mounted display (HMD) to present a composite image of the real world and CG to the user, and enable interaction between the user and CG, thereby providing an immersive experience. Hand gesture operation is one of the operation means for the user to interact with CG. Various sensors such as a camera mounted on the head-mounted display are used to detect the user's hand, and when the hand forms a predetermined gesture, CG is displayed in accordance with the detected hand, thereby making it possible to make it appear as if the user is grasping the object. For example, Patent Document 1 discloses a technology for recognizing hand gestures from a plurality of images including the user's hand.

Prior Art Literature

Patent Literature

[0003]

Patent Document 1

Summary of the Invention

Problem to be Solved by the Invention

[0004] To make it appear as if you are grasping a CG object using hand gestures, it is crucial to display the CG object's orientation precisely to match the actual hand orientation. However, depending on the hand orientation, parts of the hand, such as fingertips, may not be visible in the image used for hand gesture recognition, making it difficult to accurately calculate the orientation in which the CG object is grasped by the hand gesture (hereinafter referred to as the gesture posture).

[0005] This invention was made in view of the above-mentioned problems, and aims to enable the calculation of the gesture posture even when it is difficult to calculate the gesture posture depending on the orientation of the hand when grasping a CG object with a hand gesture. [Means for solving the problem]

[0006] To achieve the above objective, the information processing device of the present invention is characterized by comprising: a storage means for storing information based on a first posture in a storage unit when, in a first captured image, the first hand included in the first captured image is in a first posture that shows a specific gesture; and an estimation means for estimating the posture of the second hand when it shows the specific gesture, based on the information and the posture of the specific part, when, in a second captured image taken after the first captured image, the second hand included in the second captured image does not show the specific gesture, but the second captured image includes a specific part of the second hand. [Effects of the Invention]

[0007] According to the present invention, it is possible to provide a technology that can calculate the gesture posture even when it is difficult to calculate the gesture posture depending on the orientation of the hand when grasping a CG object with a hand gesture. [Brief explanation of the drawing]

[0008] [Figure 1] This is a diagram illustrating the information processing system according to Embodiment 1. [Figure 2] This is a diagram illustrating the internal configuration of the HMD according to Embodiment 1. [Figure 3] This figure illustrates the gesture posture in the pinch gesture and the grab gesture of the hand gesture according to Embodiment 1. [Figure 4] This is a flowchart illustrating the process of calculating gesture posture and generating a display image according to Embodiment 1. [Figure 5] This diagram illustrates the gesture posture and reference posture according to Embodiment 1. [Figure 6] This flowchart illustrates a method for calculating the gesture posture in an captured image using the amount of change from a stored reference posture to the reference posture in the captured image, according to Embodiment 1. [Figure 7] This flowchart illustrates a method for calculating the gesture posture in an captured image from the reference posture in the captured image, using the amount of change from the stored reference posture to the stored gesture posture, according to Embodiment 1. [Figure 8] This figure illustrates the captured image acquired by the HMD100 and the generated display image according to Embodiment 1. [Modes for carrying out the invention]

[0009] The embodiments will be described below with reference to the drawings. The same or equivalent components, members, and processes shown in each drawing will be denoted by the same reference numerals, and redundant explanations will be omitted as appropriate. Furthermore, some components, members, and processes will be omitted from the drawings.

[0010] (Embodiment 1) <System Configuration> Referring to Figure 1, the information processing system 1 according to Embodiment 1 will be described. The information processing system 1 includes a head-mounted display (HMD) 100 and a PC (personal computer) 110.

[0011] HMD 100 is a head-mounted display device (electronic device) that can be worn on the user's head. HMD 100 comprises a camera for capturing an image of an area in front of the user, and a display for displaying an image to the user. The display of HMD 100 displays a composite image obtained by combining a captured image of the area in front of the user captured by HMD 100 with content such as CG (computer graphics) in a form corresponding to the posture of HMD 100. This allows the user to experience virtual reality with the user's own eyes. In addition, system 1 is provided with a function for detecting the user's hand from an image captured by HMD 100, acquiring information related to the position and orientation of the hand as the posture of the hand, and thereby causing the movement of the hand to act on a virtual object. This allows the user to perform intuitive operations on the virtual object using the user's own hand.

[0012] PC 110 controls HMD 100. PC 110 is connected to HMD 100 via a wired connection such as a USB cable, or a wireless connection such as Bluetooth® (registered trademark) or Wi-Fi® (Wireless Fidelity, registered trademark). PC 110 and HMD 100 can communicate with each other via wireless or wired communication, and transmit and receive images and other necessary information to and from each other. PC 110 generates a composite image by combining the image captured by HMD 100 with CG generated by PC 110, and transmits the composite image to HMD 100. Note that although a PC is described herein as an example of the information processing apparatus, the information processing apparatus is not limited thereto. For example, the information processing apparatus may be a smartphone or a tablet terminal, and each configuration of PC 110 may be included in HMD 100.

[0013] <Internal Configuration of HMD> With reference to Figure 2, the internal configuration of HMD 100 will be described. HMD 100 comprises an HMD control unit 201, an imaging unit 202, an image display unit 203, an attitude sensor unit 204, a non-volatile memory 205, and a working memory 206.

[0014] The HMD control unit 201 controls each component of the HMD 100. The HMD control unit 201 has at least one CPU that executes a program stored in the non-volatile memory 205, and at least one other circuit. When the HMD control unit 201 acquires a composite image (an image obtained by the imaging unit 202 capturing the space in front of the user, combined with computer graphics) from the PC 110, it displays the composite image on the image display unit 203. Alternatively, instead of the HMD control unit 201 controlling the entire device, multiple hardware components may share the processing to control the entire device.

[0015] The imaging unit 202 includes two cameras (imaging devices). The two cameras are positioned near the user's left and right eyes when the HMD 100 is worn on the user's head. Therefore, the two cameras can capture the same space as the space seen by the user wearing the HMD 100. Images captured by the imaging unit 202 are output to the HMD control unit 201, which transmits the images from the imaging unit 202 to the PC 110. As described later, the PC 110 combines the captured images transmitted from the HMD 100 with CG to generate a composite image. In addition, the imaging unit 202 acquires a first image with parallax between them and a second image different from the first image at the same time using the two cameras. Therefore, distance information from the HMD 100 to the subject can be obtained using the images from the two cameras in the imaging unit 202. The imaging unit 202 may also record and output video.

[0016] The image display section 203 displays the composite image transmitted from the PC 110 when a composite image is transmitted from the PC 110, as described later. The image display section 203 includes a display such as a liquid crystal panel or an organic EL panel. When the user wears the HMD 100, a display such as an organic EL panel is disposed in front of each of the user's eyes. Note that a device using a transflective half mirror can also be used for the image display section 203. In this case, for example, the image display section 203 may display an image by a technique generally called AR (Augmented Reality) such that CG appears to be directly superimposed on the real space visible through the half mirror. Further, the image display section 203 may display an image of a complete virtual space without using a captured image by a technique generally called VR (Virtual Reality).

[0017] The attitude sensor section 204 acquires attitude (and position) information of the HMD 100. Then, the attitude sensor section 204 acquires attitude information of the user (the user wearing the HMD 100) that corresponds to the attitude (and position) of the HMD 100. The attitude sensor section 204 includes an inertial measurement unit (IMU) constituted by an acceleration sensor, an angular acceleration sensor, and a geomagnetic sensor. When the user wears the HMD 100, the attitude sensor section 204 acquires information on the user's attitude (attitude information). The HMD control section 201 outputs the user's attitude information (attitude information) detected by the attitude sensor 204 to the PC 110.

[0018] The non-volatile memory 205 is an electrically erasable and recordable non-volatile memory, and stores programs and the like to be executed by the HMD control section 201.

[0019] The volatile memory 206 is used as a buffer memory that temporarily holds image data captured by the imaging section 202, an image display memory for the image display section 203, a work area for the HMD control section 201, and the like.

[0020] <Internal Configuration of PC> With reference to Figure 2, the internal configuration of the PC 110 will be described. The PC 110 includes a control unit 211, a non-volatile memory 212, and a working memory 213.

[0021] The control unit 211 is a CPU including at least one processor or circuit. The control unit 211 executes each process of a flowchart described below by executing a program stored in the non-volatile memory 212. Note that, instead of the control unit 211 controlling the entire apparatus, a plurality of pieces of hardware may share processing to control the entire apparatus. The control unit 211 receives, from the HMD 100, an image (captured image) acquired by the imaging unit 202 and posture information acquired by the posture sensor unit 204. The control unit 211 composites the captured image with an arbitrary CG based on the received information to generate a composite image. The control unit 211 transmits the composite image to the HMD control unit 201 in the HMD 100.

[0022] The non-volatile memory 212 is an electrically erasable and recordable non-volatile memory, and stores a program to be executed by the control unit 211, which is described below, and information such as CG. Note that the control unit 211 can switch CG read from the non-volatile memory 212 (that is, CG used for generating a composite image).

[0023] The working memory 213 is a storage unit used as a work area or the like of the control unit 211, such as a buffer memory that temporarily holds image data captured by the imaging unit 202.

[0024] <Description of Gesture Posture of Pinch Gesture> In the information processing system 1 of the present embodiment, when a hand gesture is detected from an image captured by the HMD 100, the system has a function of compositing a virtual object corresponding to the detected hand gesture onto the captured image.

[0025] Figures 3(a) and 3(b) are diagrams for describing a gesture posture of a pinch gesture as a hand gesture in the present embodiment.

[0026] Figure 3(a) illustrates the gesture posture in a pinch gesture.

[0027] The user's hand 301 has its index finger and thumb touching, forming a shape as if pinching something. A gesture in this hand position is called a pinch gesture. When the user's hand 311 is performing a pinch gesture, it is assumed that the user grasps an object between their touching index finger and thumb. The direction in which the object is grasped at this time is called the gesture posture. The gesture posture 302 of a pinch gesture may be set in the direction of the arrow shown in Figure 3(a), for example. Here, the arrow indicating the gesture posture 302 is composed of three orthogonal three-dimensional vectors.

[0028] By using the 3D position 303 of the feature points of the user's hand 301 acquired by the control unit 211, it is possible to determine whether or not the user's hand 301 is making a pinch gesture. In this case, by calculating the distance between the feature points of the index finger and the thumb among the 3D position 303 of the feature points, it is possible to determine that it is a pinch gesture if that distance is below a predetermined threshold. Alternatively, it is also possible to determine whether or not the user's hand 301 is making a pinch gesture by using image recognition technology such as machine learning. In this case, the control unit 211 can use an image recognition model that outputs whether or not the hand included in the image is making a pinch gesture. The image recognition model is created, for example, by training it with images that include a hand in the shape of a pinch gesture. Furthermore, such an image recognition model can perform a classification process to determine whether or not the user's hand 301 is making a pinch gesture. Note that the method of using the 3D position 303 of the feature points of the user's hand 301 acquired by the control unit 211 and the method of using image recognition technology such as machine learning can be used in combination as methods for determining whether or not the user's hand 301 is making a pinch gesture.

[0029] In machine learning models, for example, captured images are input to a Convolutional Neural Network (CNN). The CNN outputs features used to identify the type of subject or the type of scene being photographed. These features are then used to identify the type of subject or scene being photographed. Examples of machine learning models used for detecting and classifying objects or people include YOLO, MobileNet, VGG16, and SSD.

[0030] Next, as a method for determining the gesture posture 302 in a pinch gesture, one can consider using the 3D position 303 of the hand's feature points calculated by the control unit 211. By calculating a vector from multiple feature points among the feature points of the 3D position 303, it is possible to determine the gesture posture 302. For example, the gesture posture 302 can be calculated from feature points corresponding to the tip of the thumb, the base of the thumb, and the base of the index finger. By accurately calculating the gesture posture 302 in accordance with the orientation of the hand, it becomes possible to display an object as if it were being grasped without any sense of incongruity when the object is displayed in the direction of the pinch gesture posture 302.

[0031] Figure 3(b) shows an example of displaying an object according to the gesture position of a pinch gesture.

[0032] Object 314 is a pen-shaped CG (virtual object), and when the user's hand 311 is in a pinch gesture state, the pen-shaped object 314 is displayed along the vector of the gesture posture 312. By displaying it in this way, the user can experience as if they are actually pinching object 314. Furthermore, if the orientation of the user's hand 311 changes, the gesture posture 312 and the posture of object 314 will change accordingly, enabling a more realistic experience.

[0033] <Explanation of the gesture posture for the gripping gesture> Figure 3(c) is a diagram illustrating the gesture posture in the grip gesture as a hand gesture in this embodiment.

[0034] The user's hand 321 has all its fingers bent, forming a hand shape as if grasping something. A gesture in which the hand is in this state is called a grab gesture. When the user's hand 321 is performing a grab gesture, it is assumed that the object will be grasped by wrapping the entire hand around it with all of the fingers and palm. Therefore, the gesture pose 322 of a grab gesture may be set in a direction such as that shown in Figure 3(c). Here, the arrow indicating the gesture pose 322 is composed of three orthogonal three-dimensional vectors.

[0035] The control unit 211 can determine whether the user's hand 321 is a grip gesture by using the three-dimensional position 323 of the feature points of the user's hand 321 acquired by the control unit 211. In this case, the angle between two vectors is calculated using a vector from a predetermined feature point on a predetermined finger to two adjacent feature points on the same finger, with the origin at the predetermined feature point 323 of the feature points. If the angle between the two vectors is below a predetermined threshold, it can be determined that it is a grip gesture. Alternatively, the control unit 211 can use image recognition technology such as machine learning to determine whether the user's hand 321 is a grip gesture. In this case, the control unit 211 can use an image recognition model that outputs whether the hand included in the image is a grip gesture. The image recognition model is created, for example, by training it with images that include a hand in the shape of a grip gesture. Furthermore, such an image recognition model can perform a classification process to determine whether the user's hand 321 is a grip gesture. Furthermore, as a method for determining whether or not the user's hand 321 is a grip gesture, it is possible to use both the method of using the 3D position 323 of the feature points of the user's hand 321 acquired by the control unit 211 and the method of using image recognition technology such as machine learning.

[0036] One method for determining the gesture posture 322 is to use the three-dimensional position 323 of the hand's feature points calculated by the control unit 211. By calculating a vector from multiple feature points among the feature points of the three-dimensional position 323, it is possible to determine the gesture posture 322. For example, the gesture posture 322 can be calculated from feature points corresponding to the wrist, the base of the index finger, the base of the middle finger, and the base of the little finger. By accurately calculating the gesture posture 322 in accordance with the orientation of the hand, it becomes possible to display an object as if it were being grasped without any sense of incongruity when the object is displayed in the direction of the gesture posture 322.

[0037] <How to determine gesture posture and reference posture> Figures 5(a) and 5(b) illustrate the gesture posture and reference posture according to this embodiment.

[0038] The captured image 500 in Figure 5(a) is one frame from a series of frames in a video captured by the imaging unit 202. The captured image 500 includes the user's hand 501, and the control unit 211 acquires the three-dimensional position of each feature point of the user's hand 501 included in the captured image 500. Furthermore, the control unit 211 calculates the reference pose 502 and gesture pose 503 of the user's hand 501 based on the three-dimensional position of each acquired feature point. Here, the reference pose 502 can be calculated by using the pose of a part of the hand, such as the palm or the back of the hand. Even if the fingertips are hidden, the palm or back of the hand has a high detection accuracy. The pose of the palm or back of the hand can be calculated using feature points such as the base of the index finger, the base of the little finger, or the wrist. Also, in Figure 5(a), the gesture pose 503 is shown as a pinch gesture, but other gestures such as a grab gesture may also be used. For example, the gesture could be to bring two fingers together, or to extend one or more fingers. The gesture posture and the reference posture may each include information about position and information about orientation, or they may include only information about orientation.

[0039] The captured image 504 in Figure 5(b) is a single frame from a video captured by the imaging unit 202, and is a frame later in time than the captured image 500. The captured image 504 includes the user's hand 501, and the control unit 211 acquires the three-dimensional position of each feature point of the user's hand 501 included in the captured image 504. Furthermore, based on the acquired three-dimensional position of each feature point, the control unit 211 calculates the reference pose 505 of the user's hand 501. Here, the reference pose 505 is the pose of the same part as the reference pose 502, and in this case, the pose of the palm or the back of the hand is defined as the reference pose 502.

[0040] Based on the reference posture 502 and gesture posture 503 calculated from the captured image 500, and the reference posture 505 calculated from the second captured image 504, the gesture posture 506 in the second captured image 504 is calculated. The specific method for calculating the gesture posture 506 in the second captured image 504 will be the calculation method described later.

[0041] <Explanation of the flowchart for the process of determining gesture posture and generating a display image> Figure 4 is a flowchart of the process for determining a gesture posture and generating a display image in this embodiment. This process is realized by the control unit 211 executing a program stored in the non-volatile memory 212 after it has been loaded into the working memory 213. Note that the timing of the execution of this flowchart is not limited to the timing when a virtual object is displayed in the mixed reality space. For example, it may be the timing when the user starts up the HMD 100 or when the user starts up a predetermined application on the HMD 100. A predetermined application is, for example, an application that the user selects on the home screen (home space) after starting up the HMD, and is an application that allows the user to interact with virtual objects using hand gestures. Furthermore, the following processing may be performed in the virtual space, not just in the mixed reality space. The process in Figure 4 is executed each time a frame of moving image is acquired from the imaging unit 202. In addition, the process in Figure 4 describes a case where a pinch gesture shown in Figure 4 is detected as a specific gesture of the user's hand, and a virtual object corresponding to the pinch gesture is displayed.

[0042] In step S401, the control unit 211 acquires the captured image taken by the imaging unit 202 and proceeds to step S402.

[0043] In step S402, the control unit 211 detects the three-dimensional position of the user's hand feature points based on the image acquired in step S401, and proceeds to step S403. As described above, the processing in S402 is performed on one frame of the moving image captured by the imaging unit 202. The feature points of the user's hand include joint points, which are points that estimate the position of at least one of the joints and fingertips of the hand. For example, 21 points may be acquired as joint points, including 20 points at the fingertips, first joints, second joints, and bases of each finger, and 1 point at the wrist.

[0044] In step S403, the control unit 211 determines whether or not it was able to detect the three-dimensional position of the feature points (joint points) of the user's hand. If the control unit 211 was able to detect the three-dimensional position of the feature points (joint points) of the user's hand, it proceeds to step S404. If it was not able to detect the three-dimensional position of the feature points (joint points) of the user's hand, it proceeds to step S409. For example, if the hand is not visible in the captured image, the three-dimensional position of the feature points (joint points) of the user's hand cannot be detected, and the process proceeds to step S409.

[0045] In step S404, the control unit 211 determines whether the user's hand is performing a specific gesture. If the control unit 211 determines that the hand is performing a specific gesture, it proceeds to step S405; if it determines that the hand is not performing a specific gesture, it proceeds to step S407. In other words, if the hand included in the captured image does not show a specific gesture, or if the hand included in the captured image is performing a gesture different from the specific gesture, it is determined that the hand is not performing a specific gesture, and the process proceeds to step S407.

[0046] In step S405, the control unit 211 calculates the reference pose and gesture pose in the captured image and proceeds to step S406.

[0047] In step S406, the control unit 211 calculates correction information for the gesture posture from the reference posture and gesture posture in the captured image calculated in step S405, stores it in the working memory 213, and proceeds to step S409. Note that the reference posture and gesture posture may be stored in the working memory 213 without calculating the correction information, or only the gesture posture may be stored in the working memory 213. Here, the correction information includes, for example, information regarding the difference in rotation amount and position of the reference posture and gesture posture in the captured image calculated in step S405. Using such a difference and the reference posture, the gesture posture that could not be fully estimated from the image alone can be calculated (estimated) using the method described later.

[0048] In step S407, the control unit 211 reads (acquires) correction information from the working memory 213 and proceeds to step S408. If correction information cannot be acquired, or if correction information is not stored in the working memory 213, the control unit 211 proceeds to step S408 without acquiring correction information in step S407. Here, if correction information is stored in the working memory 213, the control unit 211 reads the correction information from the working memory 213. If the reference posture and gesture posture are stored in the working memory 213, the control unit 211 reads the reference posture and gesture posture from the working memory 213 and calculates the correction information. If the gesture posture is stored in the working memory 213 but the reference posture is not, the control unit 211 reads the gesture posture from the working memory 213, extracts the reference posture from the gesture posture, and then calculates the correction information from the reference posture and gesture posture. Furthermore, if neither the reference posture nor the gesture posture is stored in the working memory 213, the system proceeds to step S408 without reading (acquiring) any correction information.

[0049] In step S408, the control unit 211 calculates the gesture posture based on the correction information acquired in step S407 and reference postures such as the palm and back of the hand detected from the captured image. The method for calculating the gesture posture in step S408 will be described later. Furthermore, even if correction information is acquired in step S407, if the gesture posture can be detected from the captured image without using the correction information, the gesture posture will be detected without using the correction information.

[0050] In step S409, the control unit 211 draws a virtual object and proceeds to step S410. Here, the virtual object is drawn as if the user is holding it with their thumb and index finger using a specific gesture, a pinch gesture. Therefore, based on the calculated gesture pose, the control unit 211 calculates the orientation of the virtual object when the user is holding it with their thumb and index finger using a pinch gesture. Then, it draws the virtual object in the calculated orientation. If there are virtual objects placed in the mixed reality space, they are also drawn in step S409. If the gesture pose has not been calculated, virtual objects placed in the mixed reality space are drawn.

[0051] In step S410, the control unit 211 generates a display image to be displayed on the image display unit 203 of the HMD 100 using the virtual object drawn in step S409 (performs image generation), and proceeds to step S411. After generating the display image, the control unit 211 may further send the display image to the HMD 100 and perform display control to display the display image on the image display unit 203 of the HMD 100.

[0052] In step S411, the control unit 211 determines whether or not to terminate the process. If the control unit 211 determines to terminate the process, it proceeds to step S412; otherwise, it proceeds to step S401.

[0053] In step S412, if the control unit 211 has stored the correction information in the working memory 213, it deletes the correction information and terminates the process. Alternatively, the correction information may be deleted when the power to the HMD100 is turned off, and not deleted when a specific application is terminated.

[0054] According to the above flow, for example, if the captured image includes a hand with the index finger and thumb separated, it is determined that the hand is not performing a specific gesture, and the process proceeds to step S407. In step S407, if correction information is stored, the correction information is acquired. However, in step S408, the gesture posture is detected from the hand with the index finger and thumb separated without using the correction information. Furthermore, if an image containing a hand with the index finger and thumb separated is acquired after an image of a hand performing a specific gesture, the virtual object is drawn in a way that indicates it has not moved. In other words, the virtual object is drawn in a way that indicates it has not moved from the position it was in when the image of the hand performing the specific gesture was acquired.

[0055] Here, if correction information is calculated from the reference pose and gesture pose in the captured image in step S406 and stored in the working memory 213, the computational load in step S407 is reduced compared to the case where the correction information is not stored. When real-time video is important, it is preferable to calculate the reference pose and gesture pose in the captured image and store them in the working memory 213 in advance.

[0056] Furthermore, if the gesture posture in the captured image is stored in step S406, but the reference posture is not stored, the reference posture is calculated from the stored gesture posture in step S407. Storing the gesture posture in the captured image in step S406 reduces the amount of data stored in the working memory 213 compared to the case where correction information is not calculated and both the gesture posture and the reference posture are stored. If the working memory 213 in this system is experiencing a shortage of data, it is preferable to store the gesture posture in the captured image but not the reference posture.

[0057] In the flow shown in Figure 4, correction information is stored every time a specific gesture is performed. However, the flow is not limited to this; correction information may be stored only in the first frame in which a gesture is detected. Furthermore, if multiple correction information sets are stored, the average of the multiple correction information sets may be used to calculate the gesture posture when part of the hand is hidden. If the gesture is detected multiple times, the correction information may be overwritten each time and stored in the working memory 213.

[0058] Alternatively, correction information may be acquired from the first frame when the first gesture is detected and stored in the working memory 213, but not stored until the end of the first gesture. Then, when the second gesture is detected, the correction information acquired from the first frame when the second gesture is detected may be overwritten and stored in the working memory 213. In other words, the working memory 213 may only acquire and store correction information for the first frame detected in a series of gestures, and the correction information may only be used within that series of gestures.

[0059] As shown in Figure 4, calculating the gesture pose of the captured image based on multiple image frames makes it possible to determine the gesture pose more stably than calculating it from information from only the second captured image.

[0060] <First method for calculating gesture posture when a reference posture and a gesture posture are stored> Referring to Figure 6, a flowchart illustrating a method for calculating the gesture posture is shown, which uses the rotation amount θ from the orientation of the stored reference posture to the orientation of the reference posture detected from the captured image to calculate the gesture posture in the captured image. This process is executed, for example, in step S407 when the gesture posture and reference posture are stored in the working memory 213 in step S406 of Figure 4.

[0061] In step S601, the control unit 211 calculates the rotation amount θ1, which is the difference between the orientation of the reference orientation stored in the working memory 213 and the orientation of the reference orientation detected from the captured image, and proceeds to step S602. Note that the orientation of each reference orientation may be represented by a 3x3 rotation matrix or quaternion, and the rotation amount θ1 may be calculated accordingly.

[0062] In step S602, the control unit 211 calculates the orientation of the gesture posture in the captured image by rotating the gesture posture stored in the working memory 213 by the rotation amount θ1 obtained in step S601.

[0063] In step S603, the control unit 211 calculates the change in position (Δx1, Δy1, Δz1), which is the difference between the position of the reference posture stored in the working memory 213 and the position of the reference posture detected from the captured image, and proceeds to step S604. Note that each of the reference postures and the change in position may be expressed in a three-dimensional coordinate system or in a predefined two-dimensional coordinate system.

[0064] In step S604, the control unit 211 calculates the position of the gesture posture in the captured image by moving the gesture posture stored in the working memory 213 by the amount of position change (Δx1, Δy1, Δz1) obtained in step S603.

[0065] As shown in Figure 6, even if a part of the user's hand used for calculating the gesture posture is hidden in the captured image, it is possible to reliably determine the gesture posture.

[0066] <Second method for calculating gesture posture when the reference posture and gesture posture are stored> Referring to Figure 7, a flowchart illustrating a method for calculating the gesture pose in an captured image using the relationship between the position and orientation from a stored reference pose to a stored gesture pose is described. This process is executed, for example, in step S407 of Figure 4, when the gesture pose and reference pose are stored in the working memory 213 in step S406. Alternatively, the process in step S701 below may be performed in step S406 of Figure 4 to calculate the correction information and store it in the working memory, and the process in step S702 below may be performed in step S407 of Figure 4 to calculate the gesture pose.

[0067] In step S701, the control unit 211 calculates the rotation amount θ2, which is the difference between the orientation of the reference posture stored in the working memory 213 and the orientation of the gesture posture stored in the working memory 213, and proceeds to step S702. Note that the orientation of each reference posture may be represented by a 3x3 rotation matrix or quaternion, and the rotation amount θ2 may be calculated from these. In this case, if there are multiple combinations of reference posture and gesture posture stored in the working memory 213, the rotation amount from each reference posture to the gesture posture may be calculated and then the average value of these values ​​may be used.

[0068] In step S702, the control unit 211 calculates the orientation of the gesture posture in the captured image by rotating the reference posture in the captured image by the amount of rotation θ2 obtained in step S701.

[0069] In step S703, the control unit 211 calculates the change in position (Δx2, Δy2, Δz2), which is the difference between the position of the reference posture stored in the working memory 213 and the position of the gesture posture stored in the working memory 213, and proceeds to step S704. The reference posture, gesture posture, and change in position may be expressed in a three-dimensional coordinate system or in a predefined two-dimensional coordinate system. In this case, if there are multiple combinations of reference posture and gesture posture stored in the working memory 213, the change in position from each reference posture to the gesture posture may be calculated and then the average value of these values ​​may be used.

[0070] In step S704, the control unit 211 calculates the position of the gesture posture in the captured image by moving it relative to the reference posture in the captured image by the amount of change in position (Δx2, Δy2, Δz2) determined in step S703.

[0071] As shown in Figure 7, even if a part of the user's hand used for calculating the gesture posture is hidden in the captured image, it is possible to reliably determine the gesture posture.

[0072] <Relationship between captured images and workflow> Referring to Figures 8(a), 8(b), 8(c), 8(d), 8(e), 8(f), and 8(g), the relationship between the images acquired by the HMD100 and the processing performed in the flow shown in Figure 4 will be explained.

[0073] Figure 8(a) shows a scene in which the captured image 800 includes (captures) a hand 801 demonstrating a specific gesture. If an image like the captured image 800 is acquired in step S401 of Figure 4, the control unit 211 assumes in step S404 that the hand is demonstrating a specific gesture and proceeds to step S405.

[0074] Figure 8(b) shows the reference posture 802 and gesture posture 803 calculated in step S405. When such a reference posture 802 and gesture posture 803 are obtained, in step S406, the correction information obtained by calculating the difference between the gesture posture 803 and the reference posture 802 is stored in the working memory 213.

[0075] Figure 8(c) shows the display image 820 generated from the captured image 800 in step S410. In the display image 820, a virtual object 804 drawn based on the gesture pose 803 in step S409 is superimposed.

[0076] Figure 8(d) shows a frame in the captured image 840 that includes (shows) a hand 841 that does not exhibit a specific gesture, assuming that the captured image 840 is acquired after the captured image 800. In the captured image 840, it is assumed that the hand's orientation has changed while it is moving while maintaining the same shape as the hand 801 in the captured image 800. That is, it is assumed that the user is still in a pinch pose. However, in the captured image 840, the back of the hand is visible in the image, but the thumb and index finger are not. Therefore, the position of the fingers cannot be confirmed in the image, and it is not possible to determine whether all the fingers are folded or whether the hand is in a pinch pose similar to the hand 801. Thus, if an image like the captured image 840 is acquired in step S401 of Figure 4, the control unit 211 assumes in step S404 that the hand does not exhibit a specific gesture and proceeds to step S407.

[0077] Figure 8(e) shows the reference posture 852 in the captured image 840, calculated in step S407 or step S408. Based solely on the information in this image, it is possible to estimate the position of the fingers, but it is not possible to verify whether the estimation result is truly correct.

[0078] Figure 8(f) shows the gesture pose 863 calculated in step S408. Here, the gesture pose 863 is calculated from the correction information calculated based on the captured image 800 and stored in the working memory 213, and from the reference pose 852 obtained from the captured image 840. If the captured image 840 is obtained, in step S409, the orientation for displaying the virtual object is calculated based on the gesture pose 863, and the virtual object is drawn.

[0079] Figure 8(g) shows the display image 870 generated in step S410 by superimposing the virtual object 874 drawn in step S409 onto the captured image 840. Here, the control unit 211 acquires depth information of the hand and expresses the depth relationship between the virtual object and the hand 841, so it is assumed that the part of the virtual object 874 hidden behind the hand is not superimposed on the captured image 840.

[0080] In this way, after a specific gesture is detected and the state of grasping a virtual object is detected, even if part of the user's hand is obscured and the specific gesture is no longer shown, if the specific part is still included, the posture (shape) of the specific gesture is estimated. Then, based on the estimated posture (shape) of the specific gesture, the virtual object is drawn and a display image is generated, so that the state of grasping a virtual object can be represented even if part of the hand is hidden. In other words, although the captured image 840 alone does not show a specific gesture (it is not possible to know that a specific gesture is being shown), the posture of hand 841 is estimated assuming that hand 841 is showing a specific gesture. To put it another way, based on the information obtained from the captured image 800, it can be said that the posture of hand 841 is estimated by assuming that hand 841 in captured image 840 is showing a specific gesture.

[0081] (Other embodiments) Furthermore, the present invention can also be realized by performing the following process: that is, supplying software (program) that realizes the functions of the above-described embodiment to a system or device via a network or various storage media, and having the computer (or control unit or MPU, etc.) of the system or device read and execute the program code. In this case, the program and the storage medium storing the program constitute the present invention.

[0082] Although the present invention has been described in detail above based on its preferred embodiments, the present invention is not limited to these specific embodiments, and various forms that do not depart from the spirit of the invention are also included in the present invention. Some of the above embodiments may be combined as appropriate.

[0083] Furthermore, each functional unit in each of the above embodiments (each modified example) may or may not be individual hardware. The functions of two or more functional units may be implemented by common hardware. Each of the multiple functions of a single functional unit may be implemented by individual hardware. Two or more functions of a single functional unit may be implemented by common hardware. In addition, each functional unit may or may not be implemented by hardware such as an ASIC, FPGA, or DSP. For example, the device may have a processor and a memory (storage medium) in which a control program is stored. The functions of at least some of the functional units of the device may be implemented by the processor reading and executing the control program from the memory.

[0084] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0085] Furthermore, in each of the examples described above, "processor" refers to a processor in a broad sense, including general-purpose processors (e.g., CPUs) and specialized processors (e.g., GPUs, ASICs, FPGAs, and programmable logic devices, etc.).

[0086] This embodiment includes the following configurations, methods, and programs.

[0087] [Configuration 1] A storage means for storing information based on a first posture in a storage unit when, in a first captured image, the first hand included in the first captured image is in a first posture that indicates a specific gesture, In a second image captured after the first image, even if the second hand included in the second image does not exhibit the specific gesture, if the second image includes a specific part of the second hand, the system includes estimation means for estimating the posture of the second hand when exhibiting the specific gesture, based on the information and the posture of the specific part. An information processing device characterized by the following:

[0088] [Configuration 2] The storage means stores the first posture as information, If the second captured image includes the specific body part, even if the second hand does not demonstrate the specific gesture in the second captured image, the estimation means estimates the posture of the second hand demonstrating the specific gesture based on the first posture and the posture of the specific body part of the second hand. The information processing device according to configuration 1, characterized by the above.

[0089] [Configuration 3] The storage means further stores the posture of the specific part of the first hand as information, If the second captured image includes the specific body part, even if the second hand does not demonstrate the specific gesture in the second captured image, the estimation means estimates the posture of the second hand demonstrating the specific gesture based on the first posture, the posture of the specific body part of the first hand, and the posture of the specific body part of the second hand. An information processing apparatus according to configuration 1 or 2, characterized by the above.

[0090] [Structure 4] The storage means stores, as information, the difference in position and orientation between the first posture and the posture of the specific part of the first hand. If the second captured image includes the specific part, even if the second hand does not demonstrate the specific gesture in the second captured image, the estimation means estimates the posture of the second hand demonstrating the specific gesture based on the difference and the posture of the specific part of the second hand. An information processing device according to any one of configurations 1 to 3.

[0091] [Composition 5] The system further includes determination means for determining whether a hand included in the captured image is performing the specific gesture, If the determination means determines that the first hand is performing the specific gesture, the storage means stores the information in the storage unit. Even if the determination means does not determine that the second hand is performing the specific gesture because a part of the hand is hidden in the second captured image, if the second captured image includes the specific part, the estimation means estimates the posture of the second hand performing the specific gesture based on the information and the posture of the specific part. An information processing apparatus according to any one of configurations 1 to 4, characterized by the above.

[0092] [Composition 6] The determination means determines whether the hand performs the specific gesture based on the positions of a plurality of joint points, which are points that estimate the position of at least one of the joints of the hand and the fingertips. The information processing apparatus according to configuration 5, characterized by the features described herein.

[0093] [Composition 7] The determination means determines whether the hand performs the specific gesture by classifying it into a class. The information processing apparatus according to configuration 5 or 6, characterized by the above.

[0094] [Structure 8] The system further includes detection means for detecting the positions of multiple joint points, which are points that estimate the position of at least one of the joints of the hand and the fingertips from the captured image. The storage means stores the information in the storage unit based on the positions of a plurality of joint points, which are points that estimate the position of at least one of the joints of the hand and the fingertips. An information processing device according to any one of configurations 1 to 7, characterized by the above.

[0095] [Composition 9] The aforementioned specific area is a part of the hand that offers a high degree of detection accuracy. An information processing device according to any one of configurations 1 to 8.

[0096] [Configuration 10] The aforementioned specific gesture is a pinch gesture, where the thumb and index finger are brought together, or a gripping gesture, where the hand is clenched. An information processing apparatus according to any one of configurations 1 to 9, characterized by the above.

[0097] [Composition 11] A storage means for storing information based on a first posture in a storage unit when, in a first captured image, the first hand included in the first captured image is in a first posture that indicates a specific gesture, The system includes a display control means that controls the display unit to display a first image in which a virtual object with an orientation corresponding to the first posture is synthesized when the first hand is in the first posture in the first captured image, The display control means controls the display unit to display a second image on the display unit in which the virtual object is synthesized in an orientation corresponding to the posture of the second part of the second hand, if the second image, captured after the first image, includes a specific part of the second hand, even if the second hand included in the second image does not perform the specific gesture. An information processing device characterized by the following:

[0098] [Composition 12] The storage means stores, as information, the difference in position and orientation between the first posture and the posture of the specific part of the first hand. The display control means controls the display unit to display the second image, which is generated based on the difference and the posture of the specific part. The information processing apparatus according to configuration 11, characterized by the features described above.

[0099] [Composition 13] The second captured image further includes estimation means for estimating the posture of the second hand when performing the specific gesture, based on the information and the posture of the specific part. An information processing apparatus according to configuration 11 or 12, characterized by the above.

[0100] [Composition 14] The system further comprises a generation means for generating a virtual object based on the second hand pose estimated by the estimation means. The information processing device according to configuration 13, characterized by the above.

[0101] [Composition 15] The system further includes image generation means that synthesizes the virtual object oriented to the first posture with the first captured image to generate the first image, and synthesizes the virtual object oriented to the posture of the specific part with the second captured image to generate the second image. An information processing device according to any one of the configurations 11 to 14, characterized by the above.

[0102] [program] A program for causing a computer to function as one of the means of an information processing device described in any one of items 1 to 15.

[0103] [method] A storage step in which, in the first captured image, if the first hand included in the first captured image is in a first posture that indicates a specific gesture, information based on the first posture is stored in the storage unit; If a second image captured after the first image capture does not show the specific gesture, but the second image capture includes a second part of the second hand, the estimation step includes estimating the posture of the second hand when it shows the specific gesture based on the information and the posture of the specific part. A control method for an information processing device characterized by the following features.

[0104] [system] A storage device that stores information based on the first posture when, in the first captured image, the first hand included in the first captured image is in a first posture that indicates a specific gesture, The system includes an estimation device that, in a second image captured after the first image, even if the second hand included in the second image does not exhibit the specific gesture, if the second image includes a specific part of the second hand, estimates the posture of the second hand when exhibiting the specific gesture based on the information and the posture of the specific part. An information processing system characterized by the following:

[0105] [method] A storage step in which, in the first captured image, if the first hand included in the first captured image is in a first posture that indicates a specific gesture, information based on the first posture is stored in the storage unit; The system includes a display control step that controls the display unit to display a first image in which a virtual object with an orientation corresponding to the first posture is synthesized when the first hand is in the first posture in the first captured image, In the display control step, if a second image captured after the first image capture includes a specific part of the second hand, even if the second hand in the second image capture does not perform the specific gesture, the display unit is controlled to display a second image in which the virtual object is composited in an orientation corresponding to the posture of the specific part. A control method for an information processing device characterized by the following features.

[0106] [system] Display device and A storage device that stores information based on the first posture when, in the first captured image, the first hand included in the first captured image is in a first posture that indicates a specific gesture, The system includes a display control device that controls the display device to display a first image in which a virtual object with an orientation corresponding to the first posture is synthesized when the first hand is in the first posture in the first captured image, The display control device controls the display device to display a second image on the display device in which the virtual object is synthesized in an orientation corresponding to the posture of the second part of the second hand, if the second image captured after the first image capture image includes a specific part of the second hand, even if the second hand included in the second image capture image does not exhibit the specific gesture. An information processing system characterized by the following:

Claims

1. A storage means for storing information based on a first posture in a storage unit when, in a first captured image, the first hand included in the first captured image is in a first posture that indicates a specific gesture, In a second image captured after the first image, even if the second hand included in the second image is not shown because at least part of the fingers are hidden by a specific part of the second hand, the system includes estimation means for estimating the posture of the second hand as if it were performing the specific gesture in the second image, based on the information and the posture of the specific part of the second hand. An information processing device characterized by the following:

2. The storage means stores the first posture as information, If the second captured image includes the specific part of the second hand, even if the second hand does not exhibit the specific gesture, the estimation means estimates the posture of the second hand exhibiting the specific gesture based on the first posture and the posture of the specific part of the second hand. The information processing apparatus according to feature 1.

3. The storage means further stores the posture of the specific part of the first hand as information, If the second captured image includes the specific part of the second hand, even if the second hand does not exhibit the specific gesture, the estimation means estimates the posture of the second hand exhibiting the specific gesture based on the posture of the first hand, the posture of the specific part of the first hand, and the posture of the specific part of the second hand. The information processing apparatus according to feature 1.

4. The storage means stores, as information, the difference in position and orientation between the first posture and the posture of the specific part of the first hand. If the second captured image includes the specific part of the second hand, even if the second hand does not exhibit the specific gesture, the estimation means estimates the posture of the second hand exhibiting the specific gesture based on the difference and the posture of the specific part of the second hand. The information processing apparatus according to feature 1.

5. The system further includes determination means for determining whether a hand included in the captured image is performing the specific gesture, If the determination means determines that the first hand is performing the specific gesture, the storage means stores the information in the storage unit. Even if the determination means does not determine that the second hand is performing the specific gesture because at least part of the fingers of the second hand are obscured by a specific part of the second hand in the second captured image, if the second captured image includes the specific part of the second hand, the estimation means estimates the posture of the second hand performing the specific gesture based on the information and the posture of the specific part of the second hand. The information processing apparatus according to feature 1.

6. The determination means determines whether the hand performs the specific gesture based on the positions of a plurality of joint points, which are points that estimate the position of at least one of the joints of the hand and the fingertips. The information processing apparatus according to feature 5.

7. The determination means determines whether the hand performs the specific gesture by classifying it into a class. The information processing apparatus according to feature 5.

8. The system further includes detection means for detecting the positions of multiple joint points, which are points that estimate the position of at least one of the joints of the hand and the fingertips from the captured image. The storage means stores the information in the storage unit based on the positions of the plurality of joint points. The information processing apparatus according to feature 1.

9. The aforementioned specific area is a part of the hand that offers a high degree of detection accuracy. The information processing apparatus according to feature 1.

10. The aforementioned specific gesture is a pinch gesture, where the thumb and index finger are brought together, or a gripping gesture, where the hand is clenched. The information processing apparatus according to feature 1.

11. A storage means for storing information based on a first posture in a storage unit when, in a first captured image, the first hand included in the first captured image is in a first posture that indicates a specific gesture, The system includes a display control means that controls the display unit to display a first image in which a virtual object with an orientation corresponding to the first posture is synthesized when the first hand is in the first posture in the first captured image, The display control means controls the display unit to display a second image in which, in a second image captured after the first image, the second hand included in the second image is not shown by the specific gesture because at least part of the fingers are hidden by a specific part of the second hand, but the virtual object is composited in an orientation corresponding to the posture of the specific part of the second hand as if the second hand were performing the specific gesture in the second image. An information processing device characterized by the following:

12. The storage means stores, as information, the difference in position and orientation between the first posture and the posture of the specific part of the first hand. The display control means controls the display unit to display the second image, which is generated based on the difference and the posture of the specific part of the second hand. The information processing apparatus according to feature 11.

13. The second captured image further includes estimation means for estimating the posture of the second hand when performing the specific gesture, based on the information and the posture of the specific part of the second hand. The information processing apparatus according to feature 11.

14. The system further includes a generation means for generating a virtual object based on the second hand pose estimated by the estimation means. The information processing apparatus according to feature 13.

15. The system further includes image generation means that synthesizes the virtual object oriented to the first posture with the first captured image to generate the first image, and synthesizes the virtual object oriented to the posture of the specific part of the second hand with the second captured image to generate the second image. The information processing apparatus according to feature 11.

16. A program for causing a computer to function as one of the means of the information processing apparatus described in claim 1 or 11.

17. A storage step in which, in the first captured image, if the first hand included in the first captured image is in a first posture that indicates a specific gesture, information based on the first posture is stored in the storage unit, The method includes an estimation step of estimating the posture of the second hand as if it were performing the specific gesture in the second image, based on the information and the posture of the specific part of the second hand, even though the second hand included in the second image is not performing the specific gesture because at least part of the fingers are hidden by a specific part of the second hand, in a second image taken after the first image. A control method for an information processing device characterized by the following features.

18. In a first captured image, if the first hand included in the first captured image is in a first posture that indicates a specific gesture, a storage device stores information based on the first posture, The second image, captured after the first image, includes an estimation device that estimates the posture of the second hand as if it were performing the specific gesture in the second image, even if the second hand included in the second image is not shown because at least part of the fingers are hidden by a specific part of the second hand, based on the information and the posture of the specific part of the second hand. An information processing system characterized by the following:

19. A storage step in which, in the first captured image, if the first hand included in the first captured image is in a first posture that indicates a specific gesture, information based on the first posture is stored in the storage unit, The system includes a display control step that controls the display unit to display a first image in which a virtual object with an orientation corresponding to the first posture is synthesized when the first hand is in the first posture in the first captured image, In the display control step, the system controls the display unit to display a second image in which the virtual object is synthesized in a orientation corresponding to the posture of the specific part of the second hand as if the second hand were performing the specific gesture in the second image, even though the second hand included in the second image is not shown by having at least a portion of its fingers hidden by a specific part of the second hand, in the second image. A control method for an information processing device characterized by the following features.

20. Display device and In a first captured image, if the first hand included in the first captured image is in a first posture that indicates a specific gesture, a storage device stores information based on the first posture, The system includes a display control device that controls the display device to display a first image on the display device in which a virtual object with an orientation corresponding to the first posture is synthesized when the first hand is in the first posture in the first captured image, The display control device controls the display device to display a second image in which, in a second image captured after the first image, the second hand included in the second image is not shown by the second hand because at least part of the fingers are hidden by a specific part of the second hand, but the virtual object is composited in an orientation corresponding to the posture of the specific part of the second hand as if the second hand were performing the specific gesture in the second image. An information processing system characterized by the following:

Citation Information

Patent Citations

  • Selective hand occlusion over virtual projections onto physical surfaces using skeletal tracking

    EP2691938B1

  • JP71048A

  • Learning-based estimation of hand and finger pose

    US20130236089A1

  • Tracking hand / body pose

    US20170116471A1

  • Apparatus and method for estimating hand position utilizing head mounted color depth camera, and bare hand interaction system using same

    US20170140552A1