Information processing device, information processing system, information processing program, and information processing method
The information processing device addresses the mismatch between perceived and actual hand poses by using image classification and joint point detection to enhance user interaction in mixed and augmented reality systems, improving operational comfort.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies struggle to match the user's perceived hand pose with the actual interaction with virtual objects, leading to discrepancies and discomfort in mixed reality and augmented reality systems.
An information processing device that uses an acquisition means to capture images, an estimation means to classify hand shapes, and a detection means to identify joint points, determining the user's hand pose accurately by combining class classification with joint point analysis to align hand movements with virtual object interactions.
This approach reduces the discrepancy between the user's perceived hand pose and the actual interaction, enhancing the sense of operation and improving user comfort in mixed reality and augmented reality systems.
Smart Images

Figure 2026040927000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device that determines whether a hand pose is being performed by a user. [Background technology]
[0002] Mixed reality (MR) and augmented reality (AR) are technologies that combine the real world with computer graphics (CG) generated by a computer in real time. Mixed reality and augmented reality present a user with a composite image of the real world and CG using a device worn on the user's head called a head-mounted display (HMD), enabling interaction between the user and the CG, thereby providing an immersive experience. Hand gesture manipulation is one of the operation methods used by a user to interact with CG. To realize hand gesture manipulation, a technology is used that determines the hand pose performed by the user using a machine learning model. For example, Patent Document 1 discloses a technology that captures an image of a user's hand with a camera mounted on an HMD, analyzes the captured image using a machine learning model, and classifies the image into predefined hand pose classes to determine the hand pose. Using such a technology, for example, a user can grasp a virtual object by performing a hand pose to grasp the virtual object near the virtual object to operate the virtual object, and move the virtual object in accordance with the user's hand movement. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2019-71048 A Summary of the Invention [Problem to be solved by the invention]
[0004] However, with the technology disclosed in Patent Document 1, it may be difficult to match a specific hand pose with the operation feel of each user. For example, in a hand pose known as a pinch, in which a virtual object is grasped with the thumb and index finger, the user needs to make a pose that resembles grasping space in real space. Therefore, the distance between the thumb and index finger that makes the user feel like they have grasped the virtual object varies from person to person. For example, as shown in FIG. 7, when grasping a virtual object 701 with a pinch hand pose, some people are expected to make a pinch hand pose 603 in which the fingers touch near the virtual object 701, as shown in FIG. 7(a). In addition, other people are expected to make a pinch hand pose 604 in which the thumb and index finger grasp the handle portion of the virtual object 701 with the thumb and index finger, as shown in FIG. 7(b). Therefore, for a person attempting to grasp a virtual object with the hand shape of FIG. 7(a), if the operation feel is such that the virtual object can be grasped with the hand shape of FIG. 7(b), the virtual object may be determined to be grasped earlier than expected, which may increase the number of erroneous operations. In this way, the user may feel uncomfortable interacting with the virtual object, which may reduce the sense of operation.
[0005] Therefore, the present invention aims to reduce the discrepancy between the hand posture that the user feels is a specific hand pose and the hand pose that actually results in interaction with a virtual object, thereby improving the user's sense of operation. [Means for solving the problem]
[0006] In order to achieve the above object, the information processing device of the present invention is characterized by having an acquisition means for acquiring an image, an estimation means for estimating whether the user's hand in the image acquired by the acquisition means has a specific shape based on class classification, a detection means for detecting from the image the positions of a plurality of joint points which are points that estimate the positions of at least one of the joints and fingertips of the user's hand in the image, and a determination means for determining whether the user's hand in the image has the specific shape based on the positions of the plurality of joint points detected by the detection means when the estimation means estimates that the user's hand in the image has the specific shape. [Effects of the Invention]
[0007] According to the present invention, it is possible to provide a technology that reduces the discrepancy between the hand posture that a user feels as if they have performed a specific hand pose and the hand pose that actually results in interaction with a virtual object, thereby improving the user's sense of operation. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an information processing system according to a first embodiment. [Figure 2] 1 is a diagram illustrating the internal configuration of an HMD according to a first embodiment. [Figure 3] 10 is a flowchart of a hand pose determination process according to the first embodiment. [Figure 4] 10 is a flowchart of a hand pose determination process according to the second embodiment. [Figure 5] 10 is a flowchart of a virtual object operation process according to the first embodiment. [Figure 6] 1A to 1C are diagrams illustrating examples of hand poses according to the first embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of a pinch pose according to the first embodiment. [Figure 8] 3A to 3C are diagrams illustrating an example of a composite image to be drawn according to the first embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Each embodiment will be described below with reference to the drawings. The same or equivalent components, members, and processes shown in each drawing will be assigned the same reference numerals, and duplicate descriptions will be omitted where appropriate. In addition, some of the components, members, and processes will be omitted in each drawing.
[0010] (Embodiment 1) <System configuration> An information processing system 1 according to the first embodiment will be described with reference to Fig. 1. The information processing system 1 includes a head-mounted display (HMD) 100 and a PC (personal computer) 110.
[0011] The HMD 100 is a head-mounted display device (electronic device) that can be worn on the user's head. The HMD 100 includes a camera for capturing an image of the area in front of the user and a display for displaying the image to the user. The display of the HMD 100 displays a composite image that combines an image captured by the HMD 100 of the area in front of the user with content such as CG (computer graphics) in a format that corresponds to the posture of the HMD 100. This allows the user to experience virtual reality with their eyes. The system 1 also has a function for detecting the user's hands from the image captured by the HMD 100 and acquiring information related to the position and orientation of the hands as the posture of the hands, thereby affecting virtual objects with the movements of the hands. This allows the user to intuitively operate virtual objects using their hands.
[0012] The PC 110 controls the HMD 100. The PC 110 is connected to the HMD 100 by wire such as a USB cable or wirelessly such as Bluetooth (registered trademark) or Wi-Fi (Wireless Fidelity) (registered trademark). The PC 100 and the HMD 101 can communicate with each other by wireless or wired communication and transmit and receive images and other necessary information. The PC 110 generates a composite image by synthesizing the image captured by the HMD 100 and the CG generated by the PC 110, and transmits the composite image to the HMD 100. Here, although a PC is described as an example of the information processing device, the information processing device is not limited to this. For example, the information processing device may be a smartphone or a tablet terminal, and each component of the PC 110 may be possessed by the HMD 100.
[0013] <Internal Structure of HMD> Referring to FIG. 2, the internal structure of the HMD 100 will be described. The HMD 100 includes an HMD control unit 201, an imaging unit 202, an image display unit 203, an attitude sensor unit 204, a nonvolatile memory 205, and a working memory 206.
[0014] The HMD control unit 201 controls each component of the HMD 100. The HMD control unit 201 includes at least one CPU that executes a program stored in the nonvolatile memory 205 and at least one other circuit. When the HMD control unit 201 acquires a composite image (an image obtained by synthesizing the captured image of the space in front of the user captured by the imaging unit 202 and the CG) from the PC 110, the HMD control unit 201 displays the composite image on the image display unit 203. Instead of the HMD control unit 201 controlling the entire device, a plurality of hardware may share the processing to control the entire device.
[0015] The imaging unit 202 includes two cameras (imaging devices). The two cameras are disposed near the positions of the user's left and right eyes when the HMD 100 is worn on the user's head. Therefore, the two cameras can capture a space similar to the space seen by the user wearing the HMD 100. The images captured by the imaging unit 202 are output to the HMD control unit 201, which transmits the images from the imaging unit 202 to the PC 110. As described below, the PC 110 combines the captured images transmitted from the HMD 100 with CG to generate a composite image. Furthermore, the imaging unit 202 simultaneously captures a first image having a parallax with respect to each other and a second image different from the first image using the two cameras. Therefore, information on the distance from the HMD 100 to a subject (distance information) can be acquired using the images from the two cameras in the imaging unit 202. The imaging unit 202 may also capture and output a video.
[0016] As described below, when a composite image is transmitted from the PC 110, the image display unit 203 displays the composite image transmitted from the PC 110. The image display unit 203 has a display such as a liquid crystal panel or an organic EL panel. When the user is wearing the HMD 100, a display such as an organic EL panel is disposed in front of each of the user's eyes. Note that a device using a semi-transparent half mirror may also be used for the image display unit 203. In this case, for example, the image display unit 203 may use a technology generally called AR (Augmented Reality) to display an image so that CG appears to be directly superimposed on the real space visible through the half mirror. Furthermore, the image display unit 203 may use a technology generally called VR (Virtual Reality) to display an image of a completely virtual space without using captured images.
[0017] The attitude sensor unit 204 acquires the attitude (and position) information of the HMD 100. Then, the attitude sensor unit 204 acquires the attitude information of the user (the user wearing the HMD 100) corresponding to the attitude (and position) of the HMD 100. The attitude sensor unit 204 has an inertial measurement unit (IMU; Inertial Measurement Unit) composed of an acceleration sensor, an angular acceleration sensor, and a geomagnetic sensor. When the user is wearing the HMD 100, the attitude sensor unit 204 acquires the information (attitude information) of the user's attitude. The HMD control unit 201 outputs the information (attitude information) of the user's attitude detected by the attitude sensor 204 to the PC 110.
[0018] The non-volatile memory 205 is an electrically erasable and recordable non-volatile memory, and stores programs and the like executed by the HMD control unit 201.
[0019] The volatile memory 206 is used as a buffer memory that temporarily holds the image data captured by the imaging unit 202, an image display memory for the image display unit 203, a working area of the HMD control unit 201, and the like.
[0020] <Internal Configuration of the PC> Referring to FIG. 2, the internal configuration of the PC 110 will be described. The PC 110 has a control unit 211, a non-volatile memory 212, and a working memory 213.
[0021] The control unit 211 is a CPU composed of at least one processor or circuit. The control unit 211 realizes each process of the flowchart described later by executing the program stored in the non-volatile memory 212. Instead of the control unit 211 controlling the entire device, a plurality of hardware may share the processing to control the entire device. The control unit 211 receives from the HMD 100 the image (captured image) acquired by the imaging unit 202 and the attitude information acquired by the attitude sensor unit 204. The control unit 211 synthesizes the captured image and an arbitrary CG based on the received information to generate a synthesized image. The control unit 211 transmits the synthesized image to the HMD control unit 201 in the HMD 100.
[0022] The nonvolatile memory 212 is an electrically erasable and recordable nonvolatile memory, and stores information such as programs to be described later and CG executed by the control unit 211. The control unit 211 can switch the CG read from the nonvolatile memory 212 (i.e., the CG used to generate a composite image).
[0023] The working memory 213 is a storage unit used as a working area for the control unit 211, such as a buffer memory that temporarily stores image data captured by the imaging unit 202.
[0024] <Flow explaining the process of determining hand pose> The process of determining a hand pose will be described with reference to the flowchart of Fig. 3. The process of the flowchart of Fig. 3 is executed from the timing when the user starts an application on the HMD, and is executed each time the imaging unit 202 acquires an image (each time an image is captured). An application is, for example, an application that the user selects on the home screen (home space) after starting the HMD, and includes an app that allows the user to interact with a virtual object using hand gestures. Note that the timing when this flowchart is executed is not limited to the timing when the user starts an application on the HMD. For example, it may be the timing when the user starts the HMD or the timing when a virtual object is displayed in the mixed reality space.
[0025] In step S301, the control unit 211 acquires an image captured by the imaging unit 202, and the process proceeds to step S302.
[0026] In step S302, the control unit 211 estimates the hand pose by classifying the image captured by the imaging unit 202. Examples of hand pose classes include a grip pose, a pinch pose, and no pose. A grip pose is a hand that makes a fist, and a pinch pose is a hand that pinches something by bringing the tips of the thumb and index finger close together.
[0027] In the embodiment, the grip pose and pinch pose are used as hand poses when grasping a virtual object. "No pose" indicates a hand shape that does not fit into the grip pose or pinch pose. For example, a class classification process is performed using a deep learning model that has learned these hand poses as classes. Note that the pinch pose is not limited to bringing the tips of the thumb and index finger closer together, but may also be an action that changes multiple fingers from a separated state to a state where they are brought closer together, such as bringing the thumb and middle finger closer together, or bringing the thumb, index finger, and middle finger closer together.
[0028] In a deep learning model, for example, a captured image is input into a CNN (convolutional neural network). The CNN outputs features used to identify the type of subject or the type of scene. The features output from the CNN are then used to identify the type of subject or scene. Deep learning models such as YOLO, MobileNet, VGG16, and SSD are examples of detectors and classifiers for target objects or people.
[0029] The control unit 211 uses a deep learning model to estimate which hand pose class the subject of an image captured by the imaging unit 202 belongs to, and records the hand pose of the assigned class in the working memory 213 as the hand pose being performed by the user. That is, when a captured image is input to the deep learning model, an area in which the subject of the captured image includes a hand is detected. Furthermore, if an area in which the subject of the captured image includes a hand is detected, it estimates (classifies) whether the shape of the subject is a specific shape. Here, in the class classification process, if a captured image that does not include a hand is input, nothing is detected from the input image. An example of such a hand pose will be described with reference to FIG. 6.
[0030] FIG. 6(a) is an example of a grip pose. Hand 601 is clenched with each finger visible. FIG. 6(b) is an example of a grip pose. Hand 602 has only the thumb and index finger visible, with the other fingers hidden. FIG. 6(c) is an example of a pinch pose. Hand 603 has the tips of the thumb and index finger touching together. FIG. 6(d) is an example of a pinch pose. Hand 604 has the tips of the thumb and index finger not touching together.
[0031] In step S303, the control unit 211 estimates the three-dimensional positions of the joint points of the hand based on the image captured by the imaging unit 202. Here, the "joint points" include the joint points of each finger, points indicating a predetermined position of the wrist, and points at the tips of each finger. In step S303, the positions of at least one of the joints and fingertips of the user's hand are estimated as joint points. For example, a deep learning model is used to estimate the three-dimensional positions of the joint points of the hand.
[0032] The control unit 211 estimates a three-dimensional position by combining the two-dimensional position of each joint point in the image space of the image captured by the imaging unit 202 with the depth position, with the depth of the wrist as the reference value 0, along the axis in the direction perpendicular to the camera of the imaging unit 202. The estimated value is recorded in the working memory 213.
[0033] An example of such estimation of the three-dimensional positions of hand joint points will be described with reference to Fig. 6. Fig. 6(e) shows an example of joint points when the three-dimensional positions of the joint points of the hand 603 in Fig. 6(c) are estimated. A total of 21 joint points 605 are estimated, including four points for each finger and a point on the wrist. Note that the deep learning model used to estimate the hand pose and the deep learning model used to estimate the three-dimensional positions of the hand joint points may be the same model or different models.
[0034] In step S304, the control unit 211 determines whether the hand pose determined as a result of the classification in step S302 is a grip pose or a pinch pose. If the hand pose being performed by the user is a grip pose or a pinch pose, the process proceeds to step S305; otherwise, the process proceeds to step S307. For example, if the hand is not shown in the captured image, or if the hand is shown in the captured image but it is estimated that the hand is not a specific hand pose (hand gesture) such as a grip pose or a pinch pose, the process proceeds to step S307.
[0035] In step S305, the control unit 211 determines whether the pose is a grip pose or a pinch pose using the joint points estimated in step S303. If the user's hand pose is estimated to be a grip pose in step S302, it determines whether the grip pose is being performed based on the angle of the finger joint points. The angle of the finger joint points indicates the angle formed by three adjacent joint points. If the angle of the finger joint points is smaller than a predetermined angle, it determines that the grip pose is being performed, and sets the determination result to TRUE.
[0036] The angle of the finger joint point used for the judgment may be any joint of the finger. Also, only one joint may be used for the judgment, or multiple joints may be used for the judgment. Any one of multiple fingers may be used for the judgment, or multiple fingers may be used.
[0037] If the user's hand pose is estimated to be a pinch pose in step S302, a determination is made as to whether a pinch pose is being made based on the distance between the joint points of the thumb and index finger. If this distance is smaller than a predetermined distance (threshold), it is determined that a pinch pose has been made, and the determination result is set to TRUE. If the hand pose determination result based on the joint points is TRUE, proceed to step S306; if not, proceed to step S307.
[0038] Note that when determining the hand pose based on joint points after the hand pose has been estimated by class classification using a deep learning model in this way, the hand pose may be determined based only on characteristic parts according to the estimated hand pose. For example, when determining whether a pinch pose is being made based on the distance between the joint points at the tips of the thumb and index finger, only the joint points at the tips of the thumb and index finger may be extracted from the joint points, and then a determination may be made as to whether a pinch pose is being made. In this way, by first estimating the hand pose by class classification, the processing load for determining the hand pose based on joint points can be reduced.
[0039] In step S306, if the determination result of the grip pose in step S305 is TRUE, the control unit 211 determines that a grip pose is being performed and records TRUE in the working memory 213. Also, if the determination result of the pinch pose in step S305 is TRUE, the control unit 211 determines that a pinch pose is being performed and records TRUE in the working memory 213, and proceeds to step S308.
[0040] In step S307, the control unit 211 determines that neither a grip pose nor a pinch pose is being performed, and records FALSE in the working memory 213, and the process proceeds to step S308.
[0041] In step S308, the control unit 211 performs virtual object operation processing based on the determination result recorded in the working memory 213 in step S306 or step S307, and proceeds to step S309. The virtual object operation processing performed in step S308 will be described later with reference to FIG.
[0042] In step S309, the control unit 211 determines whether or not the application has been terminated. If the application has been terminated, the process proceeds to termination, and if not, the process repeats from step S301.
[0043] 3, steps S302 and S303 may be performed in reverse order or simultaneously. For example, a single deep learning model may be used to estimate a hand pose based on class classification and determine the hand pose based on joint points. Furthermore, steps S302 and S303 may be performed on both the right-eye image and the left-eye image, or on only one of the right-eye image and the left-eye image.
[0044] According to the first embodiment, the control unit 211 estimates the hand pose by classifying and determines the hand pose by the joint points, thereby determining the hand pose by distinguishing even the finer details of the fingers, thereby improving the user's sense of control over the virtual object.
[0045] <Flow of virtual object operation processing> The virtual object operation process performed in step S308 of FIG. 3 will be described with reference to the flowchart of FIG.
[0046] In step S501, control unit 211 determines whether a virtual object exists near the hand. If control unit 211 determines that a virtual object exists near the hand, the process proceeds to step S502, and if control unit 211 does not determine that a virtual object exists near the hand, the process ends the flow of the virtual object operation process.
[0047] In step S502, control unit 211 determines whether the determination result of the user's hand pose is set to TRUE. If control unit 211 determines that the determination result of the hand pose is set to TRUE, the process proceeds to step S502, and if control unit 211 determines that the determination result of the hand pose is set to FALSE, the flow of the virtual object operation process ends.
[0048] In step S503, the control unit 211 moves the virtual object based on the position of the hand, superimposes it on the captured image, and ends the flow of the virtual object operation process.
[0049] In this way, when a virtual object is present near the hand and the hand pose determination result is set to TRUE, it is assumed that the user has selected and moved the virtual object, and the virtual object is drawn so as to follow the position of the hand.
[0050] Furthermore, if it is not determined that the virtual object is present in the vicinity of the hand, or if it is determined that the hand pose determination result is set to FALSE, the virtual object is not moved based on the position of the hand.
[0051] <Explanation of the scene where you stop manipulating a virtual object> An example of a scene in which the operation of a virtual object is stopped will be described with reference to Fig. 8. In this scene, a captured image 801 at a first time point shows a hand 603 making a pinch pose, and a virtual object 701 is placed in accordance with the position of the hand. At this point, the user has selected (grabbed) the virtual object. Furthermore, a composite image 802 generated after the captured image 801 shows a hand 803 not making a pinch pose, and the virtual object 701 is not selected.
[0052] Here, assume that in an image captured at a first time point, a user makes a pinch pose to select a virtual object, and in an image captured at a second time point, which is acquired after the image captured at the first time point, the user stops making the pinch pose. In this case, in the flow when the image captured at the first time point is acquired, the process of step S503 is executed, and a composite image 801 is generated. In addition, in the flow when the image captured at a second time point is acquired, it is determined that the hand pose determination result is set to FALSE, and the virtual object operation process is terminated with the position of the virtual object in the flow at the first time point remaining, and a composite image 802 is generated. In other words, at the second time point, an image 802 in the mixed reality space is generated while maintaining the position of the virtual object drawn at the first time point. Note that if the position of the angle of view of the captured image changes due to the user moving their head, for example, the position of the virtual object is adjusted in accordance with the change.
[0053] (Embodiment 2) Because hand pose estimation using class classification includes erroneous estimation, it may be estimated that a hand pose is not being performed even when it is actually being performed. After a virtual object is grasped, the hand pose at the time of grasping is maintained until the virtual object is released. Therefore, this erroneous estimation may lead to a determination that a hand pose is not being performed, causing the virtual object to be released and resulting in a decrease in operability. In the second embodiment, in order to reduce such a decrease in operability, a case will be described in which only joint points are used to determine the hand pose after a virtual object is grasped.
[0054] The processing after a hand pose is performed will be described with reference to Fig. 4. The description of the same processing as in Fig. 3 of the first embodiment will be omitted, and steps S404, S409, and S410 will be described.
[0055] Steps S401 to S403 are the same as steps S301 to S303 in Fig. 3, and therefore description thereof will be omitted. After performing the process of step S403, control unit 211 proceeds to step S404.
[0056] In step S404, the control unit 211 determines whether or not a grip pose or a pinch pose is being performed based on the setting information of these poses recorded in the working memory 213. If a grip pose or a pinch pose is being performed, the process proceeds to step S409; otherwise, the process proceeds to step S405.
[0057] Steps S405 to S408 are the same as steps S305 to S309 in Fig. 3, and therefore description thereof will be omitted. After performing the process of step S407, control unit 211 proceeds to step S411. Also, after performing the process of step S408, control unit 211 proceeds to step S411.
[0058] In step S409, the control unit 211 performs the same process as S305 in Fig. 3. If the determination result of the grip pose or pinch pose is TRUE, the process proceeds to step S411, and if not, the process proceeds to step S410.
[0059] In step S410, the control unit 211 records FALSE in the working memory 213, since neither a grip pose nor a pinch pose has been performed, and the process proceeds to step S411.
[0060] Steps S411 and S412 are the same as steps S308 and S309 in FIG. 3, and therefore a description thereof will be omitted.
[0061] In this way, in embodiment 2, while the user's hand is determined to have a specific shape, it is determined from the estimation result of the hand joint points, regardless of the estimation result of the hand pose by class classification, whether the user's hand has a specific shape. Note that it may be determined from the estimation result of the hand joint points, regardless of the estimation result of the hand pose by class classification, whether the user's hand has a specific shape until it is no longer determined that the user's hand has a specific shape.
[0062] In the flow of FIG. 4 , the hand pose is estimated by class classification regardless of whether the hand pose estimation result by class classification is also used to determine whether the user's hand has a specific shape. However, if the hand pose estimation result by class classification is not also used to determine whether the user's hand has a specific shape, the hand pose estimation by class classification does not need to be performed. That is, while the user's hand is determined to have a specific shape, it may be determined whether the user's hand has a specific shape from the hand joint point estimation result without performing the hand pose estimation by class classification. Furthermore, it may be determined whether the user's hand has a specific shape from the hand joint point estimation result without performing the hand pose estimation by class classification until the user's hand is no longer determined to have a specific shape.
[0063] As described above, according to the second embodiment, when a grip pose or a pinch pose is performed, the control unit 211 determines the hand pose only from the joint points, thereby preventing a decrease in operability due to an erroneous estimation of class classification. As a result, the user's sense of operability of the virtual object can be improved.
[0064] (Other embodiments) The present invention can also be realized by executing the following process: software (program) that realizes the functions of the above-described embodiments is supplied to a system or device via a network or various storage media, and the computer (or control unit, MPU, etc.) of the system or device reads and executes the program code. In this case, the program and the storage medium storing the program constitute the present invention.
[0065] Although the present invention has been described in detail above based on preferred embodiments thereof, the present invention is not limited to these specific embodiments, and various forms within the scope of the gist of the present invention are also included in the present invention. Parts of the above-described embodiments may be combined as appropriate.
[0066] Note that each functional unit in each of the above embodiments (variations) may or may not be individual hardware. The functions of two or more functional units may be realized by common hardware. Each of multiple functions of one functional unit may be realized by individual hardware. Two or more functions of one functional unit may be realized by common hardware. Furthermore, each functional unit may or may not be realized by hardware such as an ASIC, FPGA, or DSP. For example, an apparatus may have a processor and a memory (storage medium) in which a control program is stored. Then, the functions of at least some of the functional units of the apparatus may be realized by the processor reading and executing the control program from the memory.
[0067] The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.
[0068] In addition, in each of the examples described above, the term "processor" refers to a processor in a broad sense, and includes general-purpose processors (e.g., CPUs) and dedicated processors (e.g., GPUs, ASICs, FPGAs, and programmable logic devices, etc.).
[0069] The disclosure of this embodiment includes the following configuration, method, and program.
[0070] [Configuration 1] an acquisition means for acquiring a captured image; an estimation means for estimating whether the user's hand in the captured image acquired by the acquisition means has a specific shape based on the classification; a detection means for detecting, from the captured image, positions of a plurality of joint points, which are points that estimate the positions of at least one of the joints and fingertips of the user's hand in the captured image; and a determination means for determining, when the estimation means has estimated that the user's hand in the captured image has the specific shape, whether the user's hand in the captured image has the specific shape based on the positions of the plurality of joint points detected by the detection means. 1. An information processing device comprising:
[0071] [Configuration 2] The determining means determines that the user's hand has the specific shape when the hand shape formed by the positions of the plurality of joint points is the specific shape. 2. The information processing device according to configuration 1,
[0072] [Configuration 3] When the hand shape formed by the positions of the plurality of joint points is not the specific shape, the determining means does not determine that the user's hand has the specific shape. 3. The information processing device according to configuration 2.
[0073] [Configuration 4] The present invention further includes a control unit that, when the determination unit determines that the user's hand has the specific shape, controls the device to perform processing according to the specific shape. 4. The information processing device according to configuration 2 or 3.
[0074] [Configuration 5] The determining means does not determine whether the user's hand has the specific shape when the estimating means has not estimated that the user's hand has the specific shape. 5. The information processing device according to any one of configurations 1 to 4.
[0075] [Configuration 6] The determining means determines whether the user's hand has the specific shape based on angles formed by three adjacent joint points among the plurality of joint points. 6. The information processing device according to any one of configurations 1 to 5.
[0076] [Configuration 7] When the estimation means estimates that the user's hand has the specific shape and the determination means determines that the user's hand has the specific shape, the determination means determines whether the user's hand has the specific shape regardless of the estimation result of the estimation means until the determination means no longer determines that the user's hand has the specific shape. 7. The information processing device according to any one of configurations 1 to 6.
[0077] [Configuration 8] When the estimation means estimates that the user's hand has the specific shape and the determination means determines that the user's hand has the specific shape, the estimation means does not estimate whether the user's hand has the specific shape until the determination means no longer determines that the user's hand has the specific shape. 8. The information processing device according to any one of configurations 1 to 7.
[0078] [Configuration 9] The specific shape is a hand shape in which the distance between the thumb and index finger is smaller than a threshold value. 9. The information processing device according to any one of configurations 1 to 8.
[0079] [Configuration 10] The specific shape is a shape when a hand grasps a virtual object, or a shape when a hand grasps a virtual object. 10. The information processing device according to any one of configurations 1 to 9.
[0080] [Configuration 11] The acquisition means acquires a first image having a parallax with respect to each other and a second image different from the first image. 11. The information processing device according to any one of configurations 1 to 10.
[0081] [Configuration 12] The estimation means estimates whether the user's hand shown in the first image and the second image acquired by the acquisition means has a specific shape. 12. The information processing device according to configuration 11.
[0082] [Control method] an acquisition step of acquiring a captured image; an estimation step of estimating whether the user's hand shown in the captured image acquired in the acquisition step has a specific shape based on the class classification; a detecting step of detecting positions of a plurality of joint points, which are points that estimate the positions of at least one of the joints and fingertips of the user's hand, from the captured image acquired by the acquiring step; a determining step of determining whether the user's hand has the specific shape based on the positions of the plurality of joint points detected in the detecting step, when the user's hand is estimated to have the specific shape in the estimating step. 2. A method for controlling an information processing apparatus comprising:
[0083] [program] 13. A program for causing a computer to function as each of the means of the information processing device according to any one of configurations 1 to 12.
[0084] [system] an acquisition device that acquires a captured image; an estimation device that estimates whether a user's hand shown in a captured image acquired by the acquisition device has a specific shape based on the classification; a detection device that detects positions of a plurality of joint points, which are points that estimate the positions of at least one of the joints and fingertips of the user's hand, from the captured image acquired by the acquisition device; a determination device that, when the estimation device estimates that the user's hand has the specific shape, determines whether the user's hand has the specific shape based on the positions of the plurality of joint points detected by the detection device. An information processing system comprising:
Claims
1. an acquisition means for acquiring a captured image; an estimation means for estimating whether the user's hand in the captured image acquired by the acquisition means has a specific shape based on the classification; a detection means for detecting, from the captured image, positions of a plurality of joint points, which are points that estimate the positions of at least one of the joints and fingertips of the user's hand in the captured image; and a determination means for determining, when the estimation means has estimated that the user's hand in the captured image has the specific shape, whether the user's hand in the captured image has the specific shape based on the positions of the plurality of joint points detected by the detection means.
1. An information processing device comprising:
2. The determining means determines that the user's hand has the specific shape when the hand shape formed by the positions of the plurality of joint points is the specific shape.
2. The information processing apparatus according to claim 1, wherein:
3. When the hand shape formed by the positions of the plurality of joint points is not the specific shape, the determining means does not determine that the user's hand has the specific shape.
3. The information processing apparatus according to claim 2, wherein:
4. The present invention further includes a control unit that, when the determination unit determines that the user's hand has the specific shape, controls the device to perform processing according to the specific shape.
3. The information processing apparatus according to claim 2, wherein:
5. The determining means does not determine whether the user's hand has the specific shape when the estimating means has not estimated that the user's hand has the specific shape.
2. The information processing apparatus according to claim 1, wherein:
6. The determining means determines whether the user's hand has the specific shape based on angles formed by three adjacent joint points among the plurality of joint points.
2. The information processing apparatus according to claim 1, wherein:
7. When the estimation means estimates that the user's hand has the specific shape and the determination means determines that the user's hand has the specific shape, the determination means determines whether the user's hand has the specific shape regardless of the estimation result of the estimation means until the determination means no longer determines that the user's hand has the specific shape.
2. The information processing apparatus according to claim 1, wherein:
8. When the estimation means estimates that the user's hand has the specific shape and the determination means determines that the user's hand has the specific shape, the estimation means does not estimate whether the user's hand has the specific shape until the determination means no longer determines that the user's hand has the specific shape.
2. The information processing apparatus according to claim 1, wherein:
9. The specific shape is a hand shape in which the distance between the thumb and index finger is smaller than a threshold value.
2. The information processing apparatus according to claim 1, wherein:
10. The specific shape is a shape when a hand grasps a virtual object, or a shape when a hand grasps a virtual object.
2. The information processing apparatus according to claim 1, wherein:
11. The acquisition means acquires a first image having a parallax with respect to each other and a second image different from the first image.
2. The information processing apparatus according to claim 1, wherein:
12. The estimation means estimates whether the user's hand shown in the first image and the second image acquired by the acquisition means has a specific shape.
12. The information processing apparatus according to claim 11,
13. an acquisition step of acquiring a captured image; an estimation step of estimating whether the user's hand shown in the captured image acquired in the acquisition step has a specific shape based on the class classification; a detecting step of detecting positions of a plurality of joint points, which are points that estimate the positions of at least one of the joints and fingertips of the user's hand, from the captured image acquired by the acquiring step; a determining step of determining whether the user's hand has the specific shape based on the positions of the plurality of joint points detected in the detecting step, when the user's hand is estimated to have the specific shape in the estimating step.
2. A method for controlling an information processing apparatus comprising:
14. A program for causing a computer to function as each of the means of the information processing apparatus according to claim 1.
15. an acquisition device that acquires a captured image; an estimation device that estimates whether a user's hand shown in a captured image acquired by the acquisition device has a specific shape based on the classification; a detection device that detects positions of a plurality of joint points, which are points that estimate the positions of at least one of the joints and fingertips of the user's hand, from the captured image acquired by the acquisition device; a determination device that, when the estimation device estimates that the user's hand has the specific shape, determines whether the user's hand has the specific shape based on the positions of the plurality of joint points detected by the detection device. An information processing system comprising:
Citation Information
Patent Citations
Hand posture estimation method and system based on visual and inertial information fusion
CN113221726A
Three-dimensional gesture tracking method based on RGB camera
CN115810219A
Information processing device and information processing method
WO2022137901A1
JP71048A