Information processing system, information processing method and program

A smartwatch with an imaging unit and deep learning network estimates three-dimensional finger postures from the back of the hand, addressing mobility and comfort issues in existing systems, enabling accurate and comfortable wearable interaction.

JP7790697B2Active Publication Date: 2025-12-23INSTITUTE OF SCIENCE TOKYO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021143847
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-08
Filing Date
2021-09-03
Publication Date
2025-12-23
Estimated Expiration
2041-09-03

AI Technical Summary

Technical Problem

Existing information processing systems, such as those described in Patent Document 1, require a camera to capture the movement of a glove-type input device, limiting user mobility and are cumbersome for constant wear due to the weight of the HMD.

Method used

A smartwatch equipped with an imaging unit captures images of the back of the user's hand, using a deep learning network to estimate the three-dimensional posture of the fingers based on multiple frames, allowing for accurate estimation of finger movements without the need for constant camera alignment.

Benefits of technology

Enables accurate estimation of three-dimensional finger postures for natural and comfortable wear, eliminating the need for camera alignment and reducing physical interference, and allowing seamless integration into wearable devices like smartwatches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007790697000004
    Figure 0007790697000004
  • Figure 0007790697000005
    Figure 0007790697000005
  • Figure 0007790697000006
    Figure 0007790697000006
Patent Text Reader

Abstract

To provide an information processing device, information processing method, and program, which allow for estimating three-dimensional posture of user's fingers.SOLUTION: A smart watch 100, which is an information processing system, has a control unit. The control unit is configured to extract an image of an area of the back of a user's hand from a captured image containing the back of the user's hand, and estimate three-dimensional posture of user's fingers based on the extracted image.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, an information processing method, and a program. [Background technology]

[0002] Patent Document 1 discloses a system that includes a head-mounted display (HMD) and a glove-type input device, and inputs information in accordance with the movement of the glove-type input device. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-207949 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the system in Patent Document 1 requires a camera to capture the movement of a glove-type input device, so the user must move their hand within the camera's capture range. Also, the HMD is heavy and unsuitable for constant wear. [Means for solving the problem]

[0005] According to one aspect of the present invention, there is provided an information processing system including a control unit that extracts an image of the back of a user's hand from a captured image including the back of the hand, and estimates a three-dimensional posture of the user's fingers based on the extracted image. [Brief explanation of the drawings]

[0006] [Figure 1] FIG. 1 is a diagram illustrating an example of an overview of a smart watch. [Figure 2] FIG. 2 is a diagram illustrating an example of the hardware configuration of a smartwatch. [Figure 3] FIG. 3 is a diagram showing an example of a captured image captured by the imaging section. [Figure 4] FIG. 4 is a diagram illustrating an example of the functional configuration of a smartwatch. [Figure 5] FIG. 5 is a diagram illustrating an example of preprocessing. [Figure 6] FIG. 6 is a diagram illustrating an example of information processing by a smartwatch. [Figure 7] FIG. 7 is a diagram illustrating an example of a 3D hand model. [Figure 8] FIG. 8 shows the results of estimation of static gestures (ASL0-9). [Figure 9] FIG. 9 shows the results of estimation of dynamic gestures (tapping 0-4). [Figure 10] FIG. 10 shows an example of a data set of different grasp types. [Figure 11] Figure 11 shows an example of a confusion matrix and heatmap for 50 tests of each grasp type. [Figure 12] FIG. 12 is an activity diagram showing an example of information processing in a smartwatch. [Figure 13] FIG. 13 is a diagram illustrating an example of an information processing system according to the first modification. DETAILED DESCRIPTION OF THE INVENTION

[0007] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention will be described below with reference to the accompanying drawings. Various features shown in the following embodiments can be combined with each other.

[0008] In this specification, the term "unit" may include, for example, a combination of hardware resources implemented by a circuit in the broad sense and software information processing that can be specifically realized by these hardware resources. Also, in this embodiment, various information is handled, and this information is represented by high or low signal values ​​as a binary bit collection consisting of 0 or 1, and communication and calculation can be performed on a circuit in the broad sense.

[0009] In addition, a circuit in the broad sense is a circuit realized by at least appropriately combining a circuit, circuitry, a processor, a memory, etc. That is, it includes an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)), etc.

[0010] <Embodiment 1> 1. Smartwatch Overview FIG. 1 is a diagram showing an example of an overview of a smartwatch. As shown in FIG. 1, the smartwatch 100 is attached to the user's fingers. The smartwatch 100 is assumed to be worn on the user's wrist with the dial facing the back of the hand. The crown of the smartwatch 100 is provided with an imaging unit 203, which will be described later. The smartwatch 100 estimates the three-dimensional posture of the user's fingers based on an image of the back of the user's hand captured by the imaging unit 203. The smart watch 100 is an example of an information processing system as claimed in the claims, and is also an example of a wearable device worn on the wrist.

[0011] 2. Hardware Configuration FIG. 2 is a diagram showing an example of the hardware configuration of the smartwatch 100. The smartwatch 100 includes, as its hardware configuration, a control unit 201, a storage unit 202, an imaging unit 203, a display unit 204, and a communication unit 205. The control unit 201 is a CPU (Central Processing Unit) or the like, and controls the entire smartwatch 100. The storage unit 202 is a ROM (Read Only Memory), a RAM (Random Access Memory), or a combination thereof, and stores programs and data used when the control unit 201 executes processing based on the programs. The storage unit 202 stores, for example, a deep learning network (described later).

[0012] The control unit 201 executes processing based on a program stored in the storage unit 202, thereby realizing the functional configuration of the smartwatch 100 shown in FIG. 4 and the processing of the activity diagram shown in FIG. 12, which will be described later. The imaging unit 203 captures an image of a subject, in this embodiment, the back of the user's hand. FIG. 3 is a diagram showing an example of an image captured by the imaging unit 203. As shown in FIG. 3, the captured image includes the back of the user's hand, but does not include the user's fingers. The control unit 201 estimates the three-dimensional posture of the user's fingers based on the captured image captured by the imaging unit 203.

[0013] In this embodiment, the imaging unit 203 continuously captures images at an interactive rate of approximately 20 frames per second, and the captured images are passed to the control unit 201. However, if the smartwatch 100 is required to recognize high-speed finger posture, such as when recognizing finger movements while playing the piano, the imaging unit 203 may capture images at a faster frame rate. The captured images include multiple frames. The control unit 201 estimates the three-dimensional posture of the user's fingers based on changes in the multiple frames. Using multiple frames allows for more accurate estimation of the three-dimensional posture of the user's fingers than using a single frame, as it is possible to extract changes in the state of the back of the hand for each frame. In particular, the multiple frames refers to 3 to 20 frames, specifically, for example, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 frames, and may be within a range between any two of the values ​​exemplified here. However, the more frames used, the greater the processing load, so five frames is preferable. In the following, this embodiment will be described using five frames as an example. The display unit 204 is a touch panel display or the like, and displays the time, the results of processing by the control unit 201, etc. The communication unit 205 connects the smartwatch 100 to a network and manages communication with other devices. The memory unit 202 is an example of a storage medium or memory area.

[0014] 3. Functional Configuration 4 is a diagram showing an example of the functional configuration of the smartwatch 100. The functional configuration of the smartwatch 100 includes a preprocessing unit 301, a posture estimation unit 302, a three-dimensional posture construction unit 303, and a recognition result use unit 304. The smartwatch 100 reads out a deep learning network (described later) stored in the memory unit 202, and estimates the three-dimensional posture of the user's fingers based on the captured image captured by the imaging unit 203 and the read deep learning network.

[0015] (Preprocessing unit 301) The preprocessing unit 301 receives a captured image from the imaging unit 203 and performs preprocessing when inputting the received captured image to the posture estimation unit 302. More specifically, the preprocessing unit 301 extracts the back of the hand region from the color image, which is the captured image received from the imaging unit 203. The preprocessing unit 301 also binarizes the extracted image of the back of the hand region to generate a binary image. The preprocessing unit 301 then inputs the color image of the back of the hand region and this binary image to the posture estimation unit 302. By binarizing the image using an appropriate threshold, the state of the muscles, bones, and tendons of the hand can be grasped in more detail. The preprocessing of the preprocessing unit 301 will be explained in more detail using Figure 5. Figure 5 is a diagram showing an example of the preprocessing. The preprocessing unit 301 extracts color images 601 to 605, which are images of the back of the hand region. Color image 601, color image 602, color image 603, color image 604, and color image 605 are each N consecutive mask images (five in this example). When creating a binary image 701 corresponding to color image 601, preprocessing unit 301 uses color images 601 to 605 to create binary image (hereinafter also referred to as movement history image) 701. Similarly, when creating movement history image 702 corresponding to color image 602, preprocessing unit 301 creates movement history image 702 using the color image (not shown) immediately preceding color image 602 to color image 605. Then, preprocessing unit 301 inputs five frames of color images of the back of the hand region and five frames of movement history image to posture estimation unit 302. Fig. 6 is a diagram showing an example of information processing by the smartwatch 100. The Preprocessing part in Fig. 6 corresponds to the processing performed by the preprocessing unit 301. A captured image 501 is a captured image received from the imaging unit 203, and is the same as the captured image in Fig. 3.

[0016] (Posture estimation unit 302) The posture estimation unit 302 uses a deep learning network to regress the joint angles of the user's fingers from the input image of the back of the hand and the movement history image. The Hand Pose Estimation section in Figure 6 corresponds to the deep learning network. Res18 502 and Res18 503 are ResNet (Residual Network) 18. FC 504 and FC 508 are Fully Connected layers. LSTM 505 is Long-Short Term Memory. KF 506 is a Kalman filter. The ultimate goal of the pose estimation unit 302 is to build a deep learning network for learning how to regress visual features of the back of the hand, such as deformation of the skin, veins, and tendons, onto finger movements. The 3D representation of the fingers uses a representation based on the relative joint angles of the four fingers excluding the thumb, rather than position-based 3D coordinates. Figure 7 shows an example of a 3D hand model. As shown in Figure 7, the 3D hand model of this embodiment represents the thumb with a single 3D vector, and the other four fingers using joint angles. M, P, and D represent the MCP joint (metacarpophalangeal joint), PIP joint (proximal interphalangeal joint), and DIP joint (distal interphalangeal joint) of a specific finger. v and h represent the vertical and horizontal flexion of the MCP.

[0017] For each training sequence of length T (herein T=5), a masked hand image I 1:T , Movement History Image (MHI)X 1:T , hand pose label y 1:T We create preprocessed data consisting of the pose labels y 1:T The joint angles α of the index finger, middle finger, ring finger, and little finger are t 1 , α t 2 , α t 3 , α t 4 Each hand position y t The joint angles α of the index finger, middle finger, ring finger, and little finger are t 1 , α t 2 , α t 3 , α t 4and the three-dimensional position t of the thumb. As shown in Figure 7, the joint angle α t i is a set of four elements (M v i , M h i , P i , D i ) and M v i , M h i is the vertical and horizontal rotation of the first joint, and P i , D i correspond to the rotation angles of the second and third joints, respectively. Our goal is to find the optimal rotation angle for the input mask image I 1:T , MHI X 1:T and a neural network-based regressor y ~ 1:T =f(I 1:T , X 1:T ) to learn.

[0018] To this end, we propose a two-stream LSTM-based network, DorsalNet, as shown in Figure 6. At each time step t, a masked hand image I t and MHI X t Visual features are extracted using two ResNet18s (Res18 502 and Res18 503) from the respective ResNets. That is, the deep learning network includes two ResNets. An image of the back of the hand region extracted from the captured image is input to Res18 502 of the two ResNets. A historical image is input to Res18 503 of the two ResNets. Res18 502 is an example of a first ResNet. Res18 503 is an example of a second ResNet.

[0019] Next, the deep learning network concatenates the two visual features and passes them through FC 504 to form a single visual feature φt. That is, the deep learning network includes a first fully connected layer. The feature related to the finger joint angles output from Res18 502 and the feature related to the finger joint angles output from Res18 503 are combined and input to FC 504. FC 504 is an example of a first fully connected layer.

[0020] In addition, in the deep learning network, LSTM505 is used to generate the visual feature sequence φ 1:T is processed into a temporal feature sequence, which is useful for 3D posture estimation. In addition, since most finger movements are simple linear movements, the deep learning network uses LSTM-based KF506, which uses LSTM to learn a motion model and a noise model. This allows for a more stable temporal feature sequence ψ 1:T can be obtained. We also add residual connections to the deep learning network to bypass KF506 for more direct feature learning.

[0021] For each frame t, the temporal feature ψ t now contains information from past frames that helps make hand pose predictions. Finally, the temporal features ψ t is the estimated hand posture y ~ t The FC 508 determines whether to use the output of the KF 506 or the output of the LSTM 505 via the path 507 in learning. Because the two-stream CNN architecture is computationally intensive, we focus on making the entire network lightweight to achieve real-time inference times, which is why we use ResNet18 but not deeper CNN architectures such as ResNet50 or ResNet101.

[0022] To supervise the training of DorsalNet, we define the following loss function:

number

number

number

[0023] (3D posture configuration section 303) The 3D posture construction unit 303 estimates the 3D posture of the user's fingers based on the joint angles of the user's fingers output from the deep learning network. The 3D Reconstruction part in Fig. 6 corresponds to the processing performed by the 3D posture construction unit 303. In order to visualize the hand posture sequence, the 3D posture construction unit 303 ~ t The hand simulator is capable of mapping the joint angle of the thumb onto a 3D mesh. The 3D posture construction unit 303 uses the joint angle of the thumb, so the estimated value e ~ tWe employ inverse kinematics (IK) to solve for the thumb joint angle using the algorithm. Finally, the 3D pose construction unit 303 employs a smoothing function in the simulator to handle noisy results. The simulator's inference time is 10 milliseconds, and the inference time with the network is 41 milliseconds, which is considered to be excellent for real-time performance in our use case. At this point, we have been able to reconstruct the 3D hand pose excluding finger rotation.

[0024] The estimation results are divided into static gestures (ASL 0-9) and dynamic gestures (tapping 0-4), and are shown in Figures 8 and 9. Figure 8 shows the estimation results for static gestures (ASL 0-9). Figure 9 shows the estimation results for dynamic gestures (tapping 0-4). NearestN is a method developed by Padraing Cunningham and Sarah Jane Delany (2007). k-Nearest Neighbor Classifiers. Yeo et al. is a methodology of Hui-Shyong Yeo, Erwin Wu, Juyoung Lee, Aaron Quigley, and Hideki Koike. 2019. Opisthenar: Hand Poses and Finger Tapping Recognition by Observing Back of Hand Using Embedded Wrist Camera. In Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology (UIST '19). Association for Computing Machinery, New York, NY, USA, 963-971. DOI:

[0025] In static results, both methods outperformed Yeo et al.'s work. For the general model, both methods achieved the same result of 88.8%. This is because static gestures have fewer periodic movements, which may cause errors in the KF predictions. However, both methods achieved high accuracy in detecting hand gestures, and no obvious overtraining was observed. The dynamic results showed the highest accuracy in all three conditions (89.4% for the individual model, 86.8% for the general model, and 79.8% for the Leave-1-User model). These results demonstrate that it is possible to detect dynamic hand postures, which may broaden the scope of use of the system. Furthermore, the dynamic results were largely consistent with the static results, indicating that hand back deformation is suitable for both dynamic and static classification.

[0026] During the development of this system, we also confirmed the possibility of recognizing various types of grips by utilizing the features on the back of the hand. Therefore, we collected a dataset of different grip types, as shown in Figure 10. Figure 10 shows an example of a dataset of different grip types. Figure 11 shows an example of a confusion matrix and heat map of 50 tests for each grip type. The overall accuracy was 75.1%, with the best result being the extended grip type (holding a plate), which achieved 49 correct classifications (98%). In other words, the 3D pose construction unit 303 can estimate the 3D pose of the user's fingers and, based on the estimation results, estimate the object the user is holding.

[0027] (Recognition result usage section 304) The recognition result using unit 304 uses the result of recognition by the three-dimensional posture constructing unit 303. For example, the recognition result using unit 304 may input information to the smartwatch 100 based on a change in the three-dimensional posture of the fingers recognized by the three-dimensional posture constructing unit 303, or may input information to another device with which communication is possible via the communication unit 205. As another example, the recognition result using unit 304 may input information to the smartwatch 100 or the like based on the three-dimensional posture of the user's fingers (for example, the pointing direction or angle of a specific finger) estimated by the recognition result using unit 304. Note that the information to be input is not limited to letters and numbers, and may also be operation information for a device, such as clicking or tapping. Therefore, the smartwatch 100 can be used as a substitute for a virtual mouse and virtual keyboard, or as a substitute for VR gloves or stick-type controllers used in VR (Virtual Reality) spaces. For example, VR gloves require the user to raise their hands at a certain height to be positioned within the camera's imaging range, which can cause hand fatigue. However, with the configuration of this embodiment, the user can keep their hands down or place them on a desk, eliminating the above-mentioned problem.

[0028] 4. Information Processing FIG. 12 is an activity diagram showing an example of information processing in the smartwatch 100. In A1001, the preprocessing unit 301 determines whether a predetermined operation for starting processing has been performed. If the preprocessing unit 301 determines that a predetermined operation for starting processing has been performed, the preprocessing unit 301 proceeds to A1002. If the preprocessing unit 301 determines that a predetermined operation for starting processing has not been performed, the preprocessing unit 301 repeats the processing of A1001.

[0029] In A1002, the pre-processing unit 301 receives an image from the imaging unit 203, extracts the back of the hand area from the received color image, and inputs the image of the extracted back of the hand area and a movement history image obtained by binarizing the image of the extracted back of the hand area to the posture estimation unit 302. In A1003, the posture estimation unit 302 estimates the joint angles of the user's fingers from the input image of the back of the hand area and the movement history image.

[0030] In A1004, the three-dimensional posture construction unit 303 estimates the three-dimensional posture of the user's fingers based on the estimated joint angles of the user's fingers. In A1005, the recognition result use unit 304 inputs information to the smartwatch 100 based on, for example, changes in the three-dimensional posture of the fingers recognized by the three-dimensional posture construction unit 303. In A1006, the recognition result use unit 304 determines whether a predetermined operation for terminating the processing has been performed. If the recognition result use unit 304 determines that a predetermined operation for terminating the processing has been performed, it terminates the information processing shown in Fig. 12. If the recognition result use unit 304 determines that a predetermined operation for terminating the processing has not been performed, the processing returns to A1002.

[0031] According to this embodiment, the 3D posture of the user's fingers can be estimated more accurately than in the past. Furthermore, since the 3D posture of the user's fingers can be estimated only from an image of the back of the hand, it can be implemented seamlessly in devices such as smartwatches equipped with cameras, allowing users to wear or hold devices such as smartwatches naturally. For example, if one were to implement a system that captures an image including a user's fingers and estimates the three-dimensional posture of the user's fingers from the captured image, a device such as a smartwatch for capturing the image including the user's fingers would need to be a certain height, and would not be able to be worn naturally on the user's wrist, etc. Furthermore, even if a camera is mounted on a ring or the like, the size of the camera and other factors can cause physical interference between the finger on which the ring is worn and other fingers, making it difficult for the user to wear the ring comfortably.

[0032] <Variation 1> FIG. 13 is a diagram showing an example of an information processing system according to Modification 1. The information processing system includes, as a system configuration, a smartwatch 100 and a smartphone 1110 capable of communicating with the smartwatch 100. In Embodiment 1, the smartwatch 100 performed all processing. However, in Modification 1, the smartwatch 100 captures an image of the back of the user's hand and transmits the captured image to the smartphone 1110. The smartphone 1110 may then receive the captured image and estimate the three-dimensional posture of the user's fingers based on the captured image. The smartphone 1110 may input information to itself based on changes in the estimated three-dimensional posture of the user's fingers, or may transmit the estimation results to the smartwatch 100.

[0033] Modification 1 also makes it possible to estimate the 3D posture of a user's fingers more accurately than conventional methods. Furthermore, because it is possible to estimate the 3D posture of a user's fingers from only an image of the back of the hand, it can be seamlessly implemented in an information system that combines a smartwatch equipped with a camera and a smartphone.

[0034] <Variation 2> Another example of the recognition result using unit 304 will be described. The recognition result using unit 304 may be configured to output the result of recognition by the 3D pose constructing unit 303. The recognition result using unit 304 may output the recognition result to the display unit 204, or to another device with which communication is possible via the communication unit 205. As another example, when the 3D pose constructing unit 303 estimates an object grasped by the user, the recognition result using unit 304 may change the operation mode of the smartwatch 100 or another device with which communication is possible with the smartwatch 100, depending on the estimated object. For example, when the 3D pose constructing unit 303 estimates that the user is holding a touch pen, the recognition result using unit 304 may control the operation mode of a tablet terminal with which communication is possible with the smartwatch 100 to be changed from normal mode to touch pen input mode. As another example, the recognition result usage unit 304 may launch or terminate a specific application on the smartwatch 100 or another device that can communicate with the smartwatch 100 based on the results of estimation by the 3D posture construction unit 303.

[0035] According to the second modification, various processes can be performed based on the estimation result of the three-dimensional posture of the user's fingers by the three-dimensional posture construction unit 303.

[0036] <Additional Notes> The invention may be provided in the following aspects: The information processing system includes an imaging unit that captures an image of the back of the user's hand, and the control unit extracts an image of the back of the hand from the captured image captured by the imaging unit, and estimates the three-dimensional posture of the user's fingers based on the extracted image. In the information processing system, the captured image includes a plurality of frames, and images of the back of the hand area are extracted from the plurality of frames, and the control unit estimates the three-dimensional posture of the user's fingers based on changes in the images of the back of the hand area in the plurality of frames. In the information processing system, the control unit reads out a deep learning network stored in a memory area and estimates the three-dimensional posture of the user's fingers based on an image of the back of the hand extracted from the captured image and the deep learning network. In the information processing system, the control unit extracts an image of the back of the hand from the captured image, inputs the extracted image and a movement history image showing the movement history related to the extracted image into the deep learning network, and estimates the three-dimensional posture of the user's fingers based on the joint angles of the user's fingers output from the deep learning network. In the information processing system, the information processing system is a wearable device worn on the wrist of the user. In the information processing system, the control unit estimates a three-dimensional posture of the user's fingers and inputs information based on a change in the three-dimensional posture. In the information processing system, the control unit estimates a three-dimensional posture of the user's fingers, and estimates an object being grasped by the user based on the estimation result. In the information processing system, the control unit changes the operation mode based on the estimated object. In the information processing system, the deep learning network includes two ResNets (Residual Networks), and an image of the back of the hand area extracted from the captured image is input to a first ResNet of the two ResNets, and the movement history image is input to a second ResNet of the two ResNets. In the information processing system, the deep learning network includes a first fully connected layer, and the features related to finger joint angles that are the output of the first ResNet and the features related to finger joint angles that are the output of the second ResNet are combined and input to the first fully connected layer. In the information processing system, the deep learning network includes an LSTM (Long short-term memory), a Kalman filter, and a second fully connected layer, and data output from the first fully connected layer is input to the LSTM, and the second fully connected layer determines whether to use data output from the LSTM that is input to the Kalman filter and output from the Kalman filter, or data that is not input to the Kalman filter. An information processing method executed by an information processing system, which extracts an image of the back of the user's hand from a captured image including the back of the hand, and estimates the three-dimensional posture of the user's fingers based on the extracted image. A program for causing a computer to function as a control unit of the information processing system. Of course, this is not the case.

[0037] For example, the above-mentioned program may be provided as a computer-readable non-transitory storage medium that stores the program. Furthermore, the above-described embodiments and modifications may be combined in any desired manner. Furthermore, although the above-described embodiment and modified examples have been described using a smart watch as an example, the present invention is not limited to a smart watch. Another example may be a wristband equipped with a small camera. A wristband equipped with a small camera is also an example of a wearable device.

[0038] Finally, while various embodiments of the present invention have been described, these are presented by way of example only and are not intended to limit the scope of the invention. The novel embodiments may be embodied in various other forms, and various omissions, substitutions, and modifications may be made without departing from the spirit of the invention. The embodiments and their modifications are intended to be included within the scope and spirit of the invention, as well as within the scope of the inventions and their equivalents as defined in the appended claims. [Explanation of symbols]

[0039] 100: Smartwatch 201: Control unit 202: Storage section 203: Imaging unit 204: Display section 205: Communications Department 301: Preprocessing section 302: Posture estimation section 303:3D posture configuration part 304: Recognition result usage part 501: Captured image 1110: Smartphone

Claims

1. An information processing system, A control unit is provided. The control unit extracts an image of the back of the hand from a captured image including the back of the user's hand, inputs a color image of the back of the hand extracted from the captured image and a movement history image that is created based on the extracted color images and shows a movement history corresponding to the extracted color image to a deep learning network, and outputs the joint angles of the user's fingers from the deep learning network as an estimation result. Information processing system.

2. 2. The information processing system according to claim 1, An imaging unit is included, the imaging unit captures an image of the back of the user's hand, the control unit extracts a color image of a region of the back of the hand from the captured image captured by the imaging unit, and estimates joint angles of the user's fingers based on the extracted color image. Information processing system.

3. 3. The information processing system according to claim 1, the captured image includes a plurality of frames, the control unit extracts color images of the back of the hand region in the plurality of frames, creates the movement history image based on changes in the extracted color images of the back of the hand region in the plurality of frames, and inputs the created movement history image together with the color images into the deep learning network, thereby outputting the joint angles of the user's fingers as the estimation result. Information processing system.

4. 4. The information processing system according to claim 1, The information processing system is a wearable device worn on the user's wrist. Information processing system.

5. 5. The information processing system according to claim 1, the control unit inputs information based on the estimated change in the joint angle. Information processing system.

6. 6. The information processing system according to claim 1, the control unit estimates joint angles of the user's fingers and, based on the estimation result, estimates an object being held by the user. Information processing system.

7. 7. The information processing system according to claim 6, The control unit changes the operation mode based on the estimated object. Information processing system.

8. 8. The information processing system according to claim 1, The deep learning network It includes two ResNets (Residual Networks), A color image of the back of the hand extracted from the captured image is input to a first ResNet of the two ResNets; The exercise history image is input to a second ResNet of the two ResNets. Information processing system.

9. 9. The information processing system according to claim 8, The deep learning network a first fully connected layer; The feature amount related to the angle of the finger joints, which is the output of the first ResNet, and the feature amount related to the angle of the finger joints, which is the output of the second ResNet, are combined and input to the first fully connected layer. Information processing system.

10. 10. The information processing system according to claim 9, The deep learning network a long short-term memory (LSTM), a Kalman filter, and a second fully connected layer; The data output from the first fully connected layer is input to the LSTM; The second fully connected layer determining whether to use data output from the LSTM that has been input to the Kalman filter and output from the Kalman filter or data that has not been input to the Kalman filter; Information processing system.

11. An information processing method executed by an information processing system, An image of the back of the hand is extracted from a captured image including the back of the user's hand, a color image of the back of the hand extracted from the captured image and a movement history image showing a movement history corresponding to the extracted color image are input to a deep learning network, and the deep learning network is caused to output the joint angles of the user's fingers as an estimation result. Information processing methods.

12. A program, A program for causing a computer to function as a control unit of the information processing system according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Vein scanning device for automatic gesture and finger recognition

    CN111052047A

  • Gesture command input device, gesture command input method, gesture command input program, and image display system

    JP2017207949A

  • User input using proximity sensing

    US20130234970A1