Information processing device and information processing system

By employing intersecting cameras near the nose tip to capture hand and mouth movements, the method addresses the burdensomeness and complexity of existing technologies, enabling accurate and cost-effective real-time avatar synchronization.

WO2025158914A1PCT designated stage Publication Date: 2025-07-31SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/000464
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-25
Filing Date
2025-01-09
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing methods for synchronously detecting a user's hand and mouth movements for real-time reflection in avatars in virtual spaces are burdensome due to the need for separate sensors and cameras, leading to increased device size and cost, and require complex synchronization processes.

Method used

Utilizing two cameras positioned on the left and right near the nose tip of a user, with intersecting viewing angles to capture mouth and hand movements simultaneously, eliminating the need for separate sensors and simplifying synchronization.

Benefits of technology

This approach allows for accurate, real-time detection and reflection of hand and mouth movements in avatars without the burden of additional sensors, reducing device size and cost while enhancing estimation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025000464_31072025_PF_FP_ABST
    Figure JP2025000464_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information processing device and an information processing system that make it possible to synchronously detect the hand movement and mouth movement of a user at a low cost, and reflect it on an avatar in real time and with high accuracy. At positions in the vicinity of contact with the nose tip of a user when a HMD is worn and respectively corresponding to the right and left sides of the HMD around the nose tip are provided two cameras for capturing an image below the nose tip so that the cameras respectively capture a periphery including the lips of the user at different angles of view, and that the respective angles of view intersect at a region close to the lips as the boundary. The present invention can be applied to an HMD.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device and information processing system

[0001] The present disclosure relates to an information processing device and an information processing system, and more particularly to an information processing device and an information processing system that can detect a user's hand movements and mouth movements in synchronization and reflect them in an avatar in real time with high accuracy.

[0002] Technology that allows users to communicate smoothly in real time using avatars in virtual spaces is becoming widespread. To realize this technology, it is necessary to synchronize and detect information on the user's multiple movements, especially mouth and hand movements, and reflect this information in the avatar in real time with high accuracy.

[0003] Therefore, a technology has been proposed in which a motion sensor that detects hand movements is attached to the hand and a separate camera is installed near the mouth to capture mouth movements, thereby detecting hand and mouth movements and reflecting them in an avatar (see Patent Document 1).

[0004] Japanese Patent Application Laid-Open No. 2018-200671

[0005] However, the technology described in Patent Document 1 requires the user to wear a motion sensor on the hand, which places a burden on the user when moving the hand.

[0006] In addition, since the motion sensor that detects hand movements and the camera that captures mouth movements detect hand movements and mouth movements independently, it is necessary to synchronize the detected hand movements and mouth movements with high precision and reflect them in the avatar.

[0007] Furthermore, to achieve high-precision synchronization, it is possible to use high-precision motion sensors or cameras, but this raises concerns about the burden on users due to the increased size and cost of the device configuration.

[0008] The present disclosure has been made in consideration of such circumstances, and in particular, is directed to detecting a user's hand movements and mouth movements in synchronization at low cost and reflecting them in an avatar in real time with high accuracy.

[0009] An information processing device and an information processing system according to one aspect of the present disclosure include two imaging units located near the tip of a user's nose, at corresponding positions on the left and right of the nose tip, which capture images of an area below the tip of the nose, and the two imaging units each capture an image of the area around the user's mouth at a different angle of view, and the different angles of view intersect.

[0010] In one aspect of the present disclosure, two imaging units are provided near the tip of a user's nose, at corresponding positions on the left and right of the nose tip as a center, to capture an image below the tip of the nose, and the two imaging units each capture an image of the area including the user's mouth at different angles of view, and the different angles of view intersect.

[0011] 1 is a diagram illustrating communication in a virtual space using an HMD. FIG. 1 is a diagram illustrating an overview of the present disclosure. FIG. 2 is a diagram illustrating an example of an external configuration of a first embodiment of an HMD of the present disclosure. FIG. 3 is a diagram illustrating a range imaged by the camera of the HMD of FIG. 3. FIG. 4 is a diagram illustrating a range imaged by the camera of the HMD of FIG. 3. FIG. 4 is a diagram illustrating an example of an image imaged by the camera of the HMD of FIG. 3. FIG. 5 is a diagram illustrating an example of an image imaged by the camera of the HMD of FIG. 3. FIG. 6 is a diagram illustrating an example of a hardware configuration of the HMD of FIG. 3. FIG. 7 is a diagram illustrating landmark coordinates of a hand. FIG. 8 is a diagram illustrating a method of determining whether a hand is the palm or the back of the hand from the landmark coordinates of the hand. FIG. 9 is a diagram illustrating an example of an image imaged by the camera of the HMD of FIG. 3 when the left and right hands are crossed. FIG. 10 is a diagram illustrating an example of an image imaged by the camera of the HMD of FIG. 3 when both hands are stretched to the left. FIG. 11 is a diagram illustrating an example of an image imaged by the camera of the HMD of FIG. 3 when both hands are stretched to the right. FIG. 12 is a diagram illustrating an example of calculating landmark coordinates from an image of a hand and reflecting them on an avatar. FIG. 13 is a diagram illustrating an example of extracting feature amounts based on the landmark coordinates of the hand. FIG. 14 is a diagram illustrating landmark coordinates around the mouth. FIG. 10 is a diagram illustrating an example of reflecting landmark coordinates around the mouth in an avatar. FIG. 11 is a diagram illustrating an example of extracting feature amounts based on landmark coordinates around the mouth. FIG. 12 is a diagram illustrating an example of an action unit that makes a movement from a closed mouth to a toothy smile. FIG. 13 is a diagram illustrating an example of an action unit that makes a movement from a closed mouth to a lower lip. FIG. 14 is a diagram illustrating an example of an action unit that makes a movement from a closed mouth to a lifting of the lower jaw. FIG. 15 is a diagram illustrating an example of an action unit that makes a movement from a closed mouth to a pursing and protruding of the lips. FIG. 16 is a diagram illustrating an example of an action unit that makes a movement from a closed mouth to a pulling of the lips inward. FIG. 17 is a diagram illustrating an example of an action unit that makes a movement from a closed mouth to a toothless smile. FIG. 18 is a diagram illustrating an example of an action unit that makes a movement from a closed mouth to a puffing of the cheeks. FIG. 19 is a diagram illustrating an example of an action unit that makes a movement from a closed mouth to a pulling of the lips to the side.32. A diagram illustrating an example of an action unit that moves lips from a closed mouth to a peaked state. A diagram illustrating the operation of an image processing unit. A flowchart illustrating avatar reflection processing by the HMD of FIG. 8. A diagram illustrating an example of an external configuration of a modified example of the first embodiment of the HMD of the present disclosure. A diagram illustrating an example of an external configuration of a second embodiment of the HMD of the present disclosure. A diagram illustrating an example hardware configuration (part 1) of the HMD of FIG. 32. A flowchart illustrating avatar reflection processing by the HMD of FIG. 33. A diagram illustrating calibration data. A diagram illustrating calibration data. A diagram illustrating an example hardware configuration (part 2) of the HMD of FIG. 32. A flowchart illustrating calibration processing by the HMD of FIG. 33. A flowchart illustrating avatar reflection processing by the HMD of FIG. 33. A diagram illustrating an example configuration of a general-purpose computer.

[0012] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.

[0013] Hereinafter, embodiments of the present technology will be described in the following order.

[0014] 1. Overview of the present disclosure 2. First embodiment 3. Modification of the first embodiment 4. Second embodiment 5. Third embodiment 6. Example of execution by software

[0015] <<1. Overview of the Present Disclosure>> The present disclosure, in particular, enables a user's hand movements and mouth movements to be detected synchronously at low cost and reflected in an avatar in real time with high accuracy. Therefore, an overview of the present disclosure will first be described.

[0016] In order for users to have a smooth, real-time communication experience using virtual avatars (hereinafter simply referred to as avatars) in a virtual space, it is necessary to reflect the user's movements, particularly mouth and hand movements, in the avatar in real time and with high accuracy.

[0017] FIG. 1 shows an example in which users communicate with each other using avatars in a virtual space.

[0018] More specifically, image P1 on the left side of Figure 1 depicts users H1 and H2, consisting of Rose and Jack in real space, wearing HMDs (Head Mounted Displays) 11-1 and 11-2, respectively, and communicating by viewing images in a virtual space.

[0019] The right part of FIG. 1 shows an image V11-1 that is displayed on an HMD 11-1 worn by a user H1 who is Rose, and that shows the state of the virtual space viewed by the user H1.

[0020] In image V11-1, the facial expressions, gestures, hand movements, etc. of users H1 and H2, consisting of Rose and Jack in the real space, are reflected and displayed on avatars A1 and A2 in the virtual space.

[0021] By viewing the virtual space image V11-1 presented on the HMD 11-1, the user H1 recognizes the mouth movements Gm, which represent the facial expressions of the avatar A2 in the virtual space, and the hand movements Gh, which represent gestures and hand movements, and realizes communication with the user H2 through the avatar A2 in the virtual space.

[0022] At this time, the mouth movements Gm, which represent the facial expressions of avatar A2, and the hand movements Gh, which are gestures and hand movements, etc., must be synchronized in real time with the mouth movements Zm and hand movements Zh of user H2 in the real space.

[0023] To detect the mouth movement Zm of the user H2 in real space, for example, an image obtained by capturing an image of the mouth with a camera provided separately from the HMD 11-2 worn by the user H2 has often been used.

[0024] Furthermore, to detect the hand movement Zh of the user H2 in real space, for example, a motion sensor attached to the hand of the user H2 has often been used.

[0025] However, with such a configuration, the user H2 often feels annoyed by having the motion sensor attached to his / her hand, and as a result, there is a risk that he / she will restrict the movement of his / her hand.

[0026] Furthermore, since a motion sensor and an image sensor for detecting hand and mouth movements must be installed separately from the HMD 11, the device configuration becomes larger and the device costs also increase.

[0027] Furthermore, because sensors will be installed to detect hand and mouth movements, the detection results will need to be synchronized in real time and reflected in the avatar, which will increase the effort and cost involved in synchronization.

[0028] Therefore, in the present disclosure, as shown in Figure 2, two cameras are provided on the left and right near the contact point with the tip of the HMD's nose, with their respective angles of view (field of view) horizontally intersecting at the mouth as a boundary. This allows the area near the mouth (the area around the mouth including the mouth) to be imaged from the left and right, and also allows the area below the area near the mouth where the user's hands are moving to be imaged, thereby allowing the movement of the mouth and the movement of the hands to be acquired in synchronization.

[0029] In Figure 2, cameras 41L and 41R are provided on the left and right sides, as viewed from the front of the paper in Figure 2, near the area where the HMD 31 contacts the tip of the nose of the user H11, and each of the angles of view FOV-L and FOV-R includes a range near the mouth and is arranged so that a range farther than the range including the range near the mouth intersects on the left and right.

[0030] That is, the angle of view FOV-L of the camera 41R provided on the right side in Fig. 2 includes a range near the mouth of the user H11 and is set to a range to the right of the user H11 as viewed from the front of the page in Fig. 2 (a range to the left of the user H11). Conversely, the angle of view FOV-R of the camera 41L provided on the left side in Fig. 2 includes a range near the mouth of the user H11 and is set to a range to the left of the user H11 as viewed from the front of the page in Fig. 2 (a range to the right of the user H11).

[0031] With this configuration, the cameras 41L and 41R can capture images of the area around the mouth of the user H11 from the left and right, respectively, and it is therefore possible to detect the movement of the mouth based on the captured images.

[0032] Furthermore, the respective angles of view FOV-R, FOV-L of cameras 41L, 41R are set so that they intersect left and right with the range near user H11's mouth as the boundary, so that in a range farther than the range near the mouth as seen from each of cameras 41L, 41R, the camera 41R on the right side in the figure captures an angle of view FOV-L that is the range to the left of user H11, and the camera 41L on the left side in the figure captures an angle of view FOV-R that is the range to the right of user H11.

[0033] This allows the cameras 41R and 41L to capture images of the area extending beyond the mouth to the tips of the left and right hands, making it possible to detect the movements of the left and right hands.

[0034] Furthermore, since each of the cameras 41R and 41L can detect the mouth movement and the hand movement in synchronization, it becomes possible to reflect the detected mouth movement and hand movement in synchronization with high accuracy on the corresponding avatar.

[0035] Furthermore, since the components for detecting mouth movement and hand movement are not separate, it is possible to prevent the device configuration from becoming too large and also to prevent an increase in device costs.

[0036] <<2. First embodiment>> <External configuration example of HMD> Next, an external configuration example of the first embodiment of the HMD of the present disclosure will be described with reference to Fig. 3. The left part of Fig. 3 is a front external view of the HMD 101 worn by a user H101, and the right part of Fig. 3 is a left side external view of the HMD 101 worn by the user H101. Although not shown, the right side external view of the HMD 101 is omitted because it is a symmetrical view of the left side external view.

[0037] The HMD 101 in FIG. 3 is configured to be worn on the head of a user H101 so as to cover the left and right eyes, and is configured from cameras 111L and 111R and a display 112.

[0038] The cameras 111L and 111R are respectively provided below the left and right eyes of the user H101, at the contact point between the main body of the HMD 101 and the tip of the nose NT.

[0039] The camera 111L is provided below the left eye, near the contact point between the main body of the HMD 101 and the left side of the nose tip NT. The forward vertical field of view FOV-F of the camera 111L is a range below the nose tip NT, including the hand when the hand is stretched forward. The rear vertical field of view FOV-B of the camera 111L is a range below the nose tip NT, including the mouth LP. The horizontal field of view FOV-R of the camera 111L is a range below the nose tip NT on the right side of the body of the user H101, including the mouth LP.

[0040] Similarly, camera 111R is provided below the left eye, near the contact point between the main body of the HMD 101 and the right side of the nose tip NT. The field of view FOV-F and field of view FOV-B of camera 111R are the same as those of camera 111L. Furthermore, the horizontal field of view FOV-L of camera 111R is a range below the nose tip NT on the left side of the body of user H101, including the mouth LP.

[0041] In addition, the angles of view FOV-R and FOV-L are designated by different symbols R and L, with the symbols attached to the hyphen representing the cameras 111L and 111R that capture the images, because they are assigned based on the user H101.

[0042] That is, the angle of view FOV-L is captured by the camera 111R provided on the right side of the user H101, but since it is an angle of view on the left side as seen from the user H101, "-L" is added.

[0043] Similarly, the angle of view FOV-R is captured by the camera 111L provided on the left side of the user H101, but since it is an angle of view on the right side as seen from the user H101, "-R" is added.

[0044] Hereinafter, with regard to other reference numerals, reference numerals for distinguishing between left and right will be given based on the direction as seen from the user H101 with respect to the installation position or imaging range.

[0045] The angles of view FOV-L and FOV-R are set to cover the range below the tip of the nose NT, including the mouth LP, and to cover the range of movement of the left and right hands. For this reason, for example, whether the user H101 holds his or her left and right hands LH and RH down as shown in Fig. 4 or spreads them as shown in Fig. 5, the angles of view are set so that the left hand LH falls within the angle of view FOV-L and the right hand RH falls within the angle of view FOV-R, unless the user H101 crosses his or her left and right hands or places both hands simultaneously on the right or left side in front of the user H101.

[0046] For this reason, for example, when the left and right hands are placed with their palms facing downwards and slightly spread apart, camera 111R captures an image with an angle of view FOV-L, capturing an image of the left hand LH1 with the palm facing downwards (the back of the hand facing up) over the mouth on the right side, as shown in image PL1 on the left side of Figure 6.

[0047] Similarly, camera 111L captures an image with a field of view FOV-R, thereby capturing an image of the right hand RH1 with the palm facing downwards (the back of the hand facing up) over the mouth on the left side, as shown in image PR1 on the right side of Figure 6.

[0048] Furthermore, the angles of view FOV-F and FOV-B are each set to a range that includes the mouth LP and covers the range in the front-to-back direction, so even if the left and right hands are thrust forward, the left hand is set to fit within the angle of view FOV-L and the right hand is set to fit within the angle of view FOV-R, unless the left and right hands are crossed or both hands are simultaneously placed on the right or left side in front of the user H101.

[0049] Therefore, for example, when the left and right hands are held out in front with the palms facing up, the camera 111R captures an image with an angle of view FOV-L, thereby capturing an image of the left hand LH11 being held out in front with the palm facing up, over the mouth on the right side, as shown in the image PL11 on the left side of Figure 7.

[0050] Similarly, camera 111L captures an image with a field of view FOV-R, thereby capturing an image of the right hand RH11 with the palm facing up and thrust forward over the left side of the mouth, as shown in image PR12 on the right side of Figure 7.

[0051] In this way, cameras 111L and 111R capture the left and right mouths of user H101 in the range below and in front of the tip of the nose NT as the user's hand moves, and also capture the left-right inverted hands RH11 and LH11 of user H101.

[0052] That is, by configuring the cameras 111L and 111R in this way, it is possible to simultaneously acquire images of the mouth and the left and right hands of the user H101.

[0053] Furthermore, since cameras 111L and 111R can capture images of the left and right mouths, respectively, it is possible to detect the movement of the mouth based on the image of the mouth, and also to identify the center positions of the left and right sides of user H101's face based on the position of the mouth.

[0054] Furthermore, the user H101 does not need to wear a motion sensor or the like to detect hand movements, so hand movements are not restricted.

[0055] Furthermore, by consolidating the sensor for detecting mouth movement and the sensor for detecting hand movement into cameras 111L and 111R, it is possible to simultaneously achieve miniaturization of the device configuration and cost reduction.

[0056] As described above, the FOV-R and FOV-L fields of view are set so that they intersect left and right, covering the area near the mouth (the area around the mouth) and the area where the hands are located, thereby capturing images of the mouth and hands simultaneously. However, although the FOV-R and FOV-L fields of view are set so that they intersect left and right, the front direction as seen from the user H101 must be within both fields of view. Therefore, the FOV-R and FOV-L fields of view must overlap with each other. This configuration in which the two fields of view intersect left and right enables stereo imaging of the overlapping area, making it possible to acquire depth information. This depth information makes it possible, for example, to determine whether the chin is protruding forward, backward, left, or right, and to infer the center line of the face. Furthermore, 3D mesh information of the face can be acquired based on the depth information, which makes it possible to generate a 3D avatar based on this information. Furthermore, by using the depth information of the hand, it becomes possible to recognize the movement of the hand more appropriately, and therefore it becomes possible to improve the accuracy of recognizing gestures using the hand, for example.

[0057] The display 112 is configured to cover both eyes of the user H101 and displays images of the virtual space including the user H101 and avatars of other users. By viewing the images of the virtual space displayed on the display 112, the user can experience a pseudo-sensation as if they were actually present in the virtual space. Furthermore, communication between other users and their corresponding avatars in such a virtual space can be achieved in the virtual space in a manner equivalent to communication in real space.

[0058] <Example of Hardware Configuration of HMD in FIG. 3> Next, an example of the hardware configuration of the HMD 101 in FIG. 3 will be described with reference to FIG.

[0059] The HMD 101 is composed of a control unit 131, an input unit 132, an output unit 133, a memory unit 134, a communication unit 135, a drive 136, a removable storage medium 137, and cameras 111R and 111L, which are connected to each other via a bus 138 and can send and receive data and programs.

[0060] The control unit 131 is composed of a processor and a memory, and controls the overall operation of the HMD 101. The control unit 131 also includes a pre-processing unit 151, a recognition unit 152, and an image processing unit 153.

[0061] The preprocessing unit 151 performs preprocessing on the images captured by the cameras 111L and 111R to detect mouth movements and hand movements, and outputs the images to the recognition unit 152. The preprocessing for detecting mouth movements and hand movements will be described in detail later.

[0062] The recognition unit 152 recognizes (estimates) the movements of the mouth and the movements of the hands based on the preprocessing results for detecting the movements of the mouth and the hands preprocessed by the preprocessing unit 151, and outputs the estimation results to the image processing unit 153. Note that the recognition (estimation) of the movements of the mouth and the hands by the recognition unit 152 will be described in detail later.

[0063] The image processing unit 153 reflects the recognition results of the mouth movement and hand movement supplied by the recognition unit 152 in the mouth movement and hand movement of the avatar, and outputs and displays them on the display 112. The image processing by the image processing unit 153 will be described in detail later.

[0064] The input unit 132 is composed of input devices such as a keyboard, a mouse, and a touch panel for inputting various types of information, and supplies the control unit 131 with various signals corresponding to the input information.

[0065] The output unit 133 is controlled by the control unit 131 and includes a display 112 and an audio output unit (not shown). The display 112 is a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) display, and displays images in the virtual space generated by the image processing unit 153 and various processing results. The audio output unit (not shown) is an audio output device such as a speaker, and outputs various sounds, music, sound effects, and the like as audio.

[0066] The storage unit 134 is composed of a hard disk drive (HDD), a solid state drive (SSD), or a semiconductor memory, and is controlled by the control unit 131 to write or read various data and programs.

[0067] The communication unit 135 is controlled by the control unit 131 and realizes wired or wireless communication such as that represented by LAN (Local Area Network) or Bluetooth (registered trademark), and transmits and receives various data and programs to and from other users' HMDs 101 or other information processing devices via the network as necessary.

[0068] The drive 136 reads and writes data from and to a removable storage medium 137 such as a magnetic disk (including a flexible disk), an optical disk (including a CD-ROM (Compact Disc-Read Only Memory) and a DVD (Digital Versatile Disc)), a magneto-optical disk (including an MD (Mini Disc)), or a semiconductor memory.

[0069] <Hand Movement Recognition Method> Next, a description will be given of the preprocessing performed by the preprocessing unit 151 to recognize (estimate) hand movements, and a method for recognizing (estimating) hand movements performed by the recognition unit 152 based on the preprocessing results.

[0070] The pre-processing unit 151 identifies the area in which the hands of the user H101 are present, based on the images captured by the cameras 111L and 111R, by using, for example, semantic segmentation.

[0071] Next, the preprocessing unit 151 detects the positions of the joints set at predetermined positions as three-dimensional landmark coordinates such as (x, y, z) based on the image of the identified area of ​​the user H101's hand.

[0072] Here, for example, 21 joint positions are set as landmark coordinates for detecting hand movements, as shown in Fig. 9. The numbers assigned to the palm in Fig. 9 are codes that distinguish the respective landmark coordinates.

[0073] FIG. 9 shows an example in which the position of the wrist on the palm RH101 of the right hand is set as landmark coordinate 1, landmark coordinates 2 to 5 are set for the thumb, landmark coordinates 6 to 9 are set for the index finger, landmark coordinates 10 to 13 are set for the middle finger, landmark coordinates 14 to 17 are set for the ring finger, and landmark coordinates 18 to 21 are set for the little finger.

[0074] Note that landmark coordinates 1, 5, 9, 13, 17, 21, etc. are the wrist and fingertips, respectively, and are not organs that can be called joints that make up the human hand. However, here, the parts where landmark coordinates necessary for expressing hand movements are set will all be referred to as joints.

[0075] The wrist landmark coordinate 1 is connected to the landmark coordinates 2, 6, and 18 as adjacent joints, and the landmark coordinates 6, 10, 14, and 18 are connected to the wrist landmark coordinates 1 as adjacent joints.

[0076] In this way, using the coordinates of the 21 landmarks corresponding to the joints of the hand, for example, as shown in Fig. 10 , if the horizontal direction in the figure is set as the x-axis, the vertical direction in the figure as the y-axis, and the z-axis perpendicular to the plane of the paper in Fig. 10 is set, when the z-coordinates of landmark coordinates 2 to 5 of the thumb and landmark coordinates 18 to 21 of the little finger are smaller than landmark coordinates 10 to 13 of the middle finger, the recognition unit 152 can recognize that the tip of the little finger and the tip of the thumb in Fig. 10 are protruding forward from the center of the palm RH101 relative to the plane of the paper in Fig. 10. In this case, it can generally be considered that the front side of the plane of the paper in Fig. 10 is the palm (front side), and the back of the hand is on the back side of the plane of the paper in Fig. 10, and therefore the recognition unit 152 can recognize that the palm (front side) is displayed on the plane of the paper.

[0077] Furthermore, as shown in FIG. 10, when landmark coordinates 2 to 5 of the thumb are on the right side and landmark coordinates 18 to 21 of the little finger are on the left side, the recognition unit 152 can recognize that the thumb is on the right side and the little finger is on the left side on the front side of the palm RH101, and therefore the hand imaged in FIG. 10 is the palm of a right hand.

[0078] Conversely, although not shown, if thumb landmark coordinates 2 to 5 are on the left side and little finger landmark coordinates 18 to 21 are on the right side, and the z coordinates of thumb landmark coordinates 2 to 5 and little finger landmark coordinates 18 to 21 are smaller than middle finger landmark coordinates 10 to 13, then the thumb is on the left side and the little finger is on the right side of the front of the palm, and recognition unit 152 can recognize that the captured hand is the palm (front side) of a left hand.

[0079] Furthermore, although not shown, when the z coordinates of the thumb landmark coordinates 2 to 5 and the little finger landmark coordinates 18 to 21 are larger than the middle finger landmark coordinates 10 to 13, it can be assumed that the palm (front) is on the back side of the paper, and therefore the recognition unit 152 can recognize that the back of the hand is displayed on the front side of the paper.

[0080] As a result, when the thumb landmark coordinates 2 to 5 are on the right side and the little finger landmark coordinates 18 to 21 are on the left side, the thumb is on the right side and the little finger is on the left side on the back of the hand, so the recognition unit 152 can recognize that the imaged hand is the back of the left hand.

[0081] Furthermore, when the thumb landmark coordinates 2 to 5 are on the left side and the little finger landmark coordinates 18 to 21 are on the right side, the thumb is on the left side and the little finger is on the right side on the back of the hand, and therefore the recognition unit 152 can recognize that the captured hand is the back of the right hand.

[0082] As a result, for example, when cameras 111L and 111R capture images of a state in which the left and right hands are crossed, as shown in the left part of Fig. 11, the hand RH21 in image PL21 captured by camera 111R can be recognized as the right hand even if the image is captured with a field angle on the left side of user H101. Similarly, as shown in the right part of Fig. 11, the hand LH21 in image PR21 captured by camera 111L can be recognized as the left hand even if the image is captured with a field angle on the right side of user H101.

[0083] Furthermore, for example, when the left and right hands are aligned and extended to the left in front of the user H101 and cameras 111L and 111R each capture an image, as shown in the left part of Figure 12, the hands LH31 and RH31 in the image PL31 captured by camera 111R can be recognized as the left and right hands, respectively, even though the image is taken from a field of view on the left side of the user H101.

[0084] In this case, the image PR31 captured by the camera 111L will not contain either the left or right hand, as shown in the right part of Figure 12, so the recognition unit 152 may infer that there is a possibility that both hands are present in the image PL31, since there are no hands in the image PR31.

[0085] In either case, it can be recognized that both hands are present on the left side of the front of the user H101.

[0086] Conversely, for example, when the left and right hands are aligned and extended to the right in front of the user H101 and cameras 111L and 111R capture images, as shown in the right part of Figure 13, the hands LH41 and RH41 in the image PR41 captured by camera 111L can be recognized as the left and right hands, respectively, even though the image is taken from a field of view on the right side of the user H101.

[0087] In this case, since neither the left nor the right hand is present in the image PL41 captured by the camera 111R, the recognition unit 152 may infer that both hands may be present in the image PR41 since there are no hands in the image PL41.

[0088] In either case, it can be recognized that both hands are present on the right side of the front of the user H101.

[0089] This allows, for example, for the avatar's hands to be set to coordinates corresponding to landmark coordinates, making it possible to estimate the hand movements of user H101 from images captured by cameras 111R and 111L, and further to reflect these movements in the avatar's hands.

[0090] For example, when an image P51 showing a gesture of sticking up only the thumb as shown in the left part of FIG. 14 is captured, the corresponding landmark coordinates are obtained as shown in image P52.

[0091] Furthermore, by placing the landmark coordinates at the corresponding coordinate positions of the avatar's hand based on the respective positional relationships of the landmark coordinates as shown in image P52, it is possible to reflect the hand movement of user H101 as the hand movement of the avatar, as shown in image P53 on the right side of Figure 14.

[0092] The recognition unit 152 estimates hand movements from the time series changes in landmark coordinates representing hand movements, which are generated by the preprocessing units 151 through image-based preprocessing, and recognizes the estimated results as hand movements.

[0093] The recognition unit 152 may, for example, formulate the time-series changes in landmark coordinates to find the pattern of hand movement.

[0094] The recognition unit 152 may also be configured to be a deep neural network (DNN) that receives time-series landmark coordinate values ​​as input and outputs the estimated results of hand movements.

[0095] Note that the estimation of hand movement in the recognition unit 152 may be obtained from the time series changes in the landmark coordinates described above, but it may also be possible to extract changes in distance between landmark coordinates, etc. as time series feature quantities, estimate hand movement from the changes in the time series feature quantities, and use the estimation result as the estimation result of hand movement.

[0096] For example, as shown in image P61 in Fig. 15, the distance ED between landmark coordinate 5, which is the tip of the thumb, and landmark coordinate 9, which is the tip of the index finger, can be calculated from the respective coordinates. Similarly, the distances between other joints can also be calculated from the landmark coordinates.

[0097] Therefore, the recognition unit 152 may estimate the hand movement based on the time-series change in the distance between each joint. In this case, the preprocessing unit 151 may obtain the distance between each joint as a preprocessing step.

[0098] Note that the above description has been given of an example in which the preprocessing unit 151 obtains landmark coordinates, which are the positions of each joint of the hand, from images captured by the cameras 111L and 111R, and then the recognition unit 152 detects hand movements based on the landmark coordinates. However, the recognition unit 152 may be configured using a DNN or the like, and perform machine learning using learning data consisting of pairs of images of the hand region extracted by semantic segmentation or the like and correct hand movements before the landmark coordinates are obtained, so that the recognition unit 152 directly estimates hand movements from the images of the hand region extracted by semantic segmentation or the like, and uses the estimation results as hand movement recognition results.

[0099] Furthermore, the recognition unit 152 may be configured to directly estimate hand movements without preprocessing by the preprocessing unit 151, based on training data in which images captured by the cameras 111L and 111R are paired with correct hand movements as learning data for machine learning. Furthermore, the recognition unit 152 may estimate hand movements using both images captured by the cameras 111L and 111R and landmark coordinates as inputs through machine learning, and recognize the hand movements based on the estimation results. The recognition unit 152 may recognize information corresponding to landmark coordinates as hand movements, but may also recognize poses and gestures including body movements other than the hands that are linked to the hand movements as hand movements, and use the poses and gestures that are the recognition results to reflect them in the movements of the avatar corresponding to the user H101.

[0100] However, in this specification, the explanation will be given assuming that the pre-processing unit 151 detects the hand area from the image captured by the cameras 111L and 111R using semantic segmentation or the like, obtains the positions of multiple joints that make up the hand from the image of the detected hand area as landmark coordinates, and the recognition unit 152 estimates the hand movement from the landmark coordinates.

[0101] <Method of Recognizing Mouth Movement> Next, a description will be given of the preprocessing performed by the preprocessing unit 151 to recognize (estimate) mouth movement, and a method of recognizing (estimating) mouth movement performed by the recognition unit 152 based on the results of the preprocessing.

[0102] The pre-processing unit 151 identifies the area around the mouth of the user H101 based on the images captured by the cameras 111L and 111R, for example, by semantic segmentation.

[0103] Next, the preprocessing unit 151 detects the position of the epidermis, which is set at a predetermined position, as three-dimensional landmark coordinates such as (x, y, z), based on the image of the identified area around the mouth of the user H101. The position of the epidermis here corresponds to the joints in the hand movement described above.

[0104] Here, for example, it is assumed that landmark coordinates of 20 points near the lips are set as landmark coordinates for detecting the movement of the mouth, as shown in FIG.

[0105] In addition, in FIG. 16, landmark coordinates of positions on the skin of the user H101, which correspond to coordinates on the skin of the face set for the avatar, are expressed.

[0106] 16, landmark coordinates 1 to 17 are the positions of the epidermis that form the contour of the face around the lower jaw of user H101, landmark coordinates 18 to 22 are the positions of the epidermis that form the eyebrow of user H101's right eye, and landmark coordinates 23 to 27 are the positions of the epidermis that form the eyebrow of user H101's left eye. Furthermore, landmark coordinates 37 to 42 are the positions of the epidermis that form user H101's right eye, and landmark coordinates 43 to 48 are the positions of the epidermis that form user H101's left eye. Furthermore, landmark coordinates 28 to 31 are the positions of the epidermis that form the bridge of user H101's nose, and landmark coordinates 32 to 36 are the positions of the epidermis that form user H101's ala (bottom of the nose). Landmark coordinate 34 corresponds to the tip of the nose NT.

[0107] Furthermore, landmark coordinates 49 to 60 are the positions of the epidermis that form the outer contour of the lips of the user H101, and landmark coordinates 61 to 68 are the positions of the epidermis that form the contour of the opening of the lips of the user H101. Note that the positions of the epidermis that form the contour of the opening of the lips of the user H101, which are expressed by the landmark coordinates 61 to 68, are based on the positions of the epidermis when the mouth is closed.

[0108] Here, cameras 111R and 111L are provided at a position that abuts the nose tip NT of HMD 101, and the left and right angles of view are set to intersect so that each angle of view includes the area including the mouth LP. Therefore, of the landmark coordinates near the lips described above, only the positions of the epidermis of user H101's face that correspond to landmark coordinates 3 to 15, 32 to 36, and 49 to 68 within the rectangular dashed frame Z1 can be imaged.

[0109] Furthermore, what is necessary to detect the movement of the mouth are the landmark coordinates 49 to 68 within the rectangular solid frame Z2.

[0110] As a result, for example, as shown in image P71 in the upper right corner of Figure 16, when the lips LP101 are closed, landmark coordinates 49 to 60 are detected along the outer shape LOP101 indicated by the dashed line in the figure, and landmark coordinates 61 to 68 are detected along the outer shape LIP101 indicated by the dotted line in the figure.

[0111] In contrast, for example, as shown in image P72 in the lower right of Figure 16, when lips LP102 are open, landmark coordinates 49 to 60 are detected along the outer shape LOP102 indicated by the dashed line in the figure, and landmark coordinates 61 to 68 are detected along the outer shape LIP102 indicated by the dotted line in the figure.

[0112] In this way, the movement of the mouth can be detected according to the shapes of the landmark coordinates 49 to 60 and the landmark coordinates 61 to 68.

[0113] Based on the landmark coordinates of the mouth thus determined, for example, as shown in the upper left of Figure 17, camera 111L captures an image P81 of the left side of the mouth of user H101 with his mouth open, thereby identifying the position of the epidermis of user H101's lips.

[0114] Based on this, by making the position of the corresponding coordinates of the avatar's lips the position of the epidermis of the lips of the identified user H101, it is possible to reflect the avatar's mouth as being open, just like the user H101 captured in image P81, as shown in image P82.

[0115] Similarly, for example, as shown in the lower left of Figure 17, camera 111L captures an image P83 of the left side of the mouth of user H101 with their mouth closed, allowing the position of the epidermis of user H101's lips to be identified.

[0116] Based on this, by making the position of the corresponding coordinates of the avatar's lips the position of the epidermis of the identified user H101's lips, the avatar's mouth can be reflected to be in a closed state, similar to that of user H101 captured in image P83, as shown in image P84.

[0117] 17 shows an example in which the recognition unit 152 reflects on the avatar information recognized based only on the image captured by the camera 111L capturing the left side of the mouth, but the position of the lip epidermis can also be estimated using only the image captured by the camera 111R capturing the right side of the mouth. Naturally, it is also possible to use both the images captured by the cameras 111L and 111R, and the more images used, the higher the accuracy of detecting landmark coordinates.

[0118] The recognition unit 152 estimates mouth movement from landmark coordinates representing mouth movement, which are generated by the preprocessing unit 151 through image-based preprocessing, and recognizes the estimated result as mouth movement.

[0119] The recognition unit 152 may, for example, formulate the time-series changes in landmark coordinates to find the pattern of mouth movement.

[0120] The recognition unit 152 may also be configured with a DNN (Deep Neural Network), and may, for example, be configured to input the landmark coordinates in a time series, output the mouth movements, and recognize the output result as the mouth movements through machine learning using learning data consisting of the landmark coordinates in a time series and the mouth movements that are the correct answers.

[0121] The estimation of mouth movement in the recognition unit 152 may be performed from the time series changes in the landmark coordinates described above, but it may also be performed by extracting changes in the distance between landmark coordinates, etc. as time series feature quantities, estimating mouth movement from the changes in the time series feature quantities, and using the estimation result as the estimated result of mouth movement.

[0122] Here, since the shape of the lips changes in a complex manner, for example, as shown in image P91 of Figure 18, the horizontal length L1 and vertical length L2 of the outer shape represented by landmark coordinates 49 to 60 and landmark coordinates 61 to 68 are calculated, and the area S1 is calculated from the product of the two, thereby extracting them as time-series feature quantities representing the shape of the mouth, and the recognition unit 152 can estimate the movement of the mouth based on the feature quantities.

[0123] Similarly to hand movements, mouth movements may be estimated from time-series changes in the distance between landmark coordinates.

[0124] Therefore, the recognition unit 152 may estimate the movement of the mouth based on the above-mentioned feature amounts and time-series changes in the distance between the positions of the lip surface. In this case, the feature amounts and the distance between the positions of the lip surface may be obtained by the preprocessing unit 151 as preprocessing.

[0125] The recognition unit 152 may also estimate the mouth movements in units of action units, which are specific mouth movements grouped into units.

[0126] There are various types of action units, but for example, as shown in Figure 19, an action unit (smile) can be one that expresses a series of mouth movements from a closed mouth state as shown in image P111 to a smiling state with teeth exposed as shown in image P112.

[0127] Also, as shown in FIG. 20, there is an action unit (LowerLipDpressor) that expresses a series of mouth movements from a closed mouth state as shown in image P121 to lowering the lower lip as shown in image P122.

[0128] Furthermore, as shown in FIG. 21, there is an action unit (ChinRaiser) that expresses a series of mouth movements from a closed mouth state as shown in image P131 to raising the lower jaw as shown in image P132.

[0129] Similarly, as shown in FIG. 22, there is an action unit (Lip Puckerer) that expresses a series of mouth movements from a closed mouth state as shown in image P141 to pursing and protruding lips as shown in image P142.

[0130] As shown in FIG. 23, there is an action unit (LipSuck) that expresses a series of mouth movements from a closed mouth state as shown in image P151 to pulling the lips inward as shown in image P152.

[0131] As shown in FIG. 24, there is an action unit (SmileWithMouthCorner) that expresses a series of mouth movements from a closed mouth state as shown in image P161 to a toothless smile as shown in image P162.

[0132] As shown in FIG. 25, there is an action unit (Mouth Stretch) that expresses a series of mouth movements from a closed mouth state as shown in image P171 to an open mouth state as shown in image P172.

[0133] As shown in FIG. 26, there is an action unit (CheekPuff) that expresses a series of mouth movements from a closed mouth state as shown in image P181 to puffing out cheeks as shown in image P182.

[0134] As shown in FIG. 27, there is an action unit (LipStretch) that expresses a series of mouth movements from a closed mouth state as shown in image P191 to pulling the lips to the side as shown in image P192.

[0135] As shown in FIG. 28, there is an action unit (Digust) that expresses a series of mouth movements from a closed mouth state as shown in image P201 to an expression in which the lips are angled to express disgust as shown in image P202.

[0136] In the above, we have described an example in which the pre-processing unit 151 determines landmark coordinates representing the position of the lip epidermis from images captured by the cameras 111L and 111R, and then the recognition unit 152 recognizes the movement of the mouth based on the landmark coordinates.

[0137] However, the recognition unit 152 may be configured using a DNN or the like, and machine learning may be performed using training data consisting of pairs of images of the mouth area extracted by semantic segmentation or the like and correct mouth movements before the landmark coordinates are determined, so that the recognition unit 152 can directly estimate mouth movements from images of the mouth area extracted by semantic segmentation or the like, and use the estimation results as the recognition result of the mouth movement.

[0138] Furthermore, the learning data for machine learning in the recognition unit 152 may be based on teacher data in which images captured by the cameras 111L and 111R are paired with correct mouth movements, so that the recognition unit 152 can directly estimate mouth movements without preprocessing by the preprocessing unit 151. Furthermore, the recognition unit 152 may estimate mouth movements by machine learning using both the images captured by the cameras 111L and 111R and landmark coordinates as input, and recognize the mouth movements based on the estimation results.

[0139] However, in this specification, the explanation will be given assuming that the pre-processing unit 151 detects the mouth area from the images captured by the cameras 111L and 111R using semantic segmentation or the like, determines the positions of the multiple epidermis that make up the lips from the image of the detected mouth area as landmark coordinates, and the recognition unit 152 estimates the movement of the mouth from the landmark coordinates.

[0140] <Image Processing Unit> Next, image processing by the image processing unit 153 based on the recognition result of the recognition unit 152 will be described with reference to FIG.

[0141] The image processing unit 153 reflects the recognition result of the recognition unit 152 on the avatar and synthesizes it with the background of the virtual space to generate a virtual space image, which is displayed on the display 112 .

[0142] More specifically, for example, as shown in FIG. 29, the preprocessing unit 151 generates information P211 consisting of landmark coordinates indicating the position of the lip epidermis for estimating mouth movement based on images from the cameras 111R and 111L, and supplies the information P211 to the recognition unit 152.

[0143] Furthermore, the preprocessing unit 151 generates information P213 consisting of landmark coordinates indicating the positions of the hand joints for estimating the hand movement based on the images from the cameras 111R and 111L, and supplies the information P213 to the recognition unit 152.

[0144] The recognition unit 152 estimates the movement of the mouth based on information P211 consisting of landmark coordinates indicating the position of the lip epidermis, and supplies the estimated information on the movement of the mouth to the image processing unit 153 as the recognition result of the movement of the mouth.

[0145] The recognition unit 152 estimates the hand movement based on information P213 consisting of landmark coordinates indicating the positions of the hand joints, and supplies the estimated hand movement information to the image processing unit 153 as the hand movement recognition result.

[0146] Based on the estimated mouth movement results supplied from the recognition unit 152, the image processing unit 153 adjusts the position of the avatar's facial skin corresponding to the position of the user H101's facial skin, and generates an avatar face image P212 that reflects the user's mouth movement.

[0147] Based on the estimated hand movement results supplied from the recognition unit 152, the image processing unit 153 adjusts the positions of the avatar's hand joints to correspond to the positions of the user H101's hand joints, and generates an image P214 of the avatar's hand and the surrounding upper body.

[0148] Furthermore, the image processing unit 153 estimates the pose (movement) of the entire body based on the generated upper body image P214, and generates a non-face image P215 using the estimation result. Note that the image processing unit 153 may also generate the non-face image P215 directly based on the estimation result of the hand movement supplied from the recognition unit 152.

[0149] The image processing unit 153 then synthesizes the generated face image P212 with a non-face image P215 to generate an avatar image P216.

[0150] In addition, when generating a face image P212 and an upper body image P214 (including an image P215 other than a face image) based on the estimated results of mouth movement and hand movement, the image processing unit 153 may use a generation model that uses the estimated results of mouth movement and hand movement as input to generate the face image P212 and the upper body image P214 (as well as an image P215 other than a face image).

[0151] Through the above process, an avatar image that reflects the user's facial expression, pose, gestures, etc. is generated based on the mouth and hand movements of the user H101.

[0152] When the image processing unit 153 is supplied with the above-mentioned action units as the estimation results of the mouth movement, the image processing unit 153 may estimate the emotions of the user H101 from the names set in the action units, and reflect these emotions in the avatar as eye expressions, body poses, or gestures that cannot be captured by the cameras 111R and 111L.

[0153] In this case, the image processing unit 153 may generate an image of the avatar that reflects the estimated emotion by adding movement to the facial contours, eyes, eyebrows, etc., by controlling the coordinates set on the avatar's face that correspond to the landmark coordinates 1 to 48 shown in FIG. 16 based on the estimated emotion.

[0154] <Avatar Reflection Processing (Part 1)> Next, with reference to the flowchart in FIG. 30, an avatar reflection processing will be described in which the mouth movements and hand movements of the user H101 by the HMD 101 in FIG. 8 are reflected in an avatar and displayed in an image in the virtual space.

[0155] In step S31, cameras 111L and 111R capture images of their own angles of view FOV-R and FOV-L, respectively, to capture images of the left and right mouths of user H101 wearing HMD 101, and the range in which the hands are located in a typical crossing motion in the forward direction of user H101, and output these images to pre-processing unit 151 of control unit 131.

[0156] In step S32, the preprocessing unit 151 identifies the regions of the hands and lips by segmentation based on the image, and recognizes the positions of the hands and lips within the image.

[0157] In step S33, the preprocessing unit 151 detects landmark coordinates of the hands and mouth based on the recognized positions of the hands and lips, and sends the detected landmark coordinates to the recognition unit 152.

[0158] In step S34, the recognition unit 152 estimates (detects) the hand movements and mouth movements of the user H101 based on the landmark coordinate information of the hands and mouth, and supplies the results to the image processing unit 153 as recognition results.

[0159] In step S35, the image processing unit 153 generates an avatar image that reflects the hand and mouth movements of the avatar based on the hand and mouth movements of the user H101, which are the recognition results, and uses the generated avatar image to generate an image in the virtual space and display it on the display 112.

[0160] At this time, the image processing unit 153 may control the communication unit 135 to transmit an image of an avatar that reflects the hand movements and mouth movements of the user H101 to the HMD 101 worn by another user corresponding to another avatar that communicates with the avatar of the user H101.

[0161] This processing makes it possible to display an image of an avatar that reflects the hand and mouth movements of user H101 on an HMD101 worn by another user with whom user H101 communicates in a virtual space via an avatar.

[0162] In addition, the image processing unit 153 may acquire images of avatars that reflect the hand and mouth movements of other users, which are supplied from the HMD 101 worn by other users with whom the user H101 communicates in the virtual space via their avatars, and may generate an image of the virtual space by combining the images of the avatar of the user H101 and the avatars corresponding to the other users, and display the image on the display 112.

[0163] In step S36, it is determined whether or not an instruction to end the process has been given. If an instruction to end the process has not been given, the process returns to step S31, and the subsequent steps are repeated.

[0164] Then, in step S36, when an instruction to end the process is given, the process ends.

[0165] Through the above series of processes, user H101 and other users who communicate with user H101 in the virtual space via avatars can communicate using avatars that reflect their respective hand movements and mouth movements.

[0166] In this case, as described above, cameras 111L and 111R each capture images of user H101's mouth from the left and right, respectively, and also capture images of the area further back than the mouth, crossing left and right, thereby making it possible to capture images of both the left and right hands.

[0167] This allows the mouth and both hands to be captured by the common sensors, cameras 111L and 111R, making it possible to estimate the mouth and hand movements in real time and in a synchronized manner. As a result, it is no longer necessary to provide separate sensors for acquiring the mouth and hand movements, making it possible to reduce the size and cost of the device configuration.

[0168] Furthermore, since the estimated results of mouth movement and the estimated results of hand movement are acquired in a synchronized state, there is no need to synchronize the respective estimated results, and therefore no configuration or processing for synchronization is required, making it possible to estimate mouth movement and hand movement in a synchronized state with a simple configuration and simple processing.

[0169] Furthermore, since the mouth movements and hand movements can be estimated in a synchronized manner using a simple configuration and simple processing, it is possible to improve the accuracy of estimating the hand movements and mouth movements of the user H101, and it becomes possible to reflect the user's facial expressions and poses in the avatar with higher accuracy.

[0170] Furthermore, detecting hand movements no longer requires the wearing of motion sensors, which was previously required, eliminating the inconvenience of restricting hand movements by wearing a motion sensor.

[0171] Furthermore, although the above has described an example in which cameras 111R and 111L are optical RGB cameras, cameras 111R and 111L may also be configured to include a depth sensor capable of measuring distance on a pixel-by-pixel basis or LiDAR (Light Detection and Ranging, Laser Imaging Detection and Ranging).

[0172] Since cameras 111R and 111L are configured with depth sensors, LiDAR, etc., they are also capable of distance measurement, making it possible to detect landmark coordinates with higher accuracy and improve the accuracy of estimating mouth and hand movements, making it possible to more appropriately reflect the user's facial expressions and poses in the avatar.

[0173] In addition, the functions realized by the pre-processing unit 151, the recognition unit 152, and the image processing unit 153 in the control unit 131 may be realized by a cloud server on the network, and may be realized as a whole by an information processing system consisting of the cloud server and the HMD 101.

[0174] In this case, the HMD 101 may, for example, control the communication unit 135 to send the captured image to a cloud server on the network, where the mouth and hand movements of the user H101 are estimated and reflected in the avatar. After a series of processes up to generating an image in the virtual space are executed, an image in the virtual space including an avatar reflecting the mouth and hand movements of the user H101 may be acquired and displayed.

[0175] By using a cloud server in this way, the processing load on the HMD 101 is reduced, and the above-mentioned series of processes can be realized even in a configuration where the processing power of the control unit 131 is low, for example, so that the user's facial expressions and poses can be more appropriately reflected in the avatar at lower cost.

[0176] Furthermore, in the above, an example has been described in which, when user H101 puts on the HMD 101, cameras 111R and 111L are provided near the tip of the nose where the HMD 101 comes into contact, and the movements of the user's mouth and hands are detected based on the captured images and reflected and displayed on an avatar.

[0177] However, if it is not assumed that the user himself / herself will view his / her own avatar, but it is sufficient if the user can appear as an avatar in the images of the virtual space viewed by other users, the images of the virtual space that are displayed by detecting mouth and hand movements and reflecting them in the avatar only need to be viewable on the HMD worn by the other users, and therefore display 112 is no longer a required component.

[0178] Therefore, in such a case, it is possible to simply install cameras corresponding to cameras 111L and 111R on eyewear worn on the eyes, such as glasses, instead of on the HMD 101, for example, on the lower part of the rim where the pad that is placed on the tip of the nose or the lens is attached.

[0179] In this case, for example, images captured by the cameras 111L and 111R attached to the glasses are transmitted to a cloud server, etc. This allows the cloud server to detect the hand and mouth movements of the user wearing the glasses, reflect them in an avatar, generate an image of the virtual space, and distribute it to the HMDs 101 worn by other users.

[0180] In this way, if the user whose hand and mouth movements are detected does not need to view images in the virtual space but only needs to appear as an avatar in the virtual space viewed by other users, it is sufficient to provide cameras 111L and 111R in eyewear such as glasses and transmit the images to the cloud server. Also, the cameras 111L and 111R may be wide-angle cameras, which can widen the range in which stereo imaging is possible and enable depth information of the user's hands and surroundings to be acquired over a wide range.

[0181] <<3. Modification of the First Embodiment>> In the above, an example has been described in which the cameras 111R and 111L provided in the HMD 101 capture images in the range from the tip of the nose NT down, including the area around the mouth, with the angle of view set to intersect left and right, thereby detecting mouth and hand movements and reflecting the detected movements in the mouth and hand movements of an avatar corresponding to the user H101 in the virtual space.

[0182] However, in order to reflect the self-position in the real space as the self-position in the virtual space, the HMD 101 is often provided with a separate sensor for detecting the self-position in the real space.

[0183] For example, as shown in FIG. 31 , cameras 111sL and 111sR may be provided in addition to cameras 111L and 111R, each of which may comprise an image sensor for capturing images in the front direction and a depth sensor, to capture ranging images of the surroundings, and the vehicle's own position may be estimated using SLAM (Simultaneous Localization and Mapping) based on the captured images.

[0184] In the present disclosure, images of the area in front of, to the left and right of, and below the nose tip NT of the user H101 captured by the cameras 111R and 111L can be acquired. Therefore, the accuracy of self-location estimation may be improved by estimating the self-location by SLAM by combining the images captured by the cameras 111sL and 111sR provided for self-location estimation with the images captured by the cameras 111R and 111L.

[0185] <<4. Second embodiment>> In the above, an example has been described in which cameras 111L, 111R are provided on the left and right of the part of the HMD 101 that comes into contact with the nose tip NT, and the mouth is imaged from the left and right, and the area from the mouth to the hands is imaged so that the left and right images intersect, thereby simultaneously detecting mouth movement and hand movement.

[0186] However, if the length of the hand of the user H101 wearing the HMD 101 varies depending on the physique, etc., there is a risk that the hand may not be captured clearly with a uniform angle of view or focal length.

[0187] Therefore, cameras 111L and 111R may be provided with a focus adjustment mechanism and a mechanism for adjusting the angle of view so that images can be captured to suit the physique of user H101 wearing HMD 101, thereby enabling clear images to be captured regardless of the physique of user H101, and enabling mouth and hand movements to be detected with higher accuracy.

[0188] FIG. 32 shows an example of the configuration of an HMD in which the cameras 111L and 111R are provided with a focus adjustment function and a mechanism for adjusting the angle of view.

[0189] In the HMD 101' in FIG. 32, the same functions as those of the HMD 101 in FIG. 3 are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.

[0190] The HMD 101' in Figure 32 differs from the HMD 101 in Figure 3 in that cameras 111'L and 111'R are provided instead of cameras 111L and 111R, and further, focus adjustment mechanisms 111'L-AF and 111'R-AF and actuators 201L and 201R are provided, respectively.

[0191] The cameras 111'L and 111'R have the same basic functions as the cameras 111L and 111R, but are provided with focus adjustment mechanisms 111'L-AF and 111'R-AF, respectively, to enable focus adjustment.

[0192] Therefore, even if the length of the hand changes depending on the user's physique, it is possible to capture an image of the hand with the focus adjusted appropriately.

[0193] The cameras 111'L and 111'R are provided with actuators 201L and 201R, respectively, which enable the optical axes of the cameras 111'L and 111'R to be changed.

[0194] Therefore, as shown in the upper left of Figure 32, for example, the camera 111'L and the focus adjustment mechanism 111'L-AF can change the optical axis of the camera 111'L to optical axes Axa to Axc, as shown by actuators 201La to 201Lc, in accordance with the movement of the actuator 201L.

[0195] As a result, the angle of view of the camera 111'L can be changed from FOV-La' to FOV-La, FOV-Lb' to FOV-Lb, and FOV-Lc' to FOV-Lc in accordance with changes in the actuators 201La to 201Lc.

[0196] Although the movement of the actuator 201R of the camera 111'R is not shown, a similar function can change the angle of view from FOV-Ra' to FOV-Ra, FOV-Rb' to FOV-Rb, and FOV-Rc' to FOV-Rc.

[0197] Note that Figure 32 only shows an example of changing the angle of view in the horizontal direction, but the movement of actuators 201L and 201R, although not shown, can change the angle of view not only horizontally but also forward and backward relative to the front direction of user H101.

[0198] <Example of Hardware Configuration of HMD in Fig. 32 (Part 1)> Next, an example of the hardware configuration of the HMD 101' in Fig. 32 will be described with reference to Fig. 33. Note that in the HMD 101' in Fig. 33, components having the same functions as those in the hardware configuration of the HMD 101 in Fig. 8 are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate.

[0199] The hardware configuration of the HMD 101' in Fig. 33 differs from the hardware configuration of the HMD 101 in Fig. 8 in that cameras 111'L and 111'R are provided instead of cameras 111L and 111R, and further, focus adjustment mechanisms 111'L-AF and 111'R-AF are provided for each camera, as well as actuators 201L and 201R. Furthermore, an AF control unit 221 and an actuator control unit 222 are newly provided in the control unit 131 to control the focus adjustment mechanisms 111'L-AF and 111'R-AF and the actuators 201L and 201R.

[0200] The AF (Auto Focus) control unit 221 adjusts the focus by controlling the focus adjustment mechanisms (AF) 111'L-AF and 111'R-AF of the cameras 111'L and 111'R, respectively. That is, the AF (Auto Focus) control unit 221 controls the focus adjustment mechanisms (AF) 111'L-AF and 111'R-AF based on the images captured by the cameras 111'L and 111'R to adjust the focus.

[0201] The actuator control unit 222 controls the movement of the actuators 201L and 201R based on whether the mouth and hands are present within the angle of view, based on the images captured by the cameras 111'L and 111'R using the recognition unit 152, so as to maintain the state in which the mouth and hands are captured within the angle of view.

[0202] Furthermore, if a state is detected during image capture where the mouth and hands are not within the field of view, the image processing unit 153 may estimate and use the mouth and hand movements in the current frame from the mouth and hand movements in the immediately previous frame.

[0203] <Avatar Reflection Processing (Part 2)> Next, the avatar reflection processing by the HMD 101′ in Fig. 32 will be described with reference to the flowchart in Fig. 34. Note that the processing in steps S51, S52, and S54 to S57 is the same as the processing described with reference to the flowchart in Fig. 30, and therefore description thereof will be omitted.

[0204] That is, in steps S51 and S52, the cameras 111'L and 111'R capture images of the range below the tip of the nose NT of the user H101 wearing the HMD 101', including the left and right mouths below the tip of the nose NT, where the hands are located in a typical movement that crosses left and right in the forward direction of the user H101, and the pre-processing unit 151 recognizes the positions of the hands and lips by segmentation based on the images.

[0205] In step S53, the preprocessing unit 151 determines whether or not the positions of the hands and lips have been recognized. If the positions have been recognized, as described above, steps S54 to S57 are performed to display an avatar that reflects the mouth movements and hand movements of the user H101, and similar processes are repeated until an instruction to end the process is issued.

[0206] Also, in step S53, if the positions of the hands and lips cannot be recognized, that is, if there is a possibility that the focus adjustment or angle of view is not appropriate, the process proceeds to step S58.

[0207] In step S58, the AF control unit 221 controls the focus adjustment mechanisms 111'R-AF and 111'L-AF to adjust the focus positions, and the actuator control unit 222 controls the actuators 201R and 201L to adjust the angles of view of the cameras 111'R and 111'L.

[0208] In step S59, the image processing unit 153 generates an image of the avatar by reflecting the hand and mouth movements of the user H101 as a recognition result based on the image captured immediately before in a state where the positions of the mouth and hands can be recognized, in the hand and mouth movements of the avatar in the current frame, and generates an image of the avatar in the virtual space and displays it on the display 112.

[0209] In the process of step S53, the processes of steps S52, S53, S58, S59, and S57 are repeated until there is no state in which the positions of the hands and lips cannot be recognized. During this time, the angle of view and focus are adjusted, and when the positions of the hands and lips become recognizable in step S53, estimation of the movements of the hands and mouth is resumed and reflected in the avatar and in the virtual space image.

[0210] Through the above processing, the mouth and hands are captured appropriately regardless of the physique of the user H101 wearing the HMD 101', and the mouth and hand movements are estimated appropriately, reflected in the avatar with high accuracy, and presented in an image in the virtual space.

[0211] <<5. Third embodiment>> In the above, an example has been described in which focus adjustment mechanisms 111′L-AF, 111′R-AF are provided in the cameras 111′L, 111′R of the HMD 101′, and actuators 201L, 201R are provided so that the angle of view can be adjusted, thereby enabling clear images to be captured in accordance with the physique of the user H101 wearing the HMD 101′.

[0212] However, there is a risk that the mouth or hands may not be tracked due to inadvertent movements, so the focus and angle of view may be calibrated in advance according to the user's physique and stored as personal data to reduce the number of situations where tracking is not possible.

[0213] Here, the information required for calibration will be described. Note that the external configuration of an HMD that can be calibrated is similar to that of the HMD 101' in Fig. 32, and therefore its description will be omitted. However, for the sake of distinction, hereinafter, the HMD that can be calibrated will be referred to as HMD 101''.

[0214] For example, if user H101 is approximately 160 cm tall and has a standard build, and stretches both arms out to the left and right while wearing a calibration-capable HMD 101'', as shown in FIG. 35, the distance Lbh between the left and right fingertips is approximately 154 cm, and the distance Lsh from camera 111R' to the tip of the left hand is approximately 83 cm. Therefore, the distance from camera 111L' to the tip of the right hand can be considered to be approximately the same. Furthermore, the horizontal angles of view to the left and right are expressed by the angles of view FOV-L101 and FOV-R101.

[0215] Furthermore, when a user of a similar build extends both arms forward while wearing the HMD 101'', as shown in Figure 35, the distance Lfe from the cameras 111L, 111R to the fingertips is approximately 65 cm, the distance Lhe from the cameras 111L, 111R to the position of the arm directly below is approximately 25 cm, and the distance Lfh from the position of the arm directly below the cameras 111L, 111R to the fingertips is approximately 57 cm.

[0216] The forward field of view FOV-F formed by the position of the fingertip, the position of cameras 111'L and 111'R, and the position of the arm directly below cameras 111L and 111R is approximately 49 degrees, and the rear field of view FOV-B formed by the position of the shoulder, the position of cameras 111'L and 111'R, and the position of the arm directly below cameras 111L and 111R is approximately 43 degrees.

[0217] In this case, the farthest position from the cameras 111L' and 111R' can be considered to be the distance Lsh to the fingertips when the hands are spread out to the left and right.

[0218] The horizontal angles of view of the cameras 111L' and 111R' are FOV-L101 and FOV-R101, respectively. The front and rear angles of view of the cameras 111L' and 111R' are FOV-F and FOV-B, respectively.

[0219] 35 and 36 show reference values ​​for a user with an average build and a height of approximately 160 cm, but these values ​​will vary depending on the user's height and build, and the focal length and angle of view will need to be adjusted appropriately depending on the user's height and build.

[0220] Therefore, during calibration, it is necessary to adjust the focus adjustment mechanisms 111'L-AF, 111'R-AF and the actuators 201L, 201R so that the distance Lsh from the cameras 111'L, 111'R to the fingertips, the horizontal angle of view FOV-L101, FOV-R101, the forward angle of view FOV-F, and the rearward angle of view FOV-R correspond to the state when the user H101 assumes the pose shown in Figures 35 and 36 described above.

[0221] Therefore, in the present disclosure, when a user H101 wearing an HMD 101'' assumes a pose as shown in Figures 35 and 36, the focus adjustment mechanisms 111'L-AF, 111'R-AF and the actuators 201L, 201R are controlled so that the mouth and hands can be imaged, and the respective control parameters are stored as calibration data in association with personal data.

[0222] When the user H101 puts on the HMD 101'', the corresponding calibration data is read out by specifying personal data, and the focus adjustment mechanisms 111'L-AF, 111'R-AF and the actuators 201L, 201R are adjusted to capture images of the mouth and hands in accordance with the individual's physique, and the mouth and hand movements are appropriately reflected in the avatar.

[0223] <Hardware configuration example (part 2) of HMD in Fig. 32> Next, with reference to Fig. 37, an example of the hardware configuration of the HMD 101'' in Fig. 32 will be described. Note that in the HMD 101'' in Fig. 37, components having the same functions as those in the hardware configuration of the HMD 101' in Fig. 33 are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate.

[0224] The hardware configuration of the HMD 101 ″ in FIG. 37 differs from the hardware configuration of the HMD 101 ′ in FIG. 33 in that a calibration processing unit 251 is newly provided in the control unit 131 .

[0225] The calibration processing unit 251 executes calibration processing and requests the user to assume a pose as shown in Figures 35 and 36. At this time, the AF (Auto Focus) control unit 221 controls the focus adjustment mechanisms (AF) 111'L-AF and 111'R-AF of the cameras 111'L and 111'R, respectively, to adjust the focus. In addition, the actuator control unit 222 controls the movement of the actuators 201L and 201R based on whether the mouth and hands are present within the angle of view, based on the images captured by the cameras 111'L and 111'R, using the recognition unit 152, so that the mouth and hands are captured within the angle of view.

[0226] The calibration processing unit 251 acquires the control parameters of the focus adjustment mechanism (AF) 111'L-AF, 111'R-AF in the AF control unit 221 when the mouth and fingers are captured within the field of view, and the control parameters of the actuators 201L, 201R in the actuator control unit 222 as calibration data, and stores them in the memory unit 134 in association with personal data 261 such as a personal ID.

[0227] Furthermore, after calibration, the calibration processing unit 251 requests information for identifying the user, such as a personal ID, from the user. When a personal ID is input as information for identifying the user, for example, the calibration processing unit 251 reads out personal data 261 registered in association with the input personal ID from the storage unit 134, and supplies the control parameters that become the associated and stored calibration data to the AF control unit 221 and the actuator control unit 222, respectively.

[0228] The AF control unit 221 controls the focus adjustment mechanisms (AF) 111'L-AF and 111'R-AF in accordance with the control parameters, which are the calibration data. The actuator control unit 222 controls the actuators 201L and 201R based on the control parameters, which are the calibration data. As a result, the focal positions and angles of view of the cameras 111'L and 111'R are set to the values ​​set during the calibration process.

[0229] This allows user H101 to simply input his / her personal ID, etc., to capture an image of his / her mouth and hands at the appropriate focal length and angle of view set during the calibration process, thereby enabling the movements of his / her mouth and hands to be appropriately reflected in the avatar.

[0230] <Calibration Processing> Next, the calibration processing in the HMD 101'' in FIG. 37 will be described with reference to the flowchart in FIG.

[0231] In step S71, the calibration processing unit 251 presents information on the display 112 requesting a pose such as that shown in Figures 35 and 36 in order to acquire calibration data, and controls the cameras 111'L and 111'R to capture images and supply the captured images to the pre-processing unit 151.

[0232] In step S72, the preprocessing unit 151 recognizes the positions of the hands and lips by segmentation based on the images captured by the cameras 111'L and 111'R.

[0233] In step S73, the pre-processing unit 151 determines whether the mouth and hands can be detected from the positions of the hands and lips. If the mouth and hands cannot be detected, that is, if the focus adjustment or angle of view may not be appropriate, the process proceeds to step S74.

[0234] In step S74, the AF control unit 221 controls the focus adjustment mechanisms 111'R-AF and 111'L-AF to adjust the focus positions. The actuator control unit 222 controls the actuators 201R and 201L to adjust the angles of view of the cameras 111'R and 111'L. Then, the process returns to step S71.

[0235] That is, the processes of steps S71 to S74 are repeated until the positions of the hands and lips are recognized by segmentation based on the images captured by the cameras 111'L and 111'R, and the mouth and hands can be detected.

[0236] If it is determined in step S73 that the hand and mouth have been detected, that is, if it is determined that the focal length and angle of view have been set appropriately, the process proceeds to step S75.

[0237] In step S75, the calibration processing unit 251 acquires the control parameters of the focus adjustment mechanisms (AF) 111'L-AF, 111'R-AF in the AF control unit 221 at this time and the control parameters of the actuators 201L, 201R in the actuator control unit 222 as calibration data, associates it with personal data 261 such as a personal ID, stores it in the memory unit 134, and terminates the processing.

[0238] Through the above processing, the calibration processing is performed, and the control parameters of the focus adjustment mechanisms (AF) 111'L-AF, 111'R-AF in the AF control unit 221 in a state where it is determined that the focal length and angle of view are appropriately set, and the control parameters of the actuators 201L, 201R in the actuator control unit 222 are acquired as calibration data, and are stored in the memory unit 134 in association with personal data 261 such as a personal ID.

[0239] <Avatar Reflection Processing (Part 3)> Next, the avatar reflection processing by the HMD 101'' in FIG. 37 will be described with reference to the flowchart in FIG. 39. Note that the processing in steps S93 to S101 is similar to the processing in steps S51 to S59 described with reference to the flowchart in FIG. 34, and therefore description thereof will be omitted.

[0240] In step S91, the calibration processing unit 251 requests a personal ID or the like from the user, reads out from the storage unit 134 personal data 261 registered in association with the personal ID entered by the user, and supplies the control parameters that constitute the associated and stored calibration data to the AF control unit 221 and the actuator control unit 222, respectively.

[0241] In step S92, the AF control unit 221 controls the focus adjustment mechanisms (AF) 111'L-AF and 111'R-AF according to the control parameters, and the actuator control unit 222 controls the actuators 201L and 201R, thereby achieving the state during calibration processing.

[0242] Thereafter, similar to the process described with reference to the flowchart in Figure 34, cameras 111'L and 111'R capture images of the ranges on the opposite sides of the mouth so that the mouth and hands intersect with each other at the mouth, and based on the captured images, the movements of the mouth and hands are estimated. Based on the estimation results, an image of the virtual space reflecting the avatar's mouth movements and hand movements is generated and displayed on display 112.

[0243] Through the above process, user H101 can capture images of his or her mouth and hands at the appropriate focal length and angle of view set during the calibration process simply by entering his or her personal ID, etc., thereby enabling the movements of his or her mouth and hands to be appropriately reflected in the avatar.

[0244] <<6. Example of Execution by Software>> The above-described series of processes can be executed by hardware, but can also be executed by software. When the series of processes is executed by software, the program constituting the software is installed from a recording medium into a computer incorporated in dedicated hardware, or into, for example, a general-purpose computer that can execute various functions by installing various programs.

[0245] 40 shows an example of the configuration of a general-purpose computer. This computer has a built-in CPU (Central Processing Unit) 1001. An input / output interface 1005 is connected to the CPU 1001 via a bus 1004. A ROM (Read Only Memory) 1002 and a RAM (Random Access Memory) 1003 are connected to the bus 1004.

[0246] The input / output interface 1005 is connected to an input unit 1006 including input devices such as a keyboard and a mouse through which a user inputs operation commands, an output unit 1007 that outputs a processing operation screen and images of processing results to a display device, a storage unit 1008 including a hard disk drive or the like that stores programs and various data, and a communication unit 1009 including a LAN (Local Area Network) adapter or the like that executes communication processing via a network typified by the Internet. Also connected is a drive 1010 that reads and writes data from / to a removable storage medium 1011 such as a magnetic disk (including a flexible disk), an optical disk (including a CD-ROM (Compact Disc-Read Only Memory) and a DVD (Digital Versatile Disc)), a magneto-optical disk (including an MD (Mini Disc)), or a semiconductor memory.

[0247] The CPU 1001 executes various processes in accordance with a program stored in a ROM 1002 or a program read from a removable storage medium 1011 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, installed in a storage unit 1008, and loaded from the storage unit 1008 into a RAM 1003. The RAM 1003 also stores data necessary for the CPU 1001 to execute various processes as appropriate.

[0248] In a computer configured as described above, the CPU 1001 performs the above-described series of processes by, for example, loading a program stored in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.

[0249] The program executed by the computer (CPU 1001) can be provided by being recorded on a removable storage medium 1011 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.

[0250] In a computer, a program can be installed in the storage unit 1008 via the input / output interface 1005 by inserting a removable storage medium 1011 into the drive 1010. The program can also be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the storage unit 1008. Alternatively, the program can be installed in advance in the ROM 1002 or the storage unit 1008.

[0251] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0252] 40 realizes the functions of the control unit 131 in FIGS. 8, 33, and 37.

[0253] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device with multiple modules housed in a single housing, are both systems.

[0254] Furthermore, the embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure.

[0255] For example, the present disclosure can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.

[0256] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0257] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0258] The present disclosure may also have the following configurations. <1> An information processing device comprising: two image capture units near a user's nose tip, at corresponding positions on the left and right of the nose tip, for capturing an image of an area below the nose tip, wherein the two image capture units each capture an image of an area including the user's mouth at different angles of view, and the different angles of view intersect. <2> The information processing device described in <1>, wherein the different angles of view intersect with each other such that one angle of view captures the mouth from the left side and includes a range farther than the mouth on the right side in relation to a front direction of the user, and the other angle of view captures the mouth from the right side and includes a range farther than the mouth on the left side in relation to a front direction of the user. <3> The information processing device described in <2>, wherein a left end of one angle of view and a right end of the other angle of view overlap. <4> The information processing device according to <1>, wherein at least one of the two imaging units captures an image of the user's hands in a range that includes the mouth and is farther away than the mouth. <5> The information processing device according to <1>, further including an image processing unit that generates an image in which the user's mouth movements and hand movements are reflected in an avatar corresponding to the user from the images captured by the two imaging units. <6> The information processing device according to <5>, further including: a preprocessing unit that detects the mouth region and the user's hand region from the images captured by the two imaging units; and a recognition unit that recognizes the mouth movements and hand movements based on information about the mouth region and the user's hand region detected by the preprocessing unit, wherein the image processing unit generates an image in which the mouth movements and hand movements recognized by the recognition unit are reflected in an avatar corresponding to the user. <7> The information processing device described in <6>, wherein the preprocessing unit detects the mouth area and the user's hand area from the images captured by the two imaging units, and detects landmark coordinates for each of the detected mouth area and hand area, and the recognition unit recognizes the movement of the mouth and the movement of the hand based on the landmark coordinates of the mouth area and the hand area detected by the preprocessing unit.<8> The information processing device described in <7>, wherein the recognition unit recognizes the mouth movement and the hand movement based on the mouth region and the hand region in the image detected by the preprocessing unit, and the landmark coordinates of the mouth region and the hand region, respectively. <9> The information processing device described in <6>, wherein the recognition unit recognizes the mouth movement as a unit of action units and recognizes the hand movement as a pose or a gesture based on information of the mouth region and the hand region in the image detected by the preprocessing unit. <10> The information processing device described in <6>, wherein the image processing unit generates an image in which the mouth movement and the hand movement recognized by the recognition unit are reflected in an avatar corresponding to the user, using a generative model. <11> The information processing device of <6>, further including a focus adjustment mechanism that adjusts a focal length of the imaging unit, and an actuator that adjusts an angle of view of the imaging unit, wherein when the pre-processing unit cannot detect the mouth region and the hand region from the image, the focus adjustment mechanism adjusts the focal length of the imaging unit and the actuator adjusts the angle of view of the imaging unit until the mouth region and the hand region can be detected. <12> The information processing device of <11>, further including a calibration processing unit that causes the focus adjustment mechanism to adjust the focal length of the imaging unit in the focus adjustment mechanism and the actuator to perform a calibration process that adjusts the angle of view of the imaging unit in the actuator, so that the pre-processing unit can detect the mouth region and the hand region from the image.<14> The information processing device according to <13>, wherein the calibration processing unit, when executing the calibration process, requests the user to take a predetermined pose required for the calibration process, and causes the preprocessing unit to execute a calibration process in which the focus adjustment mechanism adjusts a focal length of the imaging unit and the actuator adjusts the angle of view of the imaging unit so that the mouth area and the hand area can be detected from the image while the user is taking the predetermined pose. <15> The information processing device according to <14>, wherein the predetermined poses are a pose in which the user's hands are spread out to the sides and a pose in which the hands are thrust forward. <16> The information processing device described in <13>, wherein the calibration processing unit stores, as calibration data, a control parameter resulting from adjustment of the focal length of the image capture unit in the focus adjustment mechanism and a control parameter resulting from adjustment of the angle of view of the image capture unit in the actuator when the calibration processing is completed, and when imaging by the image capture unit starts, reads out the calibration data and supplies control parameters corresponding to the focus adjustment mechanism and the actuator, thereby adjusting the focal length of the image capture unit in the focus adjustment mechanism and adjusting the angle of view of the image capture unit in the actuator so that the mouth area and the hand area can be detected from the image by the pre-processing unit. <17> The information processing device described in <16>, wherein the calibration data is stored in association with the user. <18> An information processing system comprising: two image capture units near a user's nose tip, at corresponding positions on the left and right of the nose tip, which capture an image of an area below the nose tip; the two image capture units each capture an area including the user's mouth at different angles of view, and the different angles of view intersect.

[0259] 101, 101', 101'' HMD, 111L, 111R, 111'L, 111'R Camera, 112 Display, 111'L-AF, 111'R-AF Focus adjustment mechanism, 151 Pre-processing unit, 152 Recognition unit, 153 Image processing unit, 201L, 201R Actuator, 221 AF control unit, 222 Actuator control unit, 72 Delay time acquisition unit, 251 Calibration processing unit, 261 Personal data

Claims

1. An information processing apparatus comprising two imaging units that image an image below the tip of the nose in the vicinity of the tip of the user's nose, at corresponding positions on the left and right centered on the tip of the nose, wherein the two imaging units image the periphery including the user's mouth area at different viewing angles, and the different viewing angles intersect each other.

2. The information processing apparatus according to claim 1, wherein for the different viewing angles, one viewing angle images the mouth area from the left side, and for a range farther from the mouth area, includes a range on the right side with respect to the front direction of the user, and the other viewing angle images the mouth area from the right side, and for a range farther from the mouth area, includes a range on the left side with respect to the front direction of the user, such that the left and right intersect.

3. The information processing apparatus according to claim 2, wherein the left end portion of one viewing angle overlaps with the right end portion of the other viewing angle.

4. The information processing apparatus according to claim 1, wherein at least one of the two imaging units includes the mouth area and images the user's hand in a range farther from the mouth area.

5. The information processing apparatus according to claim 1, further including an image processing unit that generates an image in which the movements of the user's mouth and hand are reflected in an avatar corresponding to the user, based on the images captured by the two imaging units.

6. The information processing apparatus according to claim 5, further including a preprocessing unit that detects the area of the mouth area and the area of the user's hand from the images captured by the two imaging units, and a recognition unit that recognizes the movement of the mouth and the movement of the hand based on the information of the area of the mouth area and the area of the user's hand detected by the preprocessing unit, wherein the image processing unit generates an image in which the movement of the mouth and the movement of the hand recognized by the recognition unit are reflected in an avatar corresponding to the user.

7. The information processing apparatus according to claim 6, wherein the preprocessing unit detects the area of the mouth area and the area of the user's hand from the images captured by the two imaging units, and detects landmark coordinates for each of the detected area of the mouth area and the area of the hand, and the recognition unit recognizes the movement of the mouth and the movement of the hand based on the landmark coordinates of each of the area of the mouth area and the area of the hand detected by the preprocessing unit.

8. The recognition unit recognizes the movement of the mouth and the movement of the hand based on the region of the mouth and the region of the hand in the image detected by the preprocessing unit, and the landmark coordinates of the region of the mouth and the region of the hand, respectively. The information processing apparatus according to claim 7.

9. The recognition unit recognizes the movement of the mouth in units of action units and recognizes the movement of the hand as a pose or gesture based on the information of the region of the mouth and the region of the hand in the image detected by the preprocessing unit. The information processing apparatus according to claim 6.

10. The image processing unit generates an image in which the movement of the mouth and the movement of the hand recognized by the recognition unit are reflected in the avatar corresponding to the user, using a generation model. The information processing apparatus according to claim 6.

11. The information processing apparatus according to claim 6, further comprising a focus adjustment mechanism for adjusting the focal length of the imaging unit and an actuator for adjusting the angle of view of the imaging unit. When the preprocessing unit cannot detect the region of the mouth and the region of the hand from the image, the focus adjustment mechanism adjusts the focal length of the imaging unit until the region of the mouth and the region of the hand can be detected, and the actuator adjusts the angle of view of the imaging unit.

12. When the preprocessing unit cannot detect the region of the mouth and the region of the hand from the image, the image processing unit generates an image in which the movement of the mouth and the movement of the hand of the user in the immediately previous image in which the region of the mouth and the region of the hand are detected are reflected in the avatar corresponding to the user. The information processing apparatus according to claim 11.

13. The information processing apparatus according to claim 11, further comprising a calibration processing unit that executes calibration processing for adjusting the focal length of the imaging unit in the focus adjustment mechanism and adjusting the angle of view of the imaging unit in the actuator so that the preprocessing unit can detect the region of the mouth and the region of the hand from the image.

14. The calibration processing unit, when executing the calibration processing, requests the user to take a predetermined pose required for the calibration processing, and while the user is taking the predetermined pose, the focus adjustment mechanism adjusts the focal length of the imaging unit so that the preprocessing unit can detect the area of the mouth and the area of the hand from the image, and the calibration processing for adjusting the angle of view of the imaging unit in the actuator is executed. The information processing apparatus according to claim 13.

15. The predetermined pose is a pose in which the user's both hands are opened to the left and right and a pose in which the both hands are protruded forward. The information processing apparatus according to claim 14.

16. The calibration processing unit, when the calibration processing is completed, stores, as calibration data, a control parameter that is a result of adjusting the focal length of the imaging unit in the focus adjustment mechanism and a control parameter that is a result of adjusting the angle of view of the imaging unit in the actuator. When imaging by the imaging unit is started, the calibration data is read out and control parameters corresponding to the focus adjustment mechanism and the actuator are supplied, so that the preprocessing unit can detect the area of the mouth and the area of the hand from the image. The focus adjustment mechanism adjusts the focal length of the imaging unit, and the actuator adjusts the angle of view of the imaging unit. The information processing apparatus according to claim 13.

17. The calibration data is stored in association with the user. The information processing apparatus according to claim 16.

18. An information processing system comprising two imaging units that image an image below the tip of the nose in the vicinity of the tip of the user's nose at corresponding positions on the left and right centered on the tip of the nose, wherein the two imaging units image the periphery including the user's mouth at different angles of view, and the different angles of view intersect.

Citation Information

Patent Citations

  • Stereoscopic camera device and electronic information device

    JP2012053303A

  • Device operating system and device control method

    JP2020118024A

  • Image processing method and image processing program and image processing system

    JP2021128476A

  • Facial expression recognition system, facial expression recognition method, and facial expression recognition program

    WO2017122299A1